Search NASA⌕ Search

SEARCH · Search NASA

Results for “data generation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Spaceflight Biospecimen and Data Sharing in Support of Science Discovery and Exploration

For decades, NASA and international partners have conducted biological experiments in space to understand effects of spaceflight and address potential hazards. To enable spaceflight back to the Moon, and then to Mars and beyond, it is imperative to further understand basic science and health risks associated with spaceflight, along with developing countermeasures. The sending of experiments and organisms into space is a costly endeavor. To maximize scientific return, sharing with the scientific community both space-flown biospecimens and data from completed experiments is essential. New fundamental, applied, and bioinformatic science insights can be gained from specimen and data sharing efforts. Data reuse enables spaceflight health risk modeling, analyzing adverse outcomes across spaceflight hazards, and deep space autonomous support for the flight medical officer. Space-flown biospecimens not required by mission Principal Investigators are regularly archived and made available for scientific request. The largest biorepository of these samples are found within NASA’s Institutional Scientific Collection at Ames Research Center (ISC-ARC), which stores over 32,000 specimens mostly from Shuttle and International Space Station (ISS) missions, but also some ground-based analog samples. The Ames Life Sciences Data Archive manages the ISC-ARC. Tissues are predominantly from mice and rats, though samples are also available from bacteria and quail. Only a handful of other similar collections exist worldwide. Rodent biospecimens exposed to simulated space radiation at Brookhaven National Laboratory are archived under the purview of NASA HRP Space Radiation Element. Microbial collection and analyses from 20 years of routine environmental monitoring of air, surfaces, and water systems of the ISS were performed to ensure a safe environment for astronauts. Samples from the ISC-ARC, space radiation and microbial collections are searchable and requestable through the NASA Life Sciences Data Archive (LSDA). Decades of planetary protection microbial isolates derived from spacecraft bioburden are archived in JPL’s microbial collection. Rodent biospecimens from spaceflight investigations conducted by the Japan Aerospace Exploration Agency (JAXA) are archived and available at the JAXA Biorepository at Tsukuba Space Center. The Russian Institute of Biomedical Problems also has a collection of animal, microbial, cellular, and fungi available for research from ground analog experiments. Several data repositories exist for scientists to utilize. The LSDA is the primary NASA source of life sciences research data and information. It contains decades of spaceflight and ground-analog research involving human, microbial, cellular, plant, and animal subjects. Data is collected from NASA-funded investigations through the Human Research Program and the Space Biology Program. The NASA Lifetime Surveillance of Astronaut Health collects and grants access to clinical and occupational health monitoring data from astronauts, with a list and description of data collected available for request through the LSDA. NASA GeneLab at ARC collects genomic, transcriptomic, proteomic, and metabolomic data from any species. It is a repository and platform for collaborative open-science bioinformatic approaches. JAXA is establishing an ‘omics-based repository in collaboration with the Tohoku Medical Megabank (ToMMo), called the JAXA-ToMMo Integrated Biobank for Space Life Science. Overall, the sharing of these biospecimen and data resources can assist researchers worldwide in understanding spaceflight effects on biology, along with enabling next generation data science applications for space exploration platforms. Websites: https://lsda.jsc.nasa.gov/ ; https://www.nasa.gov/ames/research/space-biosciences/isc-bsp ; https://www.nasa.gov/ames/research/space-biosciences/alsda

Ryan T. Scott↗

20-Years of Atmospheric Temperature, Water Vapor, Cloud, and Surface Temperature Anomalies and Trends Derived From Operational Hyperspectral Ir Sounders

Hyperspectral IR sounders such as AIRS on Aqua, CrIS on S-NPP, NOAA20 and JPSS-2, IASI on Metop A, B, and C provide high-quality atmospheric temperature, water, vapor, and greenhouse gas vertical profiles. Additionally, they provide atmospheric cloud properties, surface emissivity, and surface skin temperatures. We have developed two algorithms which can consistently derive these products from multiple IR sounders. The first one is a Single Field-of-view Sounder Atmospheric Product (SIFSAP) algorithm and the second one is a Climate Fingerprinting Sounder Product (ClimFiSP) algorithm. Compared to current operational AIRS and CrIS Level-2 (L2) algorithms, which perform one retrieval for each 3 by 3 field of views (FOVs) using a cloud-clearing approach, the SiFSAP algorithm, on the other hand, performs one retrieval for each FOV using an all-sky optimal estimation approach. The SiFSAP algorithm retrieves all the above-mentioned atmosphere and surface properties simultaneously including cloud properties with 3-time higher spatial resolution and 9-times more products. The core of the SiFSAP algorithm is an accurate and fast Principal Component-based Radiative Transfer Model (PCRTM), which can calculate hyperspectral radiance spectra under both clear and cloudy conditions. The PCRTM was developed in the past decade using consistent reference line-by-line radiative transfer model and spectroscopy for hyperspectral sounders such as AIRS, CrIS, IASI, NAST-I, and S-HIS. The SiFSAP retrieval algorithm also uses the same climatology a priori and associated covariances, which makes it ideal for generating high quality products for both weather and climate applications. Climate products are typically derived by performing spatial and temporal averaging of L2 products. It is a time-consuming process to generate L2 data products since AIRS, CrIS, and IASI have millions of observations each day with thousands of spectral channels for each observation. Additionally, differences in L2 retrieval algorithms for different satellite sensors can lead to errors in the climate products. Our ClimFiSP algorithm, which performs retrievals from spatiotemporally averaged L1 hyperspectral radiances directly, will be orders of magnitude faster than traditional method. he ClimFiSP algorithm uses consistent radiative kernels and a robust spectral fingerprinting method. It provides accurate data climate data fusion products from multiple satellite sensors. We have applied this method to both AIRS and CrIS (on SNPP and on NOAA 20) data and generated two decades climate data records for atmospheric temperature, water vapor, cloud, trace gases, and surface skin temperature. Both SiFSAP and ClimFiSP will be available at NASA GES DISC data center for public access.

Xu Liu↗

Two Decades of Atmospheric and Surface Temperature, Water Vapor, and Cloud Trends Derived from Satellite Remote Sensors

Satellite remote sensor such as Atmospheric Infrared Sounder (AIRS), Cross-track Infrared Sounder (CrIS), and Infrared Atmospheric Sounding Interferometer (IASI) provide high-quality atmospheric temperature, water, vapor, and greenhouse gas vertical profiles. Additionally, they provide atmospheric cloud properties, surface emissivity, and surface skin temperatures. We have developed two algorithms which can consistently derive these products from multiple IR sounders. The first one is a Single Field-of-view Sounder Atmospheric Product (SIFSAP) algorithm and the second one is a Climate Fingerprinting Sounder Product (ClimFiSP) algorithm. Compared to current operational AIRS and CrIS Level-2 (L2) algorithms, which perform one retrieval for each 3 by 3 field of views (FOVs) using a cloud-clearing approach, the SiFSAP algorithm, on the other hand, performs one retrieval for each FOV using an all-sky optimal estimation approach. The SiFSAP algorithm retrieves all the above-mentioned atmosphere and surface properties simultaneously including cloud properties with 3-time higher spatial resolution and 9-times more products. Climate products are typically derived by performing spatial and temporal averaging of L2 products. It is a time-consuming process to generate L2 data products since AIRS, CrIS, and IASI have millions of observations each day with thousands of spectral channels for each observation. Additionally, differences in L2 retrieval algorithms for different satellite sensors can lead to errors in the climate products. Our ClimFiSP algorithm, which performs retrievals from spatiotemporally averaged L1 hyperspectral radiances directly, will be orders of magnitude faster than traditional method. he ClimFiSP algorithm uses consistent radiative kernels and a robust spectral fingerprinting method. It provides accurate data climate data fusion products from multiple satellite sensors. We have applied this method to both AIRS and CrIS (on SNPP and on NOAA 20) data and generated two decades climate data records for atmospheric temperature, water vapor, cloud, trace gases, and surface skin temperature. Both SiFSAP and ClimFiSP will be available at NASA GES DISC data center for public access.

remote sensing↗

Hybrid learning techniques for scientific data reduction with performance guarantees

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Final report- UFL - RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING↗

Scheduler software for tracking and data relay satellite system loading analysis: User manual and programmer guide

A user guide and programmer documentation is provided for a system of PRIME 400 minicomputer programs. The system was designed to support loading analyses on the Tracking Data Relay Satellite System (TDRSS). The system is a scheduler for various types of data relays (including tape recorder dumps and real time relays) from orbiting payloads to the TDRSS. Several model options are available to statistically generate data relay requirements. TDRSS time lines (representing resources available for scheduling) and payload/TDRSS acquisition and loss of sight time lines are input to the scheduler from disk. Tabulated output from the interactive system includes a summary of the scheduler activities over time intervals specified by the user and overall summary of scheduler input and output information. A history file, which records every event generated by the scheduler, is written to disk to allow further scheduling on remaining resources and to provide data for graphic displays or additional statistical analysis.

Craft, R.↗

Data efficiency assessment of generative adversarial networks in energy applications

This study investigates the data requirements of generative artificial intelligence (AI), particularly generative adversarial networks (GANs), for reliable data augmentation in energy applications. Generative AI, though seen as a solution to data limitations, requires substantial data to learn meaningful distributions—a challenge often overlooked. This study addresses the challenge through synthetic data generation for critical heat flux (CHF) and power grid demand, focusing on renewable and nuclear energy. Two variants of GAN employed are conditional GAN (cGAN) and Wasserstein GAN (wGAN). Our findings include the strong dependency of GAN on data size, with performance declining on smaller datasets and varying performance when generalizing to unseen experiments. Mass flux and heated length significantly influence CHF predictions. wGAN is more robust to feature exclusion, making it suitable for constrained synthetic data generation. In energy demand forecasting, wGAN performed well for solar, wind, and load predictions. Longer lookback hours and larger datasets improved predictions, especially for load power. Seasonal variations posed challenges, with wGAN achieving a relatively high error of Root Mean Squared Error (RMSE) of 0.32 for load power prediction, compared to RMSE of 0.07 under same-season conditions. Feature exclusions impacted cGAN the most, while wGAN showed greater robustness. This study concludes that, while generative AI is effective for data augmentation, it requires substantial data and careful training to generate realistic synthetic data and generalize to new experiments in engineering applications.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Plutonium isotope ratio measurements by total evaporation-thermal ionization mass spectrometry (TE-TIMS): an evaluation of uncertainties using traceable standards from the New Brunswick Laboratory

The accuracy and precision of isotope amount ratio measurements using thermal ionization mass spectrometry (TIMS) instrumentation are described, and the measurement of Pu materials is emphasized. The mass fractionation observed for Am, Ga, Pu, and U for isotope amount ratio measurements using the total evaporation (TE) technique is compared with theoretical estimates to demonstrate the advantage of the TE methodology and to investigate systematic biases in the major isotope amount ratios of U and Pu certified reference material (CRM) standards from the U.S. provider of CRMs. The quality of the Pu isotopic data generated by TIMS instruments in an analytical laboratory is demonstrated by the application of the double ratio technique to estimate the 241 Pu half-life. Analytical data on traceable Pu CRMs from the New Brunswick Laboratory (NBL), generated as part of routine measurements supporting various programs, are used for this half-life estimation. Although the 241 Pu abundances in CRMs of 136, 137, 138, and 126-A are approximately 200–2000× smaller than those in the 241 Pu material used in the previous Institute for Reference Materials and Measurements (IRMM) evaluation of the 241 Pu half-life, the half-life value estimated in this work shows excellent agreement with the currently accepted value from the IRMM. This agreement also demonstrates the pedigree of the Pu isotopic standards from the NBL and the quality of the isotope amount ratio measurements using TIMS instrumentation. For both the major and minor Pu isotope amount ratios, this report describes the relative importance of the factors affecting the uncertainty of TIMS measurements, which are considered the gold standard in isotope ratio measurements (LA-UR-24-29199).

07 ISOTOPE AND RADIATION SOURCES↗

Quantitative measurements of Jupiter, Saturn, their rings and satellites made from Voyager imaging data

The Voyager spacecraft cameras use selenium-sulfur slow scan vidicons to convert focused optical images into sensible electrical signals. The vidicon-generated data thus obtained are the basis of measurements of much greater precision than was previously possible, in virtue of their superior linearity, geometric fidelity, and the use of in-flight calibration. Attention is given to positional, radiometric, and dynamical measurements conducted on the basis of vidicon data for the Saturn rings, the Saturn satellites, and the Jupiter atmosphere.

Collins, S. A.↗

Sister Rod Destructive Examinations (FY23) Appendix B: Segmentation, Defueling, Metallographic Data and Total Cladding Hydrogen

As a part of the DOE-NE High Burnup Spent Fuel Data Project, Oak Ridge National Laboratory (ORNL) is performing destructive examinations (DEs) of high burnup (HBU) (>45 GWd/MTU) spent nuclear fuel (SNF) rods from the North Anna Nuclear Power Station operated by Dominion Energy. The SNF rods, called sister rods or sibling rods are all HBU and include four different kinds of fuel rod cladding: standard Zircaloy-4 (Zirc-4), low-tin (LT) Zirc-4, ZIRLO ® , and M5 ® . The DEs are being conducted to obtain a baseline of the HBU rod’s condition before dry storage and are focused on understanding overall SNF rod strength and durability. Both composite fuel and defueled cladding will be tested to derive material properties. Although the data generated can be used for multiple purposes, one primary goal for obtaining the post-irradiation examination data and associated measured mechanical properties is to support SNF dry storage licensing and relicensing activities by (1) addressing identified knowledge gaps and (2) enhancing the technical basis for post-storage transportation, handling, and subsequent disposition of the SNF. This report documents the status of the ORNL Phase 1 DE activities related to: Rough segmentation (RS), Defueling (DEF), DE.02 optical microscopy (MET), and DE.03, cladding total hydrogen measurements. It is a cumulative update to the FY22 status report.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Developing ML/AI Methods for High-Throughput Characterization of Multiple-Sensor Streams of Tokamak Dynamics for High-Speed Control (Final Report)

This project evaluated and developed new mathematical and algorithmic techniques capable of handling (in real-time) the growing amounts of data generated by modern fusion research. While existing numerical linear algebra (NLA) methods provide the backbone to classical data analysis and algorithms, these methods fundamentally do not port to distributed architectures nor do they allow low-latency data reduction for control. Motivated by the needs for modern fusion reactors, this project explored and implemented new numerical methods to characterize plasma dynamics, respond in real-time to discharge evolution, and to process massive-scale data accurately and rapidly more fully. This project links expertise in multiple-sensor diagnostics of tokamak plasma dynamics from Columbia University’s Plasma Physics Laboratory with expertise in massive-scale data reduction and extreme data control algorithms at Columbia University’s Data Science Institute. This interdisciplinary project (i) applied machine learning methods, (ii) implemented a properly-trained neural-network for very fast processing of high-speed plasma videography, and (ii) developed the applied mathematical methods, based on randomized-NLA (rNLA) routines, for data analysis, reduction, and real-time control. The Columbia University High Beta Tokamak-Extended Pulse (HBT-EP) facility provided data to test new algorithms and partnership with Columbia University's Data Sciences Institute evaluated the broader use of new algorithms for many challenging control applications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Payload/orbiter contamination control requirement study, volume 2, exhibit A

The computer printout data generated during the Payload/Orbiter Contamination Control Requirement Study are presented. The computer listings of the input surface data matrices, the viewfactor data matrices, and the geometric relationship data matrices for the three orbiter/spacelab configurations analyzed in this study are given. These configurations have been broken up into the geometrical surfaces and nodes necessary to define the principal critical surfaces whether they are contaminant sources, experimental surfaces, or operational surfaces. A numbering scheme was established based upon nodal numbers that relates the various spacelab surfaces to a specific surface material or function. This numbering system was developed for the spacelab configurations such that future extension to a surface mapping capability could be developed as required.

Bareiss, L. E.↗

Retrieval of ice thickness from polarimetric SAR data

We describe a potential procedure for retrieving ice thickness from multi-frequency polarimetric SAR data for thin ice. This procedure includes first masking out the thicker ice types with a simple classifier and then deriving the thickness of the remaining pixels using a model-inversion technique. The technique used to derive ice thickness from polarimetric observations is provided by a numerical estimator or neural network. A three-layer perceptron implemented with the backpropagation algorithm is used in this investigation with several improved aspects for a faster convergence rate and a better accuracy of the neural network. These improvements include weight initialization, normalization of the output range, the selection of offset constant, and a heuristic learning algorithm. The performance of the neural network is demonstrated by using training data generated by a theoretical scattering model for sea ice matched to the database of interest. The training data are comprised of the polarimetric backscattering coefficients of thin ice and the corresponding input ice parameters to the scattering model. The retrieved ice thickness from the theoretical backscattering coefficients is compare with the input ice thickness to the scattering model to illustrate the accuracy of the inversion method. Results indicate that the network convergence rate and accuracy are higher when multi-frequency training sets are presented. In addition, the dominant backscattering coefficients in retrieving ice thickness are found by comparing the behavior of the network trained backscattering data at various incidence angels. After the neural network is trained with the theoretical backscattering data at various incidence anges, the interconnection weights between nodes are saved and applied to the experimental data to be investigated. In this paper, we illustrate the effectiveness of this technique using polarimetric SAR data collected by the JPL DC-8 radar over a sea ice scene.

Kwok, R.↗

Data-Driven Software Framework for Web-Based ISS Telescience

Software that enables authorized users to monitor and control scientific payloads aboard the International Space Station (ISS) from diverse terrestrial locations equipped with Internet connections is undergoing development. This software reflects a data-driven approach to distributed operations. A Web-based software framework leverages prior developments in Java and Extensible Markup Language (XML) to create portable code and portable data, to which one can gain access via Web-browser software on almost any common computer. Open-source software is used extensively to minimize cost; the framework also accommodates enterprise-class server software to satisfy needs for high performance and security. To accommodate the diversity of ISS experiments and users, the framework emphasizes openness and extensibility. Users can take advantage of available viewer software to create their own client programs according to their particular preferences, and can upload these programs for custom processing of data, generation of views, and planning of experiments. The same software system, possibly augmented with a subset of data and additional software tools, could be used for public outreach by enabling public users to replay telescience experiments, conduct their experiments with simulated payloads, and create their own client programs and other custom software.

Tso, Kam S.↗

Joint ESA-NASA Multi-Mission Algorithm and Analysis Platform (MAAP)

The scientific community is faced with a need for greatly improved data sharing, analysis, visualization and advanced collaboration based firmly on open science principles. Recent and upcoming launches of new satellite missions with more complex and voluminous data, as well as the ever more urgent need to better understand the global carbon budget and related ecological processes, provided the immediate rational for the ESA-NASA Multi-mission Algorithm and Analysis Platform (MAAP). This highly collaborative joint project of ESA and NASA established a framework between ESA and NASA to share data, science algorithms and compute resources in order to foster and accelerate scientific research conducted by ESA and NASA EO data users. Presented to the public in October 2021, the current version of MAAP provides a common cloud-based platform with computing capabilities co-located with the data, a collaborative coding and analysis environment, and a set of interoperable tools and algorithms developed to support the estimation and visualization of global above-ground biomass. Data from the Global Ecosystem Dynamics Investigation (GEDI) mission on the International Space Station and the Ice, Cloud, and Land Elevation Satellite-2 (ICESat-2) have been instrumental in the first products of MAAP including the first comprehensive map of Boreal above-ground Biomass and a current Global Biomass Harmonization Activity, but the platform is also being specifically designed to support the forthcoming ESA Biomass mission and incorporate data from the upcoming NASA-ISRO SAR (NISAR) mission. While these missions and the corresponding research which includes airborne, field, and calibration/validation data collection and analyses, provide a wealth of data and information relating to global biomass estimation, they also present data storing, processing and sharing challenges. The NISAR mission alone will produce about 80TB/day. These large data volumes present a challenge that would otherwise place accessibility limits on the scientific community and impact scientific progress. Other challenges being addressed by MAAP include: 1) Enabling researchers to easily discover, process, visualize and analyze large volumes of data from both agencies; 2) Providing a wide variety of data in the same coordinate reference frame to enable comparison, analysis, data evaluation, and data generation; 3) Providing a version-controlled science algorithm development environment that supports tools, co-located data and processing resources; and 4) Addressing intellectual property and sharing challenges related to collaborative algorithm development and sharing of data and algorithms. MAAP products can be explored on the MAAP Dashboard at https://earthdata.nasa.gov/maap-biomass or the joint platform entrance at scimaap.net. MAAP also can be accessed through individual NASA (https://maap-project.org) and ESA (https://esa-maap.org/) landing pages.

cloud computing↗

Why We Do What We Do: Data Reuse, Open Access, and Privacy in Data Management at the Life Sciences Data Archive

As custodian of the unique and irreplaceable collections of human subject research data generated by the Human Research Program and its predecessors throughout the agency’s history, the Life Sciences Data Archive (LSDA) is charged with protecting participants’ privacy and implementing their consent decisions as it provides retrospective data for use in new studies. This active, stewardship-focused approach to data management and preservation shapes the products that LSDA provides to researchers and the responsibilities of researchers in using the data and publishing their results. This presentation reviews how federal and agency mandates shape LSDA’s data management procedures and expectations for researchers. Topics covered will include LSDA’s movement towards implementation of the FAIR (Findable, Accessible, Interoperable, Reusable) principles and how the archive’s evolving data management practices support FAIR-ness; collaboration between LSDA and the Lifetime Surveillance of Astronaut Health (LSAH) project (the repository of astronaut medical data); LSDA’s response to the challenges of performing its stewardship role and maintaining trust given the public profiles of the subjects whose data it preserves; and the ever-increasing challenges to expectations of subject privacy stemming from the growing power and ubiquity of of data analysis and aggregation tools.

Data↗

Why We Do What We Do: Data Reuse, Open Access and Privacy in Data Management at the Life Sciences Data Archive

As custodians of the unique and irreplaceable collections of human subject research data generated by the Human Research Program and its predecessors throughout the agency’s history, the Life Sciences Data Archive (LSDA) is charged with protecting participants’ privacy and implementing their consent decisions as it provides retrospective data for use in new studies. This active, stewardship-focused approach to data management and preservation shapes the products that LSDA provides to researchers and the responsibilities of researchers in using the data and publishing their results. This presentation reviews how federal and agency mandates shape LSDA’s data management procedures and expectations for researchers. Topics covered will include LSDA’s movement towards implementation of the FAIR (Findable, Accessible, Interoperable, Reusable) principles and the archive’s evolving data management practices; collaboration between LSDA and the Lifetime Surveillance of Astronaut Health (LSAH) project (the repository of astronaut medical data); LSDA’s response to the challenges of performing its stewardship role and maintaining trust given the public profiles of the subjects whose data it preserves; and the ever-increasing challenges to expectations of subject privacy stemming from the growing power and ubiquity of data analysis and aggregation tools.

data management↗

Results of software error-data experiments

In order to evaluate existing software reliability models and proposed modeling approaches, a search was conducted for data on the software failure process. This search revealed that the data necessary for this evaluation were not available. As a result, a research effort was initiated by NASA to generate data on which to base the development of credible methods for assessing the reliability of software targeted for flight-crucial applications. Two sets of software error-data experiments were conducted by different research groups. The results of the experiments were consistent: errors caused by different faults in a program occurred at widely varying rates; program failure rates exhibited a log-linear trend with respect to the number of faults corrected; some faults were found to interact in either concealing or revealing ways; and contiguous regions of the input space which cause a program to generate errors, called error crystals, were found and characterized for some faults. Collectively, these experiments have produced information on software failure which must be accounted for in software reliability modeling approaches.

Finelli, George B.↗