Search NASA⌕ Search

SEARCH · Search NASA

Results for “data generation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

RC-SFA Data Management Templates and Guidance for Standardized, Reusable AI-Ready Data Packages

This data package provides templates and supporting documentation developed by the River Corridor Science Focus Area (RC-SFA; https://www.pnnl.gov/projects/river-corridor) to communicate its approach to managing and publishing AI-ready data. The package is intended to help data users and data producers understand the structures, metadata practices, and quality-control approaches that support consistent, reusable, and machine-actionable data products across RC-SFA studies. Rather than focusing on a single experimental dataset, this package documents the data management framework used to make RC-SFA data easier to find, ingest, navigate, and interpret. The materials in this package reflect RC-SFA practices for standardized data package organization, including the use of a human- and machine-readable README, file-level metadata, data dictionaries, descriptive file naming, method identifiers, and automated and review-based quality assurance procedures. Together, these components illustrate how RC-SFA extends FAIR data principles toward AI-readiness by prioritizing deep metadata, consistency across data packages, and support for informed downstream reuse by both humans and computational tools. This dataset is comprised of (1) readme; (2) presentation slides with an overview of RC-SFA approach and guidance; (3) document of RC-SFA best practices; (4) data dictionary (dd); (5) file level metadata (flmd); and a subfolder containing templates for dd and flmd. All files are .csv and .pdf. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.

AI-readiness↗

Imaging systems and related methods including radar imaging with moving arrays or moving targets

Imaging systems, including radio frequency, microwave and millimeter-wave arrangements, and related methods are described. According to one aspect, an imaging system includes an antenna array, a position capture system configured to generate position information indicative of locations of one of the antenna array and the target at the first and second moments in time, and wherein the one of the antenna array and the target move between the first and second moments in time, a transceiver configured to control the antenna array to emit electromagnetic energy towards the target and to generate an output that is indicative of the received electromagnetic energy, a data acquisition system configured to generate radar data, processing circuitry configured to process the position information and the radar data to generate image data regarding the target, and an interface configured to use the image data to generate visual images regarding the target.

Sheen, David M.↗

Deciphering the Scattering of Mechanically Driven Polymers Using Deep Learning

Here, we present a deep learning approach for analyzing two-dimensional scattering data of semiflexible polymers under external forces. In our framework, scattering functions are compressed into a three-dimensional latent space using a Variational Autoencoder (VAE), and two converter networks establish a bidirectional mapping between the polymer parameters (bending modulus, stretching force, and steady shear) and the scattering functions. The training data are generated using off-lattice Monte Carlo simulations to avoid the orientational bias inherent in lattice models, ensuring robust sampling of polymer conformations. The feasibility of this bidirectional mapping is demonstrated by the organized distribution of polymer parameters in the latent space. By integrating the converter networks with the VAE, we obtain a generator that produces scattering functions from given polymer parameters and an inferrer that directly extracts polymer parameters from scattering data. While the generator can be utilized in a traditional least-squares fitting procedure, the inferrer produces comparable results in a single pass and operates 3 orders of magnitude faster. This approach offers a scalable automated tool for polymer scattering analysis and provides a promising foundation for extending the method to other scattering models, experimental validation, and the study of time-dependent scattering data.

Ding, Lijie [Oak Ridge National Laboratory (ORNL),↗

When more data hurts: Optimizing data coverage while mitigating diversity-induced underfitting in an ultrafast machine-learned potential

Machine-learned interatomic potentials (MLIPs) are becoming an essential tool in materials modeling. However, optimizing the generation of training data used to parametrize the MLIPs remains a significant challenge. This is because MLIPs can fail when encountering local environments too different from those present in the training data. The difficulty of determining a priori the environments that will be encountered during molecular dynamics simulation necessitates diverse, high-quality training data. Here, this study investigates how training data diversity affects the performance of MLIPs using the Ultra-Fast force field (UF 3 ) to model amorphous silicon nitride. We employ expert and autonomously generated data to create the training data and fit four force field variants to subsets of the data. Our findings reveal a critical balance in training data diversity: insufficient diversity hinders generalization, while excessive diversity can exceed the MLIP's learning capacity, reducing simulation accuracy. Specifically, we found that the UF 3 variant trained on a subset of the training data, in which nitrogen-rich structures were removed, offered vastly better prediction and simulation accuracy than any other variant. By comparing these UF 3 variants, we highlight the nuanced requirements for creating accurate MLIPs, emphasizing the importance of application-specific training data to achieve optimal performance in modeling complex material behaviors.

ab initio molecular dynamics↗

Downscaled Earth System Model Data for Resilient Energy System Planning

The second-generation Sup3rCC dataset provides high-resolution meteorological data generated through the downscaling of multiple earth system models (ESMs) from the Coupled Model Intercomparison Project Phase 6 (CMIP6). This downscaling is performed through application of a generative machine learning approach called Super-Resolution for Renewable Resource Data (sup3r). This dataset builds on the first-generation Sup3rCC data by applying improved bias correction methods and adding downscaled precipitation to the output variables. In this presentation, we explore the output characteristics of the dataset and various validation analyses. We also present and discuss plans for the integration of this data into power system planning models using a decision-making under deep uncertainty (DMDU) methodology.

97 MATHEMATICS AND COMPUTING↗

Manuscript Workflows from and Processed Organic Matter Composition of Experimentally Burned Open Air and Muffle Furnace Vegetation Chars across Differing Burn Severity and Feedstock Types from Pacific Northwest, USA (v3)

This dataset includes processed organic matter chemistry data from an experimental study designed to compare how the chemical composition of organic matter changes across different burn conditions and vegetation materials representative of major land cover types of the Pacific Northwest, USA. Chars were created in a closed muffle furnace or on an open burn table from four different feedstock species representing vegetation commonly impacted by fire regimes across the Pacific Northwest, USA. Source data and associated metadata (including methods and geospatial information) can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1894135 (Grieger et al. 2022). This dataset provides processing scripts and processed data for both solid and dissolved phase organic matter characterization data from experimentally generated chars. These processed data can be used to compare how different burn conditions may influence resultant organic matter chemistry and help further our understanding of potential biogeochemical impacts on river corridors post-fire. The processed data were subsequently analyzed; and the results and ecological implications of the findings were published in peer-reviewed manuscripts. The scripts and workflows used to develop the manuscripts are also included in this data package.This data package was originally published June 2024. It was updated September 2024 (new and modified files) and in January 2025 (modified files). See the change history section in the readme for more details.This dataset is comprised of one data package readme, one data dictionary (dd), one file level metadata (flmd), and folders containing (A) processed data; (B) general processing scripts; and (C) additional folders with specific manuscript analysis scripts and processed data. Step-by-step instructions to assist the user in recreating the workflow used to generate the results in the manuscripts is also provided. The processed data folder includes (1) a folder of processed Parallel Factor Analysis (PARAFAC) and spectra indices outputs from excitation emissions matrix (EEM) fluorescence and absorbance data; (2) a folder of processed solid state carbon-13 (13-C NMR) integrals; (3) folder of high resolution characterization of organic matter via 21 Tesla Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) generated through the Environmental Molecular Sciences Laboratory (EMSL; https://www.pnnl.gov/environmental-molecular-sciences-laboratory) processed data outputs from Formultitude (https://github.com/PNNL-Comp-Mass-Spec/Formultitude), blank corrections and data aggregation, and calculated molecular indices. All files are .pdf, .csv, .html, .Rmd, .R, or .RData.

54 ENVIRONMENTAL SCIENCES↗

WHONDRS River Corridor Surface Water Metabolites and Geochemistry from Global Sites

This dataset supports a broader study examining the character of organic matter that may be delivered to subsurface sediments via hydrologic exchange. To implement the global survey, free stream sampling kits were provided to interested volunteers throughout the world. Samples were collected with minimal constraints in terms of location, but following strict protocols, and shipped for metabolomic analysis via Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS). In addition, basic geochemistry analyses (e.g., dissolved organic matter concentration) were conducted, standardized photos of each field system were taken, and extensive metadata were captured. Sampling began in 2018 and is ongoing as of 2025. This dataset is comprised of one folders of field photos, one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data, and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) international generic sample number (IGSN) mapping file; (6) field protocol; and (7) a subfolder with sample data. The sample data subfolder contains (1) surface water dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) methods codes; (3) surface water FTICR methods; and (4) a subfolder of 12 Tesla (12T) FTICR-MS data. This folder contains three subfolders, one containing the.xml files, one containing the CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, or .png. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.

Biogeochemistry↗

An ultra-fast method for generating synthetic down-scattered neutron data for inertial confinement fusion implosions

In inertial confinement fusion experiments at the National Ignition Facility, asymmetries are probed by a variety of neutron diagnostics, including neutron imaging systems, real-time neutron activation diagnostics (RTNADs), and neutron spectrometers. It is often useful to generate synthetic data based on these diagnostics to validate and tune models. However, current methods of doing so using Monte Carlo particle tracing are time-consuming. In this paper, an ultra-fast method is presented for generating synthetic neutron images, RTNAD data, and spectrometry data using line integrals and 3D convolutions. While it does not contain as much physics as particle tracing codes, it is thousands of times faster and produces nearly identical data. This enables analysis techniques that depend on generating large amounts of synthetic data, which will prove very useful for the study of asymmetries going forward.

Deuterium↗

Data for Mitochondrial ATP Generation is More Proteome Efficient than Glycolysis

Metabolic efficiency profoundly influences organismal fitness. Heterotrophs, from yeast to mammals, derive usable energy primarily through glycolysis and respiration. While respiration is more energy-efficient, some cells favor glycolysis even when oxygen is available (aerobic glycolysis, Warburg effect). A leading explanation is that glycolysis is more efficient in terms of ATP production per unit mass of protein (i.e. faster). Through quantitative flux analysis and proteomics, we find however that mitochondrial respiration is actually more proteome-efficient than aerobic glycolysis. This is shown across yeasts, T cells, cancer cells, and tissues and tumors in vivo. Instead of aerobic glycolysis being valuable for fast ATP production, it correlates with high glycolytic protein expression, which is valuable for hypoxic growth. Aerobic glycolytic yeasts do not excel at aerobic growth, but outgrow respiratory cells in oxygen limitation. Thus, aerobic glycolysis emerges from cells maintaining a proteome conducive to both aerobic and hypoxic growth.

Metabolomics↗

Using the power of secure Generative AI to eliminate data silos

Over time, multiple information repositories, access controls, and management practices created disparate information silos. Gaining accurate insights from lab information had become overly burdensome especially for new employees and collaborators

Purcell, Kevin↗

Alaska Observed Hydropower Generation

This dataset contains compiled observed hydropower generation for hydropower plants in Alaska. Data have been compiled from data provided to the Energy Information Administration by asset owners, data contained in annual reports produced by the Institute of Social and Economic Research at the University of Alaska Anchorage (Alaska Electric Power Statistics and Alaska Energy Statistics) and data provided to the Federal Energy Regulatory Commission by asset owners. This dataset provides available generation data from all sources in monthly and annual files, with quality flags, and generation data identifying the highest quality source in monthly and annual files.

hydropower datasets↗

Alaska Observed Hydropower Generation

This dataset contains compiled observed hydropower generation for hydropower plants in Alaska. Data have been compiled from data provided to the Energy Information Administration by asset owners, data contained in annual reports produced by the Institute of Social and Economic Research at the University of Alaska Anchorage (Alaska Electric Power Statistics and Alaska Energy Statistics) and data provided to the Federal Energy Regulatory Commission by asset owners. This dataset provides available generation data from all sources in monthly and annual files, with quality flags, and generation data identifying the highest quality source in monthly and annual files.

Broman, Daniel [Pacific Northwest National Laborat↗

A Nonergodic Ground-Motion Model for the San Francisco Bay Area for Small-Magnitude Earthquakes

ABSTRACT Recently, generative models have become a computationally efficient alternative to physics-based numerical simulations of ground motions. Neural networks can learn from existing ground-motion data to generate unobserved ground-motion data at new source and site locations. A key challenge with generative models is ensuring that predicted ground motions remain within a physically realistic range. For this purpose, we developed an empirical, nonergodic ground-motion model (GMM) for small-magnitude earthquakes in the San Francisco Bay area based on about 5000 recordings per component for Mw ≤ 4 earthquakes. The nonergodic GMM predicts spatially varying median source, site, and path effects for both the Fourier amplitude spectrum (FAS) and the Fourier phase derivative (a proxy for duration), as well as the corresponding epistemic uncertainty for each term. For FAS, our model shows above-average source and site effects in the western part of the region and below-average effects in the eastern part, with regional effects exhibiting larger spatial correlation lengths with increasing frequency. For duration, the source term is negligible for small-magnitude earthquakes, and the site term leads to site-specific variations up to 5 s. Path effects for FAS and duration depend on the source–site pair and are extrapolated spatially using recent methods for path-effect modeling. The aleatory variability of the within-site within-path residuals is similar to the variability found in previous studies for other regions. The nonergodic model provides two key contributions: first, median adjustment terms that are transferable to larger magnitude earthquakes, further reducing aleatory variability in probabilistic seismic hazard analysis; second, region-specific criteria for validating machine learning-based ground-motion generators to evaluate whether synthetic ground motions exhibit physically realistic source, site, and path effects.

Lacour, Maxime↗