Search NASA⌕ Search

SEARCH · Search NASA

Results for “data generation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Merged Aerosol Value-Added Product Report

The Merged Aerosol Value-Added Product (VAP) simplifies scientists’ use of Atmospheric Radiation Measurement (ARM) User Facility aerosol data by performing several tedious, time-consuming tasks for the users. First, the VAP identifies the best data available when multiple datastreams exist for a single geophysical quantity so that ARM users do not have to research this for themselves. Second, the VAP consolidates multiple ARM aerosol datastreams into a single file for ARM data users so that they do not have to download, open, and read multiple files for their analysis. Next, the VAP transforms all measurements onto a common one-hour timestamp. The one-hour resolution matches the time resolution of the slowest instrument. Instruments with faster sampling rates than one measurement per hour are averaged over the time interval. Finally, the VAP reads the QA/QC variables and marks data with known issues as missing, so that users do not have to spend excessive time cleaning data. This includes incorporating Data Quality Reports (DQRs) that exist at the time when the VAP data is generated. DQRs are reports filed by instrument mentors or data users that indicate a problem with the output data of individual instruments.

54 ENVIRONMENTAL SCIENCES↗

Generation of Enrichment-Dependent Thermal Neutron Scattering Data

This work details the generation of enrichment-dependent thermal neutron scattering cross sections for several crucial uranium fuel compounds. The evaluations of the thermal scattering law (TSL) and associated cross sections for uranium dioxide (UO 2 ), uranium carbide (UC), and uranium nitride (UN) were performed using standard ab initio lattice dynamics (AILD) methods. The data for uranium metal was produced using a novel hybrid approach of molecular dynamics combined with lattice dynamics methods. 235 U enrichments of 5%, 10% (LEU+), 19.75% (HALEU), 93% (HEU), and 100% were considered, in addition to natural uranium. The enrichment-dependent masses and free atom cross sections were used in the generation of elastic and inelastic thermal neutron scattering cross sections, while the calculation of the phonon density of states (DOS) and resulting TSL considered only the natural isotopic composition of uranium. The use of an identical DOS for all enrichments is expected to have minimal impact on the final data, as the small change in uranium mass should not significantly affect lattice vibrations. The cross sections are shown to exhibit significant dependence on 235 U enrichment. The submission of this data to the National Nuclear Data Center (NNDC) for release in the ENDF/B-VIII.1 database should support the design of advanced reactor concepts.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Generative learning for slow manifolds and bifurcation diagrams

In dynamical systems characterized by separation of time scales, the approximation of so called “slow manifolds”, on which the long term dynamics lie, is a useful step for model reduction. Initializing on such slow manifolds is a useful step in modeling, since it circumvents fast transients, and is crucial in multiscale algorithms (like the equation-free approach) alternating between fine scale (fast) and coarser scale (slow) simulations. In a similar spirit, when one studies the infinite time dynamics of systems depending on parameters, the system attractors (e.g., its steady states) lie on bifurcation diagrams (curves for one-parameter continuation, and more generally, on manifolds in state parameter space. Sampling these manifolds gives us representative attractors (here, steady states of ODEs or PDEs) at different parameter values. Algorithms for the systematic construction of these manifolds (slow manifolds, bifurcation diagrams) are required parts of the “traditional” numerical nonlinear dynamics toolkit. In more recent years, as the field of Machine Learning develops, conditional score-based generative models (cSGMs) have been demonstrated to exhibit remarkable capabilities in generating plausible data from target distributions that are conditioned on some given label. It is tempting to exploit such generative models to produce samples of data distributions (points on a slow manifold, steady states on a bifurcation surface) conditioned on (consistent with) some quantity of interest (QoI, observable). In this work, we present a framework for using cSGMs to quickly (a) initialize on a low-dimensional (reduced-order) slow manifold of a multi-time-scale system consistent with desired value(s) of a QoI (a “label”) on the manifold, and (b) approximate steady states in a bifurcation diagram consistent with a (new, out-of-sample) parameter value. This conditional sampling can help uncover the geometry of the reduced slow-manifold and/or approximately “fill in” missing segments of steady states in a bifurcation diagram. Finally, the quantity of interest, which determines how the sampling is conditioned, is either known a priori or identified using manifold learning-based dimensionality reduction techniques applied to the training data.

Dynamical systems↗

SDYN-GANs: Adversarial learning methods for multistep generative models for general order stochastic dynamics

We introduce adversarial learning methods for data-driven generative modeling of dynamics of nth-order stochastic systems. Our approach builds on Generative Adversarial Networks (GANs) with generative model classes based on stable m-step stochastic numerical integrators. From observations of trajectory samples, we introduce methods for learning long-time predictors and stable representations of the dynamics. Our approaches use discriminators based on Maximum Mean Discrepancy (MMD), training protocols using both conditional and marginal distributions, and methods for learning dynamic responses over different time-scales. We show how our approaches can be used for modeling physical systems to learn force-laws, damping coefficients, and noise-related parameters. Our adversarial learning approaches provide methods for obtaining stable generative models for dynamic tasks including long-time prediction and developing simulations for stochastic systems.

• Artificial intelligence (AI) / machine learning ↗

Pushing the limits of NAND technology scaling with ferroelectrics

Artificial intelligence (AI) continues to drive transformative advancements across various industries. The data-intensive nature of AI training (and inferencing) has resulted in the generation of unprecedented volumes of data with machine-generated content surpassing human-generated data by more than 100-fold in 2025. Efficiently managing this data influx necessitates advanced digital storage technologies. However, traditional NAND flash memory, which is critical for supporting data flows in AI systems—alongside high-bandwidth memory, for AI training—faces fundamental scaling limitations as it approaches the 1000-layer milestone, encompassing more than 40 trillion transistors. This article delves into the potential of hafnia-based ferroelectric materials as a breakthrough solution to these challenges. Recent advancements indicate that the intrinsic limitations of ferroelectric field-effect transistors (FEFETs) can be mitigated through material and device-level engineering. These advancements enable FEFETs to meet the stringent density, reliability, and scalability requirements of future three-dimensional NAND technology. The role of ferroelectrics in addressing NAND scaling challenges and expanding storage capabilities presents a promising avenue for meeting the storage demands of the AI-driven era.

3D NAND↗

OmicsMLMentor: A Web Application for Guided Machine Learning Analysis of Omics Data

Expression-based omics technologies (e.g. proteomics, metabolomics, transcriptomics, etc.) increasingly rely on supervised and unsupervised machine learning (ML) models to find key biomolecules distinguishing conditions, identify natural groupings in biological data, or generate predictions for outcomes of interest. Fitting ML models to omics data presents several challenges, including handling missing data, selecting a normalization method, choosing a valid model, and optimizing hyperparameters, all requiring statistical programming skills to address these challenges. Thus, the open-source web application SLOPE was designed to lower the barrier to ML modeling for omics data. SLOPE supports the fitting of 15 ML models (10 supervised and 5 unsupervised) tailored to omics datasets, such as proteomics, metabolomics, lipidomics, and transcriptomics. SLOPE offers several omics-specific features, including methods for handling missingness (imputation, conversion, removal), normalization tests, ranking of models based on the structure of a user’s data and user input, and optimal hyperparameter selections using cross-validation splits. By streamlining ML workflows for omics analysis, SLOPE address critical gaps in existing online web tools, facilitating a broader adoption of these models for omics research. Here, SLOPE is applied to data from a lignin exposure study to highlight the workflow for fitting both supervised and unsupervised models to data.

lipidomics↗

Simultaneous inference of equation of state parameters and unknown data errors with uncertainty quantification via hierarchical Bayesian posterior maximization

Equations of state (EOSs) are a key component in running hydrodynamic simulations as they relate the thermodynamic states for the material. The Davis reactants EOS is commonly used for modeling high explosives (HEs), and the EOS model parameters are calibrated using material specific data. The calibrations are often performed with uncertainty quantification via Bayesian inference to account for uncertainty in the data and generate ensembles of likely parameters. However, there are relatively few HE data sets to use for calibration and many are historical and lack error information. In this work, we simultaneously calibrate the Davis reactants EOS model parameters and unknown data error terms for the high explosive PBX 9501. To quantify the uncertainty in the models and the data, we use a Bayesian framework for the calibration and compute the hierarchical Bayesian posterior distribution with both a posteriori maximization approach and Markov Chain Monte Carlo. In general, we find that, given our assumptions, the two approaches result in similar calibrated parameters, posterior covariance matrices, and insights about the parameters but that the posterior maximization requires far less computational resources.

97 MATHEMATICS AND COMPUTING↗

LevSeq: Rapid Generation of Sequence-Function Data for Directed Evolution and Machine Learning

Sequence-function data provides valuable information about the protein functional landscape but is rarely obtained during directed evolution campaigns. Here, we present Long-read every variant Sequencing (LevSeq), a pipeline that combines a dual barcoding strategy with nanopore sequencing to rapidly generate sequence-function data for entire protein-coding genes. LevSeq integrates into existing protein engineering workflows and comes with open-source software for data analysis and visualization. The pipeline facilitates data-driven protein engineering by consolidating sequence-function data to inform directed evolution and provide the requisite data for machine learning-guided protein engineering (MLPE). LevSeq enables quality control of mutagenesis libraries prior to screening, which reduces time and resource costs. Simulation studies demonstrate LevSeq’s ability to accurately detect variants under various experimental conditions. Lastly, we show LevSeq’s utility in engineering protoglobins for new-to-nature chemistry. Widespread adoption of LevSeq and sharing of the data will enhance our understanding of protein sequence-function landscapes and empower data-driven directed evolution.

59 BASIC BIOLOGICAL SCIENCES↗

Transverse Kinematic Imbalance in MicroBooNE's New nue CC0pi Measurements

Neutrino-nucleus cross section measurements require accurate modelling of neutrino interactions. Neutrino beams are not monoenergetic, and the energy of each interaction must instead be modelled using nuclear interaction assumptions. This introduces significant systematic uncertainty into cross section measurements. Effects such as Fermi motion, nuclear correlations, and final-state interactions (FSI) smear the underlying quasi-elastic scattering signal, making it difficult to disentangle genuine quasi-elastic kinematics from nuclear effects across the full range of interaction channels (QE, MEC, RES, DIS) probed in these measurements. Transverse Kinematic Imbalance (TKI) variables, such as $\delta p_T$ and $\delta \alpha_T$, probe this same phase space by exploiting the fact that the incoming neutrino has zero transverse momentum ($\vec{p}_T^{\,\nu} = 0$). Any measured transverse imbalance in the final state therefore arises from nuclear effects rather than from uncertainty in the incident neutrino energy, allowing cross section measurements to select a phase space that is rich in quasi-elastic-like events with minimal contamination from FSI and other nuclear effects, independent of energy reconstruction. Recent unfolded MicroBooNE cross section measurements of electron-neutrino charged-current interactions with zero pions and at least one proton ($\nu_e$ CC0$\pi$, 1eNp0$\pi$) show that several leading nuclear interaction generators (including GENIE variants, NuWro, GiBUU, and NEUT) reproduce the differential cross section in electron energy reasonably well, but consistently struggle to describe the differential cross section in the cosine of the leading proton's angle, yielding lower $p$-values across all seven generators tested. This tension points to a more fundamental, kinematics-driven disagreement between data and generators that is not visible in energy-only cross section observables. This is precisely the regime TKI variables are designed to probe. Following previous TKI cross section measurements with muon-neutrino data in MicroBooNE, this poster presents the case for extending the TKI framework to electron-neutrino cross section measurements as a next step to isolate and characterize the source of the observed generator tension in proton kinematics.

Burridge, Jessica [U. Manchester (main)] (ORCID:00↗

Elastic Bayesian Model Calibration

Functional data are ubiquitous in scientific modeling. For instance, quantities of interest are modeled as functions of time, space, energy, density, etc. Uncertainty quantification methods for computer models with functional response have resulted in tools for emulation, sensitivity analysis, and calibration that are widely used. However, many of these tools do not perform well when the computer model’s parameters control both the amplitude variation of the functional output and its alignment (or phase variation). This paper introduces a framework for Bayesian model calibration when the model responses are misaligned functional data. The approach generates two types of data out of the misaligned functional responses: (1) aligned functions so that the amplitude variation is isolated and (2) warping functions that isolate the phase variation. These two types of data are created for the computer simulation data (both of which may be emulated) and the experimental data. The calibration approach uses both types so that it seeks to match both the amplitude and phase of the experimental data. The framework is careful to respect constraints that arise, especially when modeling phase variation, and is framed in a way that it can be done with readily available calibration software. In conclusion, we demonstrate the techniques on two simulated data examples and on two dynamic material science problems: a strength model calibration using flyer plate experiments and an equation of state model calibration using experiments performed on the Sandia National Laboratories’ Z-machine.

97 MATHEMATICS AND COMPUTING↗

Development and Assessment of SNP Genotyping Arrays for Citrus and Its Close Relatives

Rapid advancements in technologies provide various tools to analyze fruit crop genomes to better understand genetic diversity and relationships and aid in breeding. Genome-wide single nucleotide polymorphism (SNP) genotyping arrays offer highly multiplexed assays at a relatively low cost per data point. We report the development and validation of 1.4M SNP Axiom® Citrus HD Genotyping Array (Citrus 15AX 1 and Citrus 15AX 2) and 58K SNP Axiom® Citrus Genotyping Arrays for Citrus and close relatives. SNPs represented were chosen from a citrus variant discovery panel consisting of 41 diverse whole-genome re-sequenced accessions of Citrus and close relatives, including eight progenitor citrus species. SNPs chosen mainly target putative genic regions of the genome and are accurately called in both Citrus and its closely related genera while providing good coverage of the nuclear and chloroplast genomes. Reproducibility of the arrays was nearly 100%, with a large majority of the SNPs classified as the most stringent class of markers, “PolyHighResolution” (PHR) polymorphisms. Concordance between SNP calls in sequence data and array data average 98%. Phylogenies generated with array data were similar to those with comparable sequence data and little affected by 3 to 5% genotyping error. Both arrays are publicly available.

59 BASIC BIOLOGICAL SCIENCES↗

A Generative Model for Realistic Galaxy Cluster X-Ray Morphologies

Abstract The X-ray morphologies of clusters of galaxies display significant variations, reflecting their dynamical histories and the nonlinear dependence of X-ray emissivity on the density of the intracluster gas. Qualitative and quantitative assessments of X-ray morphology have long been considered a proxy for determining whether clusters are dynamically active or “relaxed.” Conversely, the use of circularly or elliptically symmetric models for cluster emission can be complicated by the variety of complex features realized in nature, spanning scales from megaparsecs down to the resolution limit of current X-ray observatories. In this work, we use mock X-ray images from simulated clusters from The Three Hundred project to define a basis set of cluster image features. We take advantage of the clusters’ approximate self-similarity to minimize the differences between images before encoding the remaining diversity through a distribution of high-order polynomial coefficients. Principal component analysis then provides an orthogonal basis for this distribution, corresponding to natural perturbations from an average model. This representation allows novel, realistically complex X-ray cluster images to be easily generated, and we provide code to do so. The approach provides a simple way to generate training data for cluster image analysis algorithms and could be straightforwardly adapted to generate clusters displaying specific types of features or selected by physical characteristics available in the original simulations.

79 ASTRONOMY AND ASTROPHYSICS↗

Data and scripts associated with “Non-random processes impacting organic matter chemistry are maximized in mid-order streams”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the publication “Non-random processes impacting organic matter chemistry are maximized in mid-order streams” submitted to Limnology and Oceanography (L&O) by Danczak et al. (in review). This package contains data and scripts used to investigate dissolved organic matter (DOM) molecular chemistry and diversification processes across 47 surface-water sampling sites in the Yakima River Basin, Washington, USA, during an August 2021 sampling campaign. The package contains analyses of ultrahigh-resolution Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS), geochemical measurements, geospatial attributes, molecular diversity, and meta-metabolome ecological null models needed to reproduce the main manuscript results. The underlying field data were pulled from exising data packages at https://doi.org/10.15485/1892052 (Fulton et al., 2022) and https://doi.org/10.15485/1898914 (Grieger et al., 2022). For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. We thank the following organizations for providing access to field locations for sample collection: the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, the Confederated Tribes and Bands of the Yakama Nation, and the Cowiche Canyon Conservatory. Research was conducted under Washington State Parks and Recreation Commission Scientific Research Permit #210901. We are grateful to the Yakama Nation Tribal Council and Yakama Nation Fisheries for their collaboration in facilitating sample collection and ensuring data usage aligns with their values and worldview. This data package contains an R-Markdown file for analyses and five folders: (1) Data, (2) Geospatial Data, (3) Supplemental_Files, (5) Figures_pdf, (4) and src. The Data folder contains tabular inputs and derived files used in the manuscript analysis. The Geospatial Data folder contains climate and water-balance, hydrologic, land-cover, population/regional water-use, stream, topographic, and stream-order attribute CSV files. The src folder contains scripts used to process data, run analyses, and generate figures. The Figures_pdf folder contains manuscript figure outputs. The Supplemental_Files folder contains supplemental analysis products. All files are .csv, .pdf, .html, .png, .R, .Rmd, .svg, or .tre. This data package is associated with the rcfsa-RC2-SPS_Null_Modeling repository found at https://github.com/river-corridors-sfa/rcfsa-RC2-SPS_Null_Modeling.

54 ENVIRONMENTAL SCIENCES↗

Quantum Time Dynamics Mediated by the Yang–Baxter Equation and Artificial Neural Networks

Quantum computing shows great potential, but errors pose a significant challenge. This study explores new strategies for mitigating quantum errors using artificial neural networks (ANNs) and the Yang–Baxter equation (YBE). Unlike traditional error mitigation methods, which are computationally intensive, we investigate artificial error mitigation. We developed a novel method that combines ANNs for noise mitigation combined with the YBE to generate noisy data. This approach effectively reduces noise in quantum simulations, enhancing the accuracy of the results. The YBE rigorously preserves quantum correlations and symmetries in spin chain simulations in certain classes of integrable lattice models, enabling effective compression of quantum circuits while retaining linear scalability with the number of qubits. This compression facilitates both full and partial implementations, allowing the generation of noisy quantum data on hardware alongside noiseless simulations using classical platforms. By introducing controlled noise through the YBE, we enhance the data set for error mitigation. We train an ANN model on partial data from quantum simulations, demonstrating its effectiveness in mitigating errors in time-evolving quantum states, providing a scalable framework to enhance quantum computation fidelity, particularly in noisy intermediate-scale quantum (NISQ) systems. We demonstrate the efficacy of this approach by performing quantum time dynamics simulations using the Heisenberg XY Hamiltonian on real quantum devices.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

GRid Analysis and Visualization Interface (GRAVI) [SWR-24-16]

GRAVI (GRid Analysis and Visualization Interface) is a web application for viewing and analyzing nodal Production Cost Model (PCM) and Capacity Expansion Model (CEM) simulations. The web application provides the ability to animate geospatially coupled timeseries data in an agnostic way regardless of the underlying simulation tool used to generate the data. GRAVI also provides capabilities to animate non-geospatial data relevant to a PCM or CEM model. Furthermore, this web application can be tailored as an real-time operational tool to better understand a live grid.

Webb, Micah↗

Organic Matter Concentration and Composition in November 2021 and April 2022 from 12 Streams Impacted by the 2020 Holiday Farm Fire (v2)

This dataset represents results from a field study aiming to understand storm induced transport of pyrogenic materials to streams impacted by varying degrees of burn severity. Time series samples were collected at 5 sites within the McKenzie River Watershed (Oregon, USA) whose catchment were each completely engulfed by the 2020 Holiday Farm Fire. An additional 7 sites were sampled once during the storm. The samples were collected during storm events in November 2020, January 2021, November 2021, and April 2022. Samples were characterized for benezenepolycarboxylic acids (BPCA), ultra-high resolution mass spectrometry, dissolved organic carbon and optics (absorbance and fluorescence). Fourier-transform ion cyclotron resonance mass spectrometry (FTICR) and dissolved organic carbon data from the November 2020 (referred to as “EWEB_2020”) sampling can be found in a separate data package (doi: 10.15485/1869708). NOTE: The 2020 samples were run on FTICR-MS in two unique instances. The first run can be found in the previous data package (EWEB_2020). The second run is included in this data package. These samples were run for a second time so that the data were more directly interoperable with the other samples in this data package. We have not done any investigation into the differences/similarities between these datasets and the previously ran/published data in the other data package. This data package was originally published in November 2024. It was updated in April 2025 (v2; new and modified files). See the change history section below for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset contains (1) file-level metadata; (2) data dictionary; (3) data package readme; (4) metadata; (5) methods information; (6) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data; (7) excitation emission matrix (EEM) methods; and (8) a sub-folder with processed EEM data (9) benzene polycarboxylic acid (BPCA) concentration data; (10) Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) methods; and (11) folder of high-resolution characterization of organic matter via 12 Tesla FTICR-MS generated through the Environmental Molecular Sciences Laboratory (EMSL; https://www.pnnl.gov/environmental-molecular-sciences-laboratory). The EEMs sub-folder contains two additional folders; the Absorbance and Fluorescence folders which contain the processed EEMs absorbance and fluorescence data respectively. This package contains the following file types: csv, xml, pdf.

54 ENVIRONMENTAL SCIENCES↗

Calculating a minimum overlap period for successful intercalibration of soil moisture sensors

Abstract Long‐term in situ soil moisture monitoring inevitably requires sensors to be replaced. Ensuing discontinuities in the data record can be mitigated by intercalibration, however it is unclear how long the existing sensor needs to remain alongside the newly installed before there is enough overlapping data to generate a robust intercalibration. We used 154 pairs of established and newly installed sensors within the Marena, Oklahoma, In Situ Sensor Testbed to determine if there is a minimum overlap time that should be considered when planning upcoming replacements. Hourly observations of the existing sensor were linearly calibrated to those of the newly installed sensor with coefficients determined from overlap periods incremented by 30 days until a reference period of 2 years was reached. The resulting bias, root‐mean‐square error, and correlation coefficient for sensor pairs indicate that a minimum of 6 to 9 months of overlapping data are required to generate a successful intercalibration. Extending that to a full year before decommissioning the old sensor results in a stable intercalibration with higher confidence.

Agriculture↗

Generative learning of densities on manifolds

A generative modeling framework is proposed that combines diffusion models and manifold learning to efficiently sample data densities on manifolds. The approach utilizes Diffusion Maps to uncover possible low-dimensional underlying (latent) spaces in the high-dimensional data (ambient) space. Two approaches for sampling from the latent data density are described. The first is a score-based diffusion model, which is trained to map a standard normal distribution to the latent data distribution using a neural network. The second one involves solving an Itô stochastic differential equation in the latent space. Additional realizations of the data are generated by lifting the samples back to the ambient space using Double Diffusion Maps , a recently introduced technique typically employed in studying dynamical system reduction; here the focus lies in sampling densities rather than system dynamics. The proposed approaches enable sampling high dimensional data densities restricted to low-dimensional, a priori unknown manifolds. The efficacy of the proposed framework is demonstrated through a benchmark problem and a material with multiscale structure.

Double diffusion maps↗