Search NASASearch

SEARCH · Search NASA

Results for “initial data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

A Data Science and Machine Learning Platform Supporting Large Particle Accelerator Control and Diagnostics Applications Final Report: SBIR Initial Phase II DE-SC0022583

The Machine Learning Data Platform (MLDP) is a product providing full-stack support for data science, Machine Learning, and Artificial Intelligence (ML/AI) applications at particle accelerator and large experimental physics facilities. It supports ML/AI applications from front-end, high-speed acquisition of heterogeneous, time-series data, through data archiving and management, to back-end analysis. The MLDP embodies a “data-science ready” platform for data analysis and ML/AI applications in diagnosis, modelling, control, and optimization of these facilities. It provides data scientists and applications a consistent, datacentric interface to archive data standardizing implementation and deployment of ML/AI algorithms to different operations configurations within the same facility, or between facilities. Being an open-source, public-domain project, the MLDP is intended for broadest possible impact by increasing accessibility and minimizing the required expertise for installation and operation. The MLDP can also be deployed at user facilities for experimental data collection, archiving, and analysis. It is capable of acquisition and archiving of heterogeneous data from experimental equipment (e.g., images, arrays, structures, etc.) along with system hardware configurations (e.g., scalars, tables), control system process variables, and any metadata required for provenance. Thus, the MLDP can manage experimental data through its entire lifecycle, from acquisition and archiving, through analysis and investigation, to release and final publication.

43 PARTICLE ACCELERATORS

Chimera F-Series Gravitational Wave Emission Sourced from Matter and Neutrino Anisotropy

This dataset contains gravitational wave data sourced from the the time-dependent fluid quadrupole motion as well as neutrino anisotropy, in the Chimera F-Series two-dimensional core collapse supernova simulations. Data from two models initiated from two different progenitors are presented: F15.78 and F15.79. Please see the README for more information about the data structure and progenitors.

79 ASTRONOMY AND ASTROPHYSICS

Chimera D-Series Gravitational Wave Emission Sourced from Neutrino Anisotropy

Gravitational wave data sourced from the time-dependent anisotropic neutrino emission, as well as the time-dependent fluid quadrupole motion, in the Chimera D-Series three-dimensional core collapse supernova simulations. Data from three models initiated from three different progenitors are presented: D9.6-3D, D15-3D, and D25-3D. Please see the README for more information about the data structure and progenitors.

79 ASTRONOMY AND ASTROPHYSICS

Chimera D-Series Gravitational Wave Emission Sourced from Matter

Gravitational wave data sourced from the the time-dependent fluid quadrupole motion, in the Chimera D-Series three-dimensional core collapse supernova simulations. Data from three models initiated from three different progenitors are presented: D9.6-3D, D15-3D, and D25-3D. Please see the README for more information about the data structure and progenitors.

79 ASTRONOMY AND ASTROPHYSICS

A Data-driven approach to Core Power distribution reconstruction in a Nuclear Reactor

This report presents the initial development of a data-driven approach for reconstructing the core power distribution in a nuclear reactor (power shape synthesis) using ex-core sensors. Traditional techniques rely on deploying a large number of detectors throughout the reactor core. However, this approach is not feasible for innovative reactor concepts like Advanced Reactors and Microreactors. First, the tight lattice pitch, designed to maximize power density, limits the space available for sensors. Secondly, the harsh operating conditions are not compatible with commercially available detectors. The method proposed in this work integrates high-fidelity modeling with data-driven techniques to accurately reconstruct power distribution across various reactor types, thereby reducing the reliance on in-core sensors. Purdue University Reactor One (PUR-1) was selected as the test case. The CAD model representing the latest configuration of the PUR-1 core was imported into the OpenMC simulation framework, and the model was built. Additionally, the previously developed MCNP6 model was updated. The two models were assessed against the data collected during an experimental campaign conducted in July 2024. Thirty gold foils were placed in three Irradiation Assemblies in PUR-1 core. Using the measured activity of the irradiated foils, the neutron flux at different core locations was estimated.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Final report- UFL - RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING

MIBiG 4.0: advancing biosynthetic gene cluster curation through global collaboration

Specialized or secondary metabolites are small molecules of biological origin, often showing potent biological activities with applications in agriculture, engineering and medicine. Usually, the biosynthesis of these natural products is governed by sets of co-regulated and physically clustered genes known as biosynthetic gene clusters (BGCs). To share information about BGCs in a standardized and machine-readable way, the Minimum Information about a Biosynthetic Gene cluster (MIBiG) data standard and repository was initiated in 2015. Since its conception, MIBiG has been regularly updated to expand data coverage and remain up to date with innovations in natural product research. Here, we describe MIBiG version 4.0, an extensive update to the data repository and the underlying data standard. In a massive community annotation effort, 267 contributors performed 8304 edits, creating 557 new entries and modifying 590 existing entries, resulting in a new total of 3059 curated entries in MIBiG. Particular attention was paid to ensuring high data quality, with automated data validation using a newly developed custom submission portal prototype, paired with a novel peer-reviewing model. MIBiG 4.0 also takes steps towards a rolling release model and a broader involvement of the scientific community. MIBiG 4.0 is accessible online at https://mibig.secondarymetabolites.org/.

59 BASIC BIOLOGICAL SCIENCES

A Data-Agnostic, Continuous Machine Learning Framework for Application in High Energy Physics and Beyond: Phase 1 Final Scientific/Technical Report

This Phase 1 effort has focused on the development of continual learning frameworks for use in machine learning, specifically in the applied context of High Energy Physics (HEP). Machine learning (ML) is a transformative technology by which computers, typically through the use of neural networks, are able to perform tasks with proficiency that rivals or surpasses that of human users. Model Degradation & Catastrophic Forgetting are two undesired phenomena which can occur in ML where the performance of a model degrades when either deployed on novel data streams, or trained on novel data which are sufficiently different than the data the models were initially trained on. A natural example where these sorts of effects can be observed is in the performance of detectors in harsh environments, where the detector signature may change over the lifetime of the detector as it ages and deteriorates — precisely what occurs in the experiments conducted in HEP. Real world HEP data is therefore an excellent test-ground and use-case for Continual Learning paradigms, which are techniques used in ML to counteract these problems. Ensemble learning is one such technique, where multiple smaller models are trained on subsets of the overall data and are ensembled together during inference. The intuition behind this technique is that, although there are shifts in the distributions which govern the incoming data streams, these shifts are not expected to be homogeneous or global. If a sufficient diversity in solutions within the various sub-models has been achieved, then at least one sub-model is expected to retain its performance within the overall ensemble. One further strength of this approach is that the architectures of the various models do not need to be identical, and in fact even different modalities of data can naturally be combined in this way. This work focused on applying ensemble learning techniques to derive results using two main datasets, anomaly detection in HEP data & time-series forecasting in semiconductor manufacturing data. Semiconductor manufacturing involves data with surprising similarity to that of HEP (e.g. wafer maps look very similar to digi-occupancy maps) and Cerium Lab’s prominence within the semiconductor industry makes semiconductor manufacturing a natural opportunity for commercialization of this work. Our efforts have led to two strong results. The first is that we evaluated the proposed ensembling techniques using previously proposed machine learning architectures for use in anomaly detection, namely AutoEncoder based models and their derivatives. We also developed new architectures which have not been evaluated in this context before. In fact, this work marks the first use of Vision Transformers for anomaly detection in HEP. Second, we demonstrated that ensemble learning significantly improves model performance in scenarios prone to degradation, validating its effectiveness across both HEP and semiconductor datasets. These results further support ensemble learning as a powerful strategy for mitigating catastrophic forgetting and maintaining robust performance in evolving data environments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Quantitative radiography for determining density fluctuations in HED experiments

We have developed a method to extract density fluctuation measurements from x-ray radiographs of high-energy density (HED) instability growth and turbulence experiments. We use this information to calculate density fluctuation statistics for constraining the performance of turbulent mix models in HED systems. The density calculation combines image filtering, removal of systemic effects such as backlighter variation, calculation of transmission across multiple materials, and use of tracer materials to generate an approximate single-material density field. From the density map, we calculate both average density and a variance-like moment b (density-specific-volume covariance), which we compare to our models. We infer both quantities from a single image, which is significantly more information than the historic single scalar mix width measurements. We also develop a method of analyzing simulation outputs that incorporate both the density fluctuation metric from a turbulence model and the bulk material maps from the hydrodynamic code. This analysis helps address the question of how to initialize the simulations for best comparison to data from systems with large separations of scale in the mixing perturbation initial condition. We find that our data analysis method yields 1D average density and b curves with similar morphology and amplitudes as those from preliminary simulation comparisons.

47 OTHER INSTRUMENTATION

Introducing Molecular Hypernetworks for Discovery in Multidimensional Metabolomics Data

Orthogonal separations of data from high-resolution mass spectrometry can provide insight into sample composition and address challenges of complete annotation of molecules in untargeted metabolomics. “Molecular networks” (MNs), as used in the Global Natural Products Social Molecular Networking platform, are a prominent strategy for exploring and visualizing molecular relationships and improving annotation. MNs are mathematical graphs showing the relationships between measured multidimensional data features. MNs also show promise for using network science algorithms to automatically identify targets for annotation candidates and to dereplicate features associated with a single molecular identity. Here, this paper introduces “molecular hypernetworks” (MHNs) as more complex MN models able to natively represent multiway relationships among observations. Compared to MNs, MHNs can more parsimoniously represent the inherent complexity present among groups of observations, initially supporting improved exploratory data analysis and visualization. MHNs also promise to increase confidence in annotation propagation, for both human and analytical processing. We first illustrate MHNs with simple examples, and build them from liquid chromatography- and ion mobility spectrometry-separated MS data. We then describe a method to construct MHNs directly from existing MNs as their “clique reconstructions”, demonstrating their utility by comparing examples of previously published graph-based MNs to their respective MHNs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Genesis Data Card Schema, Template and Supporting Tools

Genesis Data Cards provide a standardized template and schema for documenting scientific datasets in support of discovery, access, interoperability, reusability, governed use, and AI usability. This release of the Genesis Data Card repository includes a versioned Markdown template, a LinkML schema with generated Pydantic and JSON artifacts, schema documentation, and example completed data cards. Validation tooling is provided to ensure that completed data cards conform to the schema prior to submission. Accompanying documentation for the structured metadata is provided as a Field Reference Guide. The schema and accompanying template provided in this repository address the call for actionable context that enables humans and AI systems to find, access, interpret, cite, and reuse data, and, when appropriate, integrate it into AI and machine learning workflows. The data card is intended to serve as a common metadata artifact intended to support standardized, cross-program dataset documentation across Department of Energy (DOE)-aligned efforts, including but not limited to Genesis Mission-related implementations, the Office of Science, National Nuclear Security Administration (NNSA), and Advanced Simulation and Computing (ASC) data governance and stewardship initiatives.

data card

Measurements and kinetic modeling of O 2 vibrational kinetics in O 2 –Ar mixtures partially dissociated by a Ns pulse discharge

Vibrational kinetics of O 2 is studied during the O atom recombination in an O 2 –Ar mixture, partially dissociated by a burst of ns discharge pulses in a heated plasma flow reactor. The time-resolved temperature in the discharge afterglow is determined by Rayleigh scattering. Time-resolved O atom number density is measured by ps Two-Photon absorption Laser Induced Fluorescence, calibrated in xenon. Time-resolved vibrational level populations of molecular oxygen, O 2 (v= 8–20), are measured by ps Laser Induced Fluorescence (LIF), with the absolute calibration by NO LIF. Time-resolved ozone number density is monitored by broadband UV absorption. The results are compared with the predictions of a state-specific kinetic model. The experimental data indicate a rapid initial decay of O 2 (v) populations generated by electron impact in the discharge, due to the vibration-translation (V–T) relaxation by O atoms. This is followed by a slower population reduction, on the time scale much longer compared to that for V–T relaxation or vibration-vibration (V–V) exchange. Both O atoms and the O 2 (v) populations decay on the same time scale, indicating that chemical reactions initiated by the O atom recombination result in the generation of vibrationally excited O 2 molecules. These trends are reproduced by the kinetic model, which shows that the reaction of O atoms with ozone is the dominant pathway of O 2 (v) generation at the present conditions. The predicted relative O 2 (v) populations are close to the experimental results, but absolute number densities differ from the experimental data. This is likely due to uncertainties in the absolute calibration of LIF measurements and in the spectroscopic model used in the data reduction. The present work demonstrates the capability for the absolute, time-resolved measurements of vibrationally excited O 2 in recombining gas flows, to quantify the energy partition in the recombination reactions.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Effects of 9.5 Years of Whole-Soil Warming on the Fatty Acid and n-Alkanes Composition in Bulk Soil and Density Fractions at Blodgett Experimental Forest, California, USA

Original data of molecular data (fatty acids and n-alkanes) including concentrations and calculated molecular proxies in a whole-soil warming experiment at the Blodgett Forest Research Station after 9.5 years of warming. The study site has a Mediterranean climate with annual average temperature of 12.5 ℃ and annual average precipitation of 1774 mm. The study site is characterized by a mesic Ultic Alfisol formed from granitic parent material, corresponding to a Dystric Cambisol under the World Reference Base for Soil Resources (WRB) classification system. Experimental warming is applied throughout the soil profile to a depth of 1 m using vertically embedded heating cables that raise soil temperature by 4 °C relative to ambient conditions. Soil samples were collected on 1 May 2023, after the experiment had been operating continuously for about 9.5 years since its initiation in January 2014.The data has been processed from raw data and cross-validated by other peers. The dataset includes: - Bulk_Fattyacid_9.5-year_Soil_Warming_Blodgett, California, USA: fatty acid concentrations and proxies including Carbon Preference Index (CPI) and Average Chain Length (ACL) of bulk soil organic carbon; - Fractions_Fattyacid_9.5-year_Soil_Warming_Blodgett, California, USA: fatty acid concentrations and proxies including CPI and ACL of free particulate organic matter (fPOM) and mineral-associated organic matter (MAOM); - Bulk_Alkanes_9.5-year_Soil_Warming_Blodgett, California, USA: n-alkanes concentrations and proxies including CPI and ACL of bulk soil organic carbon; - Fractions_Alkanes_9.5-year_Soil_Warming_Blodgett, California, USA: n-alkanes concentrations and proxies including CPI and ACL of fPOM and MAOM; - n-Alkanes_All_Monomer_Concentration_9.5-year_Soil_Warming_Blodgett, California, USA: concentration of all the n-alkane monomers identified and integrated for bulk soil, fPOM and MAOM; - Fattyacid_All_Monomer_Concentration_9.5-year_Soil_Warming_Blodgett, California, USA: concentration of all the fatty acid monomers including diacids identified and integrated for bulk soil, fPOM, and MAOM. All data are provided in CSV format and can be viewed using Microsoft Excel. We specifically look at fatty acids (FA) and n-alkanes in bulk soil, fPOM and MAOM and calculated molecular proxies such as CPI and ACL to understand the source of oragnic carbon (with ACL) and degree of decomposition (CPI) of each soil fraction. Due to lack of long-chain fatty acids (carbon number ⩾ 20), microorganism-derived organic carbon is characterized by shorter ACL in comparison to plant-derived organic carbon. Fresh SOC is characterized by even-over-odd dominance for fatty acids and odd-over-even dominance for n-alkanes. Therefore, CPI indicates whether soil organic carbon (SOC) represents fresh input (CPI > 10) or is strongly decomposed (close to 1). The research questions should be then, after 9.5-year warming: 1. whether the relative contribution between microorganism-derived and plant-derived SOC in each soil fraction? 2. whether fPOM became more decomposed whereas MAOM remained relatively persistent in each soil fraction across the soil depth?

Carbon

Validation Data for Benchmarking Wire Arc Additive Manufacturing Process Simulations

Residual stresses cause geometric distortion and affect mechanical performance of additively manufactured structures, yet they are notoriously difficult to assess and predict. Distortion (warpage) can drive parts outside dimensional tolerance limits, leading to part rejection or rework. For parts that meet tolerance, locked-in residual stress fields can affect structural integrity during operation, particularly subcritical cracking by fatigue, creep, or corrosion. This work develops benchmark data for a common additive manufacturing process (Wire Arc Additive Manufacturing) that can be applied for calibration and validation of physical process models that predict residual stress fields. The work includes design of two different samples of differing geometry, detailed manufacturing records for a set of physical samples, and an extensive set of residual stress measurement data developed using two diverse techniques (the contour method and neutron diffraction). An initial application of the work is also reported, where a modeling challenge was issued to secure residual stress model predictions from two independent laboratories that were blind to residual stress measurement data. These initial blind residual stress predictions show significant discrepancies relative to the measurement data, illustrating the potential value of the underlying validation data. An open repository for this work, including the sample designs, manufacturing process records, and the residual stress data, is also provided for future application in non-blind validation efforts.

36 MATERIALS SCIENCE

Deep Design Data Portal (D3P) v0.01

The Deep Design Data Portal (D3P) tool was developed to demonstrate how readily accessible data sources, such as building energy model reports for design and baseline energy performance data for projects, can provide the data required for reporting to an industry initiative (AIA 2030 commitment), as well as more detailed data that makes the industry dataset more valuable to all stakeholders, enabling project level analysis and analysis of BEM industry trends. D3P provides an easier and less time-consuming way for firms to auto-extract data from this data source, compared to the current reporting workflows of the firms. The BEM reports are the first of several data sources that D3P could integrate. D3P also provides the ability for firms to review, compare, and evaluate the performance of their projects to not only their portfolio, but also to the larger anonymized industry dataset created each time a project is added to D3P. The intent of D3P is to become part of a data-sharing ecosystem to assist creating large anonymized industry datasets that are accessible to industry.

Regnier, Cynthia [Lawrence Berkeley National Labor

Grid-Ready Energy Analytics Training with Data (“GREAT with Data”)

GridEd is a collaborative educational initiative consisting of the Electric Power Research Institute (EPRI), 5 Partner Universities (Stony Brook University, The University of Texas at Austin, University of California – Riverside, Virginia Tech, Washington State University), and participating industry sponsors. This educational initiative focuses on developing and training the next generation of power engineers so they can help shape the electric grid of the future by anticipating and fulfilling the needs of changing electric industry requirements. GridEd is leveraging electric industry research to educate a future electric grid workforce by empowering new and continuing education students, not only to become competent and well-informed engineers, but also to participate and influence major technological, social, and policy decisions that address critical global challenges. GridEd’s activities are centered around four core pillars: Enhancement of university power systems engineering curricula; Professional development and training for a diverse electric industry workforce; Stimulating students to join the movement for the next generation of power engineers, and; Improve workforce development efforts in the electric utility industry. Major accomplishments over the course of the project were: Over 50 unique professional short courses were delivered by more than 30 instructors across the GridEd network; Over 3,600 unique learners, many who took multiple courses, received more than 27,000 professional development hours (PDH) and over 1,000 certificates of completion; Approximately 2,900 unique learners undertook a course that was offered LIVE online or in-person; Approximately 700 unique learners undertook a course that was offered as computer-based training (CBT); Over 35 unique university courses were delivered by more than 30 instructors to 1,500 university students across the GridEd network; One-hundred-and-fifty-seven (157) students were funded to completed 43 student projects in topics of power systems and data science, and; Six (6) Historically Black Colleges and Universities (HBCUs) recruited as Affiliate Universities via Utility Partners. The professional training initiative and workforce development activities launched by this project will be sustained through EPRI’s collaborative business model with industry. Stimulating students to join the power engineering workforce of the future and the enhancement of tertiary training may continue to need government support.

14 SOLAR ENERGY

Parallel-in-Time Solution of Allen-Cahn Equations by Integrating Operator Learning into the Parareal Method

While recent advances in deep learning have shown promising efficiency gains in solving time-dependent partial differential equations (PDEs), matching the accuracy of conventional numerical solvers still remains a challenge. One strategy to improve the accuracy of deep learning-based solutions for time-dependent PDEs is to use the learned model as the coarse propagator in the Parareal method and a traditional numerical method as the fine solver. However, successful integration of deep learning into the Parareal method requires consistency between the coarse and fine solvers, particularly for PDEs exhibiting rapid changes such as sharp transitions. Here, to ensure this consistency, we propose using convolutional neural networks (CNNs) to learn the fully discrete time-stepping operator defined by the same numerical scheme employed as the fine solver. We demonstrate the effectiveness of the proposed method in solving the classical and mass-conservative Allen–Cahn (AC) equations. Through iterative updates in the Parareal algorithm, our approach achieves a significant computational speedup compared to traditional fine solvers while converging to high-accuracy solutions. Our results highlight that the proposed hybrid Parareal algorithm effectively accelerates simulations, particularly when implemented on multiple GPUs, and converges to the desired accuracy in only a few iterations. Another advantage of our method is that the CNN model is trained on trajectory-based data generated from random initial conditions, such that the trained model can be used to solve the AC equations with various initial conditions without retraining. This work demonstrates the potential of integrating neural network methods into parallel-in-time frameworks for efficient and accurate simulations of time-dependent PDEs.

97 MATHEMATICS AND COMPUTING