Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Science Model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will produce unprecedented volumes of heterogeneous astronomical data (images, catalogs, and alerts) that challenge traditional analysis pipelines. The LSST Dark Energy Science Collaboration (DESC) aims to derive robust constraints on dark energy and dark matter from these data, requiring methods that are statistically powerful, scalable, and operationally reliable. Artificial intelligence and machine learning (AI/ML) are already embedded across DESC science workflows, from photometric redshifts and transient classification to weak lensing inference and cosmological simulations. Yet their utility for precision cosmology hinges on trustworthy uncertainty quantification, robustness to covariate shift and model misspecification, and reproducible integration within scientific pipelines. This white paper surveys the current landscape of AI/ML across DESC's primary cosmological probes and cross-cutting analyses, revealing that the same core methodologies and fundamental challenges recur across disparate science cases. Since progress on these cross-cutting challenges would benefit multiple probes simultaneously, we identify key methodological research priorities, including Bayesian inference at scale, physics-informed methods, validation frameworks, and active learning for discovery. With an eye on emerging techniques, we also explore the potential of the latest foundation model methodologies and LLM-driven agentic AI systems to reshape DESC workflows, provided their deployment is coupled with rigorous evaluation and governance. Finally, we discuss critical software, computing, data infrastructure, and human capital requirements for the successful deployment of these new methodologies, and consider associated risks and opportunities for broader coordination with external actors.

Aubourg, Eric [APC, Paris] (ORCID:000000025592023X↗

Scaling Laws of Graph Neural Networks for Atomistic Materials Modeling

Atomistic materials modeling is a critical task with wide-ranging applications, from drug discovery to materials science, where accurate predictions of the target material property can lead to significant advancements in scientific discovery. Graph Neural Networks (GNNs) represent the state-of-the-art approach for modeling atomistic material data thanks to their capacity to capture complex relational structures. While machine learning performance has historically improved with larger models and datasets, GNNs for atomistic materials modeling remain relatively small compared to large language models (LLMs), which leverage billions of parameters and terabyte-scale datasets to achieve remarkable performance in their respective domains. To address this gap, we explore the scaling limits of GNNs for atomistic materials modeling by developing a foundational model with billions of parameters, trained on extensive datasets in terabytescale. Our approach incorporates techniques from LLM libraries to efficiently manage large-scale data and models, enabling both effective training and deployment of these large-scale GNN models. This work addresses three fundamental questions in scaling GNNs: the potential for scaling GNN model architectures, the effect of dataset size on model accuracy, and the applicability of LLM-inspired techniques to GNN architectures. Specifically, the outcomes of this study include (1) insights into the scaling laws for GNNs, highlighting the relationship between model size, dataset volume, and accuracy, (2) a foundational GNN model optimized for atomistic materials modeling, and (3) a GNN codebase enhanced with advanced LLM-based training techniques. Our findings lay the groundwork for large-scale GNNs with billions of parameters and terabyte-scale datasets, establishing a scalable pathway for future advancements in atomistic materials modeling.

Li, Chaojian [ORNL] (ORCID:0000000340309777)↗

Bottom-up design of actinide materials from molecular clusters: Demonstration of a general-purpose simulation capability leveraging machine-learned atomic potentials

Actinide thin-film coatings such as uranium dioxide (UO 2 ) play an important role in nuclear reactors and other mission-relevant applications, but realization of their potential requires a deep fundamental understanding of the chemical vapor deposition (CVD) processes used for their growth. The slow experimental progress can be attributed, in part, to the standard safety guidelines associated with handling uranium byproducts, which are often corrosive, toxic, and radioactive. Accurate simulation techniques, when used in concert with experiment, can improve laboratory safety, material durability, and deliverable timeframes. However, state-of-the-art computational methods are either insufficiently accurate or intractably expensive. To remedy this situation, in this project we suggested a machine-learning (ML) accelerated workflow for simulating molecular clustering toward deposition. As a benchmark test case, we considered molecular clustering in steam and assessed independent components of our workflow by comparing with measured thermodynamic properties of water. After analyzing each component individually and finding no fundamental barrier to realization of the workflow, we attempted to integrate the ML component, a Sandia-developed tool called FitSNAP. As this was the first application of FitSNAP to atoms and molecules in the gas phase at Sandia, the method required more fitting data than was originally anticipated. Systematic improvements were made by including in the fit data diatomic potentials, molecular single-bond-breaking curves, and symmetry-constrained intermolecular potentials. We concluded that our strategy provides a feasible pathway toward modeling CVD and related processes, but that extensive training data must be generated before it can be of practical use.

36 MATERIALS SCIENCE↗

Electric Grid Security (EGS) FY24 Annual Report

Sandia’s Electric Grid Security program advances a national vision of a secure, resilient, and affordable electric system for all users. Our achievements reflect a strategic approach combining technology development; modeling, simulation, and data analytics; and partnered demonstrations and outreach to further the adoption of advanced grid and storage technologies. Our FY24 efforts leverage the strengths of our partnerships—spanning Sandia’s core science and technology competencies as well as external technology leaders—to develop the solutions today which enable the grid of tomorrow. Key accomplishments in this report that support our strategy span our technical program areas and include: • The advancement of energy storage technologies, including creation of a national Long Duration Energy Storage Consortium; • Applications of artificial intelligence and machine learning to enhanced grid operations and planning; • Development of solid-state power conversion technologies and a new medium-voltage research lab; • New technologies to assess wildfire vulnerabilities and mitigate potential impacts; • Advanced applications of new cybersecurity technologies with industry partners; • Contributions to understanding the impacts of electromagnetic pulses and geomagnetic disturbances on grid components; and • Digital twin development for hybrid microgrids with multiple generators, storage, and loads.

24 POWER TRANSMISSION AND DISTRIBUTION↗

FY25 Electric Grid Security Annual Report

Sandia’s Electric Grid Security program advances a national vision of energy dominance and accessibility, while applying our national security -emphasis on ensuring of a secure, resilient, and affordable electric system for all users. Our achievements reflect a strategic approach combining technology development; modeling, simulation, and data analytics; and partnered demonstrations and outreach to further the adoption of advanced grid and storage technologies. Our FY25 efforts leverage the strengths of our partnerships—spanning Sandia’s core science and technology competencies as well as external technology leaders—to develop the solutions today which enable the grid of tomorrow. Key accomplishments in this report that support our strategy span our technical program areas and include: • New open-source analytical tools for systems -level planning and optimization, including significant advances to the QuESt analytical environment; • Further advancement of artificial intelligence and machine learning to enhanced grid operations and planning as we rise to the challenge of new large loads; • Development of solid-state power conversion technologies and a new medium-voltage research lab; • New technologies to assess wildfire vulnerabilities and mitigate potential impacts; • Advanced applications of new cybersecurity technologies with industry partners; • Contributions to understanding the impacts of electromagnetic pulses and geomagnetic disturbances on grid components; and • Digital twin development for hybrid microgrids with multiple generators, storage, and loads. This report indicates key areas of research and engagement and summarizes the impact of Sandia’s contributions through notable accomplishments, journal publications, patents, and technical conferences and presentations. It is provided with the hope that readers discover ways we can further team to create our modern grid and apply the outcomes of our efforts. The bulk of work described herein is funded by several offices within the U.S. Department of Energy (USDOE), including the Office of Electricity (OE); Cybersecurity, Energy Security, and Emergency Response (CESER); former offices such as the Office of Energy Efficiency and Renewable Energy (EERE), the Grid Deployment Office (GDO), the Office of Clean Energy Demonstrations (OCED), and other key programs at USDOE. As we continue to state in these annual reports, the contributors to our successes are too numerous to name here, though our team wishes to express our deep gratitude to the numerous program and project sponsors at the US Department of Energy, who often function equally as technical collaborators; our many partners in industry, academia, utilities, and other national labs; and fellow researchers and business partners at Sandia whose leadership and creativity have enabled the accomplishments described herein.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Sensitivity of Regional WRF‐Chem Air Quality and Weather Simulations to Biomass‐Burning Emission Data Sets: A Case Study of the Impact of Canadian Wildfire on the US°

This study focuses on the period from June 26 to 29, 2023, when record‐breaking Canadian wildfires severely impacted air quality in the Midwest United States. Using the Weather Research and Forecasting Model with Chemistry (WRF‐Chem) and four biomass‐burning data sets (Fire Inventory from NCAR version 1, Fire Inventory from NCAR version 2.5, Quick Fire Emissions Data set [QFED], and Regional ABI‐VIIRS Emission), we analyzed aerosol transport from Canada to the US and assessed the model's accuracy in predicting PM 2.5 , O 3 , CO and aerosol weather feedback. Model simulations were compared with ground‐based and remote sensing observations as well as field measurements from the Community Research on Climate and Urban Science (CROCUS) project. Our findings show that the movement of a low‐pressure system from the Great Lakes to the Atlantic, combined with the high‐pressure system over the Atlantic, caused the transport of aerosols from Canadian wildfires to the US. Results show WRF‐Chem significantly underestimated key atmospheric components: aerosol optical depth (AOD) by over 50%, PM 2.5 by 65%–90% and peak O 3 concentrations by 50%–55% across four biomass burning data sets. Additionally, CO and NO 2 concentrations were underpredicted. The substantial underestimation of PM 2.5 led to an overestimation of temperature by up to 3.6 °C primarily due to excessive downward shortwave radiation, which resulted from the underestimation of direct aerosol effects and an increase in sensible heat flux. Among the biomass‐burning data sets, QFED produced the most accurate AOD and PM 2.5 predictions due to improved wildfire emission estimates, leading to a 1.0 to 1.5 °C reduction in temperature overestimation during the daytime. These findings underscore the need for improving wildfire emission estimates for trace gases and aerosols to enhance air quality and weather feedback predictions.

WRF-chem model↗

The U.S. Agrivoltaic Shading Tool: A National-Scale Interface for Modeling Light and Shade Patterns in Ten Common Agrivoltaic Configurations

Agrivoltaic systems are dual-use configurations that co-locate agriculture and photovoltaic (PV) infrastructure and require careful design to balance crop performance and energy generation. A critical element of agrivoltaic design is the spatial and temporal distribution of irradiance and shade within and around PV arrays. To support research, planning, and stakeholder decision-making, we introduce the U.S. Agrivoltaic Shading Tool, a novel web-based application that delivers high-resolution irradiance and photosynthetically active radiation (PAR) modeling for ten standardized PV configurations across the conterminous United States. The tool leverages the National Laboratory of the Rockies (NLR) System Advisor Model (SAM) to perform detailed irradiance simulations, using meteorological data from the National Solar Radiation Database (NSRDB). Outputs include seasonal, monthly, weekly, and diurnal patterns of available sunlight, amount of shade, irradiance, and PAR at ground level within agrivoltaic system footprints. For a user's selected location, these results are visualized through interactive visualizations, heatmaps, and time-series plots, designed to be accessible to both technical and non-technical users. In addition to facilitating rapid spatial exploration of agrivoltaic light environments, the tool will offer seamless integration with the InSPIRE Agrivoltaics Design and Analysis Model (ADAM). This optional workflow will allow users to port selected site and configuration parameters into a more advanced modeling environment for further customization of structural layouts, crop-system compatibility, power generation, and technoeconomic performance. Finally, to promote open science, the entire dataset will be hosted and available for open access through the OpenEI platform. By standardizing and disseminating high-quality irradiance data and design tools, the U.S. Agrivoltaic Shading Tool supports a wide range of users, including researchers, landowners, energy developers, and policymakers, in evaluating the agronomic and energetic feasibility of agrivoltaic systems across the United States.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Atmospheric Radiation Measurement (ARM) airborne field campaign data products between 2013 and 2018

Airborne measurements are pivotal for providing detailed, spatiotemporally resolved information about atmospheric parameters and aerosol and cloud properties, thereby enhancing our understanding of dynamic atmospheric processes. For 30 years, the US Department of Energy (DOE) Office of Science supported an instrumented Gulfstream 1 (G-1) aircraft for atmospheric field campaigns. Data from the final decade of G-1 operations were archived by the Atmospheric Radiation Measurement (ARM) Data Center and made publicly available at no cost to all registered users. To ensure a consistent data format and to improve the accessibility of the ARM airborne data, an integrated dataset was recently developed covering the final 6 years of G-1 operations (2013 to 2018, https://doi.org/10.5439/1999133; Mei and Gaustad, 2024). The integrated dataset includes data collected from 236 flights (766.4 h), which covered the Arctic, the US Southern Great Plains (SGP), the US West Coast, the eastern North Atlantic (ENA), the Amazon Basin in Brazil, and the Sierras de Córdoba range in Argentina. These comprehensive data streams provide much-needed insight into spatiotemporal variability in the thermodynamic quantities and aerosol and cloud properties for addressing essential science questions in Earth system process studies. This paper describes the DOE ARM merged G-1 datasets, including information on the acquisition, data collection challenges and future potentials, and quality control processes. It further illustrates the usage of this merged dataset to evaluate the Energy Exascale Earth System Model (E3SM) with the Earth System Model Aerosol–Cloud Diagnostics (ESMAC Diags) package.

54 ENVIRONMENTAL SCIENCES↗

Automated Resonance Fitting for Nuclear Data Evaluation

Global and national efforts to deliver high-quality nuclear data to users have a wide-ranging impact, affecting applications in national security, reactor operations, basic science, medicine, and more. Cross section evaluation is a major part of this effort, combining theory and experimentation to produce recommended values and uncertainties for reaction probabilities. Resonance region evaluation is a specialized type of nuclear data evaluation that can require significant manual effort and months of time from expert scientists. In this article, non-convex non-linear optimization methods are combined with concepts of inferential statistics to infer a resonance model from experimental data in an automated manner that is not dependent on prior evaluation(s). This methodology aims to enhance the workflow of a resonance evaluator by minimizing time, effort, and the potential for bias from prior assumptions, while enhancing reproducibility and documentation, thereby addressing well-known challenges in the field.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Scientific Discovery with Physics-Informed System Identification (Abbreviated Report)

My fellowship research focused on making physics-based simulations faster and more useful through machine learning. Many problems in science and engineering are governed by partial differential equations, but high-fidelity simulations are often too expensive to run repeatedly. I worked on improving Latent Space Dynamics Identification (LaSDI), a reduced-order modeling framework that compresses large simulation data sets into a smaller representation and then learns how that representation evolves over time. The motivation was to develop reduced models that remain accurate for more challenging systems, especially when predictions must remain reliable over long time intervals or when the underlying dynamics are more complicated than standard methods can easily handle. I also contributed to related work on Quandary, a high-performance software effort for simulation and control of open quantum systems, before focusing primarily on Latent Space Dynamics Identification methods. The main outcomes of the fellowship were two new algorithms (both of which were published), Rollout-LaSDI and Higher-Order LaSDI, together with supporting work on multi-stage Latent Space Dynamics Identification. Rollout-LaSDI improved long-term prediction by training the model to stay accurate over extended time horizons, and Higher-Order LaSDI broadened the method so it could model systems with higher-order time dynamics. My contributions to multistage Latent Space Dynamics Identification also helped show that its later training stages could be simplified without losing effectiveness, and that this behavior held across different model architectures and training strategies. Taken together, these advances improved the accuracy, flexibility, and practical value of reduced-order modeling tools for computational science.

97 MATHEMATICS AND COMPUTING↗

Precision measurements of EFT parameters and BAO peak shifts for the Lyman- α forest

We present precision measurements of the bias parameters of the one-loop power spectrum model of the Lyman- α (Ly- α ) forest, derived within the effective field theory (EFT) of large-scale structure. We fit our model to the three-dimensional flux power spectrum measured from the ACCEL 2 hydrodynamic simulations. The EFT model fits the data with an accuracy of below 2% up to k = 2 h Mpc − 1 . Further, we analytically derive how nonlinearities in the three-dimensional clustering of the Ly- α forest introduce biases in measurements of the baryon acoustic oscillations (BAOs) scaling parameters in radial and transverse directions. From our EFT parameter measurements, we obtain a theoretical error budget of Δ α ∥ = − 0.2 % ( Δ α ⊥ = − 0.3 % ) for the radial (transverse) parameters at redshift z = 2.0 . This corresponds to a shift of − 0.3 % (0.1%) for the isotropic (anisotropic) distance measurements. We provide an estimate for the shift of the BAO peak for Ly- α -quasar cross-correlation measurements assuming analytical and simulation-based scaling relations for the nonlinear quasar bias parameters resulting in a shift of − 0.2 % ( − 0.1 % ) for the radial (transverse) dilation parameters, respectively. This analysis emphasizes the robustness of Ly- α forest BAO measurements to the theory modeling. We provide informative priors and an error budget for measuring the BAO feature—a key science driver of the currently observing Dark Energy Spectroscopic Instrument (DESI). Our work paves the way for full-shape cosmological analyses of Ly- α forest data from DESI and upcoming surveys such as the Prime Focus Spectrograph, WEAVE-QSO, and 4MOST. Published by the American Physical Society 2025

de Belsunce, Roger (ORCID:0000000336604028)↗

Autonomous hybrid optimization of a SiO 2 plasma etching mechanism

Computational modeling of plasma etching processes at the feature scale relevant to the fabrication of nanometer semiconductor devices is critically dependent on the reaction mechanism representing the physical processes occurring between plasma produced reactant fluxes and the surface, reaction probabilities, yields, rate coefficients, and threshold energies that characterize these processes. The increasing complexity of the structures being fabricated, new materials, and novel gas mixtures increase the complexity of the reaction mechanism used in feature scale models and increase the difficulty in developing the fundamental data required for the mechanism. This challenge is further exacerbated by the fact that acquiring these fundamental data through more complex computational models or experiments is often limited by cost, technical complexity, or inadequate models. In this paper, we discuss a method to automate the selection of fundamental data in a reduced reaction mechanism for feature scale plasma etching of SiO 2 using a fluorocarbon gas mixture by matching predictions of etch profiles to experimental data using a gradient descent (GD)/Nelder–Mead (NM) method hybrid optimization scheme. These methods produce a reaction mechanism that replicates the experimental training data as well as experimental data using related but different etch processes.

36 MATERIALS SCIENCE↗

Using Data Science Tools to Reveal and Understand Subtle Relationships of Inhibitor Structure in Frontal Ring-Opening Metathesis Polymerization

The rate of frontal ring-opening metathesis polymerization (FROMP) using the Grubbs generation II catalyst is impacted by both the concentration and choice of monomers and inhibitors, usually organophosphorus derivatives. Herein we report a data-science-driven workflow to evaluate how these factors impact both the rate of FROMP and how long the formulation of the mixture is stable (pot life). Using this workflow, we built a classification model using a single-node decision tree to determine how a simple phosphine structural descriptor (V bur-near ) can bin long versus short pot life. Additionally, we applied a nonlinear kernel ridge regression model to predict how the inhibitor and selection/concentration of comonomers impact the FROMP rate. Furthermore, the analysis provides selection criteria for material network structures that span from highly cross-linked thermosets to non-cross-linked thermoplastics as well as degradable and nondegradable materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Rapid data acquisition and machine learning-assisted composition design of functionally graded alloys via wire arc additive manufacturing

Abstract The lack of high-quality datasets in materials science hinders artificial intelligence (AI)-driven alloy design. To address this challenge, wire arc additive manufacturing (WAAM) was employed to fabricate graded alloys, generating extensive data for machine learning (ML)-assisted property prediction. ML models were developed using high-throughput experiments, computational models, and genetic algorithm to optimize feature selection, successfully predicting hardness and porosity. The ML model demonstrated its efficacy by designing a gradient alloy with enhanced properties. However, scaling up revealed uncertainties in tensile property and porosity due to differences in size and thermal conditions between the designed alloy build and the gradient print used to construct the ML model. This underscores the need for uncertainty quantification and process optimization in WAAM-driven alloy design. Our work advances AI-integrated additive manufacturing, offering a rapid approach to exploring process–structure–property relationships and accelerating materials development.

Wang, Xin↗

Preparation of the Multi-Site Data Processing at the Vera C. Rubin Observatory

The Vera C. Rubin Observatory’s Legacy Survey of Space and Time (LSST) Camera is scheduled to start taking data in the summer of 2025. The Data Release Production will run the LSST Science Pipe software at data facilities in the US, France and the UK. The LSST Science Pipeline consists of complex directed acyclic graphs (DAGs) of tasks. Rubin will use the Production and Distributed Analysis (PanDA) workflow and workload management system to orchestrate this complex workflow and the distribution of workloads to the data facilities. When run end-to-end by a team of data production staff, this processing (the Science Pipelines, distributed by the workflow and workload management system) is referred to as a 'campaign'. This paper describes the central services and data facility specific services that support this multi-site data process model, including the service deployment infrastructure, the workload and workflow system, the Campaign Management tools, and connection to Rubin Data Management. This paper will also mention the experience of processing the Rubin Commissioning Camera data. All these are part of the effort to scale up the processing capabilities for the expected very large data volume from the LSST Camera.

Yang, Wei [SLAC]↗

Electromagnetic Induction (EMI) Data, 2024, Trail Creek, Colorado

This dataset contains Electromagnetic Induction (EMI) data collected at Trail Creek, Colorado, in 2024. EMI surveys were conducted to investigate the spatial distribution of electrical conductivity in the subsurface, providing insights into soil moisture and subsurface geological features. The surveys were performed along multiple transects to capture variations in conductivity influenced by changes in soil composition, moisture content, and underlying geological structures. This dataset complements other geophysical data collected in the region, including Electrical Resistivity Tomography (ERT) and Terrestrial LiDAR Scanning (TLS), providing a detailed understanding of the subsurface and its impact on surface vegetation and hydrological processes. The data are valuable for environmental geophysics, ecological research, and hydrological modeling in mountainous ecosystems. The files include: - data.zip: the raw EMI data (.csv) - inversion.zip: the inverted resistivity model (.csv and .kml) - kriging.zip: the kriging resistivity model (.csv, .tif, .kmz) - flmd.csv: file level metadata file describing all files within this dataset - dd.csv: data dictionary file describing the column headers within CSV files This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

CMD Mini-Explorer↗

Towards Resilient Near Real-Time Analysis Workflows in Fusion Energy Science

Nuclear fusion holds the promise of an endless source of energy. Several research experiments across the world and joint modeling and simulation efforts between the nuclear physics and high performance computing communities are actively preparing the operation of the International Thermonuclear Experimental Reactor (ITER). Both experimental reactors and their simulated counterparts generate data that must be analyzed quickly and in a resilient way to support decision making for the configuration of subsequent runs or prevent a catastrophic failure. However, the cost if the traditional techniques used to improve the resilience of analysis workflows, i.e., replicating datasets and computational tasks, becomes prohibitive with explosion of the volume of data produced by modern instruments and simulations. Therefore, we advocate in this paper for an alternate approach based on data reduction and data streaming. The rationale is that by allowing for a reasonable, controlled, and guaranteed loss of accuracy it becomes possible to transfer smaller amounts of data, shorten the execution time of analysis workflows, and lower the cost of replication to increase resilience. We develop our research and development roadmap towards resilient near real-time analysis workflows in fusion energy science and present early results showing that data streaming and data reduction is a promising way to speed up the execution and improve the resilience of analysis workflows.

Suter, Fred↗