Search NASA⌕ Search

SEARCH · Search NASA

Results for “Generative Machine Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

A CHIL Validation of Machine Learning-Assisted Methods for Real-Time Controls of Solar PV for Grid Services: Preprint

Recent research has highlighted the potential for solar to act as a zero-marginal-cost and zero-emission flexibility resource on the bulk power system when operated with advanced control systems. To increase the performance of these systems, leading technologies, including machine learning (ML) and hierarchical inverter set point allocation, have been proposed; however, these technologies lack comprehensive validation under real-world application scenarios. This paper addresses this gap by designing and developing a controller-hardware-in-the-loop framework to evaluate the performance of different flexible solar technologies in responding to automatic generation control signals in a closed-loop fashion. Simulation results indicate the superior performance of an ML-based approach compared to the conventional reference-control grouping-based approach, showcasing its potential to support grid stability and operational efficiency.

closed-loop validation↗

2019 Budget Request for the DOE Computational Science Graduate Fellowship (CSGF) Grant

The Department of Energy Computational Science Graduate Fellowship (DOE CSGF) is necessary to meet the continual challenging national workforce needs that arise as computational science and engineering problems continue to grow in scope and complexity. Computational science and engineering (CSE) is a multidisciplinary approach that uses scientific computing to solve practical problems methods and to supply technical tools across the scientific discovery spectrum. In particular, the DOE CSGF emphasizes high-performance computing (HPC) that enables CSE that advances science and engineering in directions important to the DOE and the economy in general. Over the past half-century, HPC has been an essential tool for DOE’s success. During this period, important missions, such as nuclear stockpile stewardship, have turned to HPC as an essential technology. Entire science disciplines, such as biology and cosmology, have been transformed through the augmentation of scientific observation via HPC. At government laboratories and in industry, DOE CSGF alumni are helping push traditional HPC boundaries while contributing to discoveries in high-energy physics, renewable energy, fusion-reactor design, additive manufacturing, nanomaterials for next-generation batteries and transistors, and turbine and advanced nuclear reactor modeling. In addition, HPC is used to address national health needs that will eventually point to cures both by helping cancer researchers manage and analyze huge troves of data, by simulating biological mechanisms, and by accelerating drug development — including continuing to rise to the challenge of pandemic-related research. A 2023 report from the ASCAC Subcommittee on American Competitiveness and Innovation to the ASCR office, “Can the United States Maintain Its Leadership in High-Performance Computing?” says of the Program, “The CSGF program provides a barometer for disciplines that will be of interest to future DOE computing.” An explosion in scientific and technological data has driven the need for increasingly sophisticated HPC to transform those data into scientific understanding. With access to more and more data and the proliferation of HPC, Machine Learning and Artificial Intelligence are experiencing a renaissance, complementing the now well-established use of computational simulation. Indeed, in its September 2020 subcommittee report on “AI/ML, Data Intensive Science and High-Performance Computing”, the DOE Advanced Scientific Computing Advisory Committee (ASCAC) explicitly called for a fellowship program to train computational and data scientists to tackle exascale and data-intensive computing challenges. This collaboration of empirical and theory-based modeling will increasingly inform federal policymakers whose decisions affect American society and future generations, and it requires highly skilled and intellectually agile computational scientists who can support the fast-moving DOE National Laboratory research environment. In fact, the DOE CSGF program has explicitly and consistently addressed this need.

97 MATHEMATICS AND COMPUTING↗

Recent advances in rational design of defect-engineered photocatalysts toward sustainable NH 3 synthesis as H 2 carrier: From fundamental and development to machine-learning

In this study, we provide a detailed overview of the fundamental mechanisms underpinning photocatalytic N 2 reduction. We also discuss advances in catalyst design for the synthesis of NH 3 . Particular emphasis is placed on the role of surface defect engineering, which includes the creation of surface defects to enhance the performance of semiconducting photocatalysts for efficient N 2 reduction. In addition, the application of a machine learning-based computational modeling approach is discussed as an important driving force for predicting and regulating NH 3 synthesis efficiency based on catalyst features and reaction conditions. Finally, existing challenges and future perspectives for improving the performance of defect-engineered photocatalysts are outlined to contribute to the ongoing discourse on sustainable ammonia generation. This review aims to clarify recent progress in the rational design of defect-containing photocatalysts for the synthesis of NH 3 and encourages innovative approaches to catalyst optimization rather than solely focusing on new materials.

08 HYDROGEN↗

Identifying Outliers in AI-based Image Compression

Image compression using artificial intelligence (AI) is becoming increasingly prevalent across various fields, including scientific research. Scientific instruments can generate hundreds of images per second, and effectively compressing these images with high compression ratios is crucial for facilitating scientific discoveries. However, automatically detecting outlier cases, where compression may not have succeeded or where interesting scientific phenomena are present, poses a significant challenge. To address this, we have developed a methodology based on unsupervised machine learning techniques for detecting outlier compressed images. This methodology utilizes metrics such as peak signal-to-noise ratio (PSNR), structural similarity index measure (SSIM), structural texture similarity index measure (STSIM), and deep image and structural texture similarity index (DISTS). We have evaluated our methodology on several unlabeled datasets, including microscopy and x-ray images, and have successfully identified multiple outlier images using our proposed approach. Furthermore, our approach has enabled us to identify image semantics that are valuable for post-experiment analysis by scientists.

Data Analysis↗

DOE FAIR Surrogate Benchmarks Supporting AI and Simulation Research (SBI Surrogate Benchmark Initiative) (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia(UVA). SBI repositories include data, code, and all relevant collateral artifacts, that the science and engineering community needs to use and reuse these data sets and surrogates. SBI repositories generate active research from both participants in SBI and the broader AI and domain science communities. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and capture them as surrogate benchmarks with a rich set of metadata, covering. Data; Model; Metrics specification; Machine specification; Science, Speed, Power Results, We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non Surrogate benchmarks that have many common features and similar issues regarding FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, benchmarks have datasets, models, and metadata, and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates, including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

Scalable learning of potentials to predict time-dependent Hartree–Fock dynamics

We propose a framework to learn the time-dependent Hartree–Fock (TDHF) inter-electronic potential of a molecule from its electron density dynamics. Although the entire TDHF Hamiltonian, including the inter-electronic potential, can be computed from first principles, we use this problem as a testbed to develop strategies that can be applied to learn a priori unknown terms that arise in other methods/approaches to quantum dynamics, e.g., emerging problems such as learning exchange–correlation potentials for time-dependent density functional theory. We develop, train, and test three models of the TDHF inter-electronic potential, each parameterized by a four-index tensor of size up to 60 × 60 × 60 × 60. Two of the models preserve Hermitian symmetry, while one model preserves an eight-fold permutation symmetry that implies Hermitian symmetry. Across seven different molecular systems, we find that accounting for the deeper eight-fold symmetry leads to the best-performing model across three metrics: training efficiency, test set predictive power, and direct comparison of true and learned inter-electronic potentials. All three models, when trained on ensembles of field-free trajectories, generate accurate electron dynamics predictions even in a field-on regime that lies outside the training set. To enable our models to scale to large molecular systems, we derive expressions for Jacobian-vector products that enable iterative, matrix-free training.

97 MATHEMATICS AND COMPUTING↗

Molecular simulation using transfer-learned potentials for the disordered nanoscale structure of nitrogen-doped nanoporous carbons

Machine learning (ML)-based molecular dynamics (MD) simulations of the formation of a class of N-doped nanoporous carbons are performed to assess their disordered partially graphitized nanoscale structure. The study is motivated by the effectiveness of so-called nitrogen assembly carbons (NACs) for catalysis applications. Benchmark simulations for pure-C disordered graphitic systems reveal the importance of reliably capturing the vdW component of the potentials in order to accurately describe the tendency for layering of disordered graphene-like sheets. In our modeling, this is achieved by a transfer learning strategy incorporating features of the energetics from the optB88-vdW DFT functional into potentials initially trained with a less expensive functional, thereby providing a superior description of the pure-C systems. Generation from MD simulations of realistic partially graphitized structures is significantly more challenging for N-doped versus for pure C systems. However, such structures are achieved by a tailored MD simulation protocol mimicking the experimental synthesis process and in particular incorporating an annealing and subsequent quenching stages. Simulated PXRD patterns effectively reproduce the features of experimental observations for NACs, including the appearance of a prominent but broad (002) peak at around 25, and the development of another weaker feature associated with in-layer ordering of mixed C-N graphene-like sheets.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

HTESP (High-throughput electronic structure package): A package for high-throughput ab initio calculations

High-throughput ab initio calculations are the indispensable parts of data-driven discovery of new materials with desirable properties, as reflected in the establishment of several online material databases. The accumulation of extensive theoretical data through computations enables data-driven discovery by constructing machine learning and artificial intelligence models to predict novel compounds and forecast their properties. Efficient usage and extraction of data from these existing online material databases can accelerate the next stage materials discovery that targets different and more advanced properties, such as electron–phonon coupling for phonon-mediated superconductivity. However, extracting data from these databases, generating tailored input files for different ab initio calculations, performing such calculations, and analyzing new results can be demanding tasks. Here, in this work, we introduce a software package named “HTESP” (High-Throughput Electronic Structure Package) written in Python and Bash languages, which automates the entire workflow including data extraction, input file generation, calculation submission, result collection and plotting. Our HTESP will help speed up future computational materials discovery processes.

36 MATERIALS SCIENCE↗

Wasatch Fault Structure from Machine Learning Arrival Times and High-Precision Earthquake Locations

Abstract On 18 March 2020, a magnitude 5.7 earthquake hit the Salt Lake valley in the state of Utah, United States. Using a dense geophone deployment and machine learning (ML), an additional several thousand events were detected and located. Currently, both the mainshock and the majority of the aftershocks are suspected to have occurred on or near a deeper portion of the Salt Lake segment of the Wasatch fault—part of a large range-bounding fault system thought to be capable of generating an Mw 7.2 earthquake. However, a small subset of aftershocks may have occurred on a portion of the more steeply, eastward dipping, and poorly understood West Valley fault. Unfortunately, the catalog locations and lack of focal mechanisms for this subset of aftershocks provide only a crude constraint on the true fault structure. To better illuminate fault structure, we relocate the ML-generated catalog with a range of magnitudes from −2 to 4.6, using: (1) NonLinLoc, a nonlinear location algorithm, (2) source-specific station terms, and (3) waveform coherence. We further compute first-motion focal mechanisms for 68 events. Results of the relocation suggest a simpler, minimally listric Wasatch fault geometry, contrary to what has been previously proposed. We also find that analysis of the focal mechanisms and waveform similarity indicates minimal event similarity throughout the Magna sequence, suggesting a highly complex and heterogeneous rupture zone, as opposed to rupture on a single plane. These findings suggest an increased seismic hazard due to the overall shallowness of the earthquake sequence and highly varied rupture mechanisms.

Geochemistry & Geophysics↗

Improving Detector Systematic Uncertainties Through Data-Driven Machine Learning

Detector simulation in liquid argon time projection chambers (LArTPCs) is a constant challenge. In particular the modeling of electrons response on wires is highly nontrivial. However, new machine learning techniques exist which can be leveraged to ameliorate these concerns. We present a novel methodology to attempt to learn from cosmic muon data in the ICARUS detector how reconstructed wire signals are influenced by features of the hits such as location, particle direction, angle relative to the wire plane, etc. A model can then be generated which can apply the learned mapping to Monte Carlo events to create a more data-like simulation sample. By creating such a sample we expect to reduce the systematic uncertainties at ICARUS due to our detector modeling.

Hausner, Harry [Fermilab] (ORCID:0000000188932280)↗

Improving Detector Systematic Uncertainties Through Data-Driven Machine Learning

Detector simulation in liquid argon time projection chambers (LArTPCs) is a constant challenge. In particular the modeling of electrons response on wires is highly nontrivial. However, new machine learning techniques exist which can be leveraged to ameliorate these concerns. We present a novel methodology to attempt to learn from cosmic muon data in the ICARUS detector how reconstructed wire signals are influenced by features of the hits such as location, particle direction, angle relative to the wire plane, etc. A model can then be generated which can apply the learned mapping to Monte Carlo events to create a more data-like simulation sample. By creating such a sample we expect to reduce the systematic uncertainties at ICARUS due to our detector modeling.

Hausner, Harry [Fermilab] (ORCID:0000000188932280)↗

Data and scripts associated with a manuscript analyzing ELM-FATES parameter sensitivity under pre-fire and postfire scenarios using machine learning

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “Fire Severity-Dependent Shifts in Vegetation Parameter Sensitivity: A Pre- and Post-Fire Analysis Using ELM-FATES and Explainable AI” submitted to Journal of Advances in Modeling Earth Systems (Zahura et al. 2026). The study examines vegetation physiological parameters controlling pre-fire and post-fire vegetation dynamics. To support this analysis, 73 vegetation parameters in Functionally Assembled Terrestrial Ecosystem Simulator (FATES) (Fisher et al., 2018) , which is coupled with E3SM (Energy Exascale Earth System Model) land model (ELM, ELM-FATES), were perturbed using a Sobol sequence to generate 1,024 ensemble members for two plant functional types: needleleaf evergreen extratropical trees (NEET) and C3 grass. Simulations were conducted for the pre-fire period (2016) and post-fire period (2018–2023). Burn severity was represented by modifying the Nesterov index in FATES to 75,000, 150,000, and 300,000 for low, moderate, and high severity, respectively. A no-fire scenario was also included. Simulations were performed for 16 grid cells in the American River Watershed across different burn severities and plant functional types. XGBoost (eXtreme Gradient Boosting) models were trained using the parameter ensembles and ELM-FATES-simulated outputs, including leaf area index (LAI), gross primary productivity (GPP), aboveground biomass, vegetation evaporation, transpiration, and soil evaporation. Models were trained separately for each year and burn severity, followed by SHAP (SHapley Additive exPlanations) analysis to identify changes in dominant parameters after fire disturbance. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. The data package contains the ELM-FATES simulation data. The scripts and data related to the analysis will be added later. The inputs and outputs from ELM-FATES are inside the “FATES” folder. “FATES_domain_surface” contains the domain and surface netcdfs that were used to run ELM-FATES in the study area. “FATES_parameters” contains the 1024 ensembles that were generated using Sobol sequence. “FATES_outputs” folder contains ELM-FATES simulated variables. All files are .csv and .nc (NetCDF).

Aboveground biomass↗

Persistent global greening over the last four decades using novel long-term vegetation index data with enhanced temporal consistency

Advanced Very High-Resolution Radiometer (AVHRR) satellite observations have provided the longest global daily records from 1980s, but the remaining temporal inconsistency in vegetation index datasets has hindered reliable assessment of vegetation greenness trends. To tackle this, we generated novel global long-term Normalized Difference Vegetation Index (NDVI) and Near-Infrared Reflectance of vegetation (NIRv) datasets derived from AVHRR and Moderate Resolution Imaging Spectroradiometer (MODIS). We addressed residual temporal inconsistency through three-step post processing including cross-sensor calibration among AVHRR sensors, orbital drifting correction for AVHRR sensors, and machine learning-based harmonization between AVHRR and MODIS. After applying each processing step, we confirmed the enhanced temporal consistency in terms of detrended anomaly, trend and interannual variability of NDVI and NIRv at calibration sites. Our refined NDVI and NIRv datasets showed a persistent global greening trend over the last four decades (NDVI: 0.0008 yr -1 ; NIRv: 0.0003 yr -1 ), contrasting with those without the three processing steps that showed rapid greening trends before 2000 (NDVI: 0.0017 yr -1 ; NIRv: 0.0008 yr -1 ) and weakened greening trends after 2000 (NDVI: 0.0004 yr -1 ; NIRv: 0.0001 yr -1 ). These findings highlight the importance of minimizing temporal inconsistency in long-term vegetation index datasets, which can support more reliable trend analysis in global vegetation response to climate changes.

54 ENVIRONMENTAL SCIENCES↗

Next-Generation Materials Design: Quantum Mechanics and Data-Driven Modeling

The future of materials design is rapidly advancing through the combination of quantum mechanics and data-driven modeling. These approaches integrate quantum principles with advanced data analysis, enabling precise insights into material behavior. This talk will highlight recent progress in using these methods for computational design, particularly in high-entropy alloy catalysts, emphasizing the role of hierarchical machine-learning architectures for accurate predictions. Additionally, I will discuss our work on developing machine learning interatomic potentials (MLPs) for single-element metals, metal oxides, and alloys under extreme conditions, focusing on melting behavior and phase properties at high temperatures and pressures. We have also refined our MLP models to capture dynamic surface interactions, such as CO2 and CO adsorption on MgO, using both static and molecular dynamics simulations. These models maintain high accuracy while significantly reducing computational costs compared to first-principles calculations. By enabling efficient and accurate simulations, this work supports broader community adoption, optimizes datasets for materials discovery, and extends the accessible time, size, and environmental conditions beyond the limits of experiments and traditional simulations.

machine learning↗

FAIR Surrogate Benchmarks Supporting AI and Simulation Research (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia (UVA). SBI repositories include data, code, and all relevant collateral artifacts that the science and engineering community need to use and reuse these data sets and surrogates. SBI repositories generate active research from both the participants in SBI and the broad community of AI and domain scientists. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and captures them as surrogate benchmarks with a rich set of metadata covering: Data; Model; Metrics specification; Machine specification; and Science, Speed, and Power Results. We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non-Surrogate benchmarks that have many common features and similar issues as regards FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, Benchmarks have datasets, models, and metadata and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

Open Power System Datasets and Open Simulation Engines: A Survey Toward Machine Learning Applications

A major factor behind the success of machine learning (ML) models in multiple domains is the availability and accessibility of large, labeled, and well-organized datasets for training and benchmarking. In comparison, power grid datasets face three major challenges: (i) real-world data is often restricted by regulatory constraints, privacy reasons, or security concerns, making it difficult to obtain and work with; (ii) synthetic datasets, which are created to address these limitations, often have incomplete information and are released using specialized tools, making them inaccessible to the broader community; and, (iii) input-output datasets are difficult to generate through simulation for non-experts because open-source simulators are not known outside the power system community. This survey addresses these challenges by serving as an entry point to publicly available datasets and simulators for researchers venturing in this area. We review the current landscape of open-source power network data, machine models, consumer demand profiles, renewable generation data, and inverter models. We also examine open-source power system simulators, which are crucial for generating high-quality, high-fidelity power grid datasets. We aim to provide a foundation for overcoming data scarcity and advance towards a structured web of datasets and simulators to support the development of ML for power systems.

42 ENGINEERING↗

An investigation on machine learning predictive accuracy improvement and uncertainty reduction using VAE-based data augmentation

The confluence of ultrafast computers with large memory, rapid progress in Machine Learning (ML) algorithms, and the availability of large datasets place multiple engineering fields at the threshold of dramatic progress. However, a unique challenge in nuclear engineering is data scarcity because experimentation on nuclear systems is usually more expensive and time-consuming than most other disciplines. One potential way to resolve the data scarcity issue is deep generative learning, which uses certain ML models to learn the underlying distribution of existing data and generate synthetic samples that resemble the real data. In this way, one can significantly expand the dataset to train more accurate predictive ML models. In this study, our objective is to evaluate the effectiveness of data augmentation using variational autoencoder (VAE)-based deep generative models. We investigated whether the data augmentation leads to improved accuracy in the predictions of a deep neural network (DNN) model trained using the augmented data. Additionally, the DNN prediction uncertainties are quantified using Bayesian Neural Networks (BNN) and conformal prediction (CP) to assess the impact on predictive uncertainty reduction. To test the proposed methodology, we used TRACE simulations of steady-state void fraction data based on the NUPEC Boiling Water Reactor Full-size Fine-mesh Bundle Test (BFBT) benchmark. Here, we found that augmenting the training dataset using VAEs has improved the DNN model’s predictive accuracy, improved the prediction confidence intervals, and reduced the prediction uncertainties.

Bayesian neural network↗