Search NASA⌕ Search

SEARCH · Search NASA

Results for “database for machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Predictive Chemical Kinetic Modeling: Where We Succeed, Where We Struggle, and What Comes Next

Chemical kinetic modeling plays a foundational role in fields ranging from energy to environmental science, pharmaceuticals, and advanced materials. The past two decades have seen remarkable progress, particularly in modeling gas-phase reactions for thermochemical processes, leading to impactful industrial applications such as steam cracking and air quality management. However, new challenges are emerging. The successful development of systematic methodologies for the description of gas-phase kinetics opens the possibility to apply the same approach to the study of more challenging systems. Here, we review recent advances, including ab initio transition state theory-based master equation estimation of elementary rates, automated mechanism generation, machine-learning-assisted kinetics, and uncertainty quantification, and discuss the advances needed to apply the same methodological approach in areas such as heterogeneous catalysis, electrochemistry, liquid-phase and solid-state reactivity, and multiscale model integration. We advocate for the development of targeted tools, especially methods that go beyond empirical tuning toward first-principles-based predictions. We highlight the need for accessible software and AIaugmented workflows to democratize modeling for industry and academia alike. In this perspective, we call attention to not only what has worked but also what remains unsolved, advocating to avoid overemphasizing successes in scientific works at the expense of realism. The next decade should focus on predictive capability, physical accuracy, and community infrastructure (e.g., databases and services) to enable innovation across diverse fields. We argue that kinetic modeling, properly equipped, can accelerate discovery far beyond its traditional domains.

ab initio calculations↗

Machine learning for the redox potential prediction of molecules in organic redox flow battery

Here, organic redox flow batteries (ORFB) are recognized as an innovative technology for the large-scale storage of renewable energy. The redox potential of organic redox-active molecules plays a vital role in their performance. Advanced screening techniques like high-throughput experiment and machine learning (ML) have significantly enhanced organic material performance and transformed the field of ORFB. However, the scarcity of experimental data poses a considerable challenge for ML model development in this domain. In our study, we developed lightweight graph-based Gaussian process regression (GPR) models with GPU-accelerated marginalized graph kernel and hybrid kernel to predict the redox potentials of organic redox-active molecules for ORFBs, specifically focusing on small datasets. To evaluate model accuracy, we created a new experimental database of organic redox-active molecules by the data from hundreds of published papers and assembled previous computational datasets. We also considered some key parameters, such as pH conditions and solvent type, to assess their impact on redox potential prediction. Our GPR model predicted redox potentials with high accuracy across all datasets using minimal training data. The study provides powerful tools for molecule screening and design and delivers valuable guidance on designing training datasets for costly experiments.

25 ENERGY STORAGE↗

Chemistry Informed Machine Learning-Based Heat Capacity Prediction of Solid Mixed Oxides

Knowing heat capacity is crucial for modeling temperature changes with the absorption and release of heat and for calculating the thermal energy storage capacity of oxide mixtures with energy applications. The current prediction methods (ab initio simulations, computational thermodynamics, and the Neumann–Kopp rule) are computationally expensive, not fully generalizable, or inaccurate. Machine learning has the potential of being fast, accurate, and generalizable, but it has been scarcely used to predict mixture properties, particularly for mixed oxides. Here, we demonstrate a method for the generalizable prediction of heat capacity of solid oxide pseudobinary mixtures using heat capacity data obtained from computational thermodynamics and descriptors from ab initio databases. Further, models trained through this workflow achieved an error (mean absolute error of 0.43 J mol –1 K –1 ) lower than the uncertainty in differential scanning calorimetry measurements, and the workflow can be extended to predict other properties derived from the Gibbs free energy and for higher-order oxide mixtures.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Results and lessons learned from accelerating radio frequency modeling using machine learning [slides]

The “advanced tokamak” reactor concept is a leading candidate for a steady state fusion pilot plant. An advanced tokamak (AT) sustains a majority of the required plasma current with effects resulting from maintenance of the peaked pressure at the device center. This current is augmented by auxiliary current drive sources. These auxiliary actuators may consist of neutral particle beams and/or radio frequency (RF) systems such as lower hybrid current drive (LHCD) and high harmonic fast wave (HHFW) current drive using radio and microwaves from antennas. The primary focus of this work is to develop models of RF current profile control suitable for use in integrated modeling frameworks and for real-time control in experiments. Direct physics models of RF current drive can be computationally intensive. In order to achieve predictive times appropriate for the thousands of calls needed in real-time control of experiments and for use in integrated models, we will apply modern machine learning (ML) techniques to accelerate these models and interpolate their results. To generate the fast and accurate models for use in control level algorithms and integrated modeling we need to replace present models with high dimensional interpolation of their results. We will perform additional simulations across a broader parameter range for EAST and other tokamaks in different physics regimes (Alcator C-Mod, DIII-D, WEST, CFETR, ARC, ITER) and combine them into a larger database for training and testing of the ML models. Further testing of the control level models with experimental current profile data from EAST and C-Mod tokamaks will provide additional confirmation of the control level model before integration in a tokamak control system or integrated modeling suite. ML will be used to optimize the selection of training data consisting of RF current driven at different values of density profile, temperature profile, plasma current, and wavenumber. ML will also be used to facilitate classification of current drive from these input data. The output of this effort will be a validated classifier capable of determining the current drive profiles for HHFW CD and LHCD on a mille-second timescale. This will provide a breakthrough capability enabling real-time control of RF driven current profiles in experiments including ITER ICRF and use integrated modeling frameworks requiring thousands of current profile calculations in discharge simulations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Shock Hugoniot calculations using on-the-fly machine learned force fields with ab initio accuracy

We present a framework for computing the shock Hugoniot using on-the-fly machine learned force field (MLFF) molecular dynamics simulations. In particular, we employ an MLFF model based on the kernel method and Bayesian linear regression to compute the free energy, atomic forces, and pressure, in conjunction with a linear regression model between the internal and free energies to compute the internal energy, with all training data generated from Kohn–Sham density functional theory (DFT). We verify the accuracy of the formalism by comparing the Hugoniot for carbon with recent Kohn–Sham DFT results in the literature. In so doing, we demonstrate that Kohn–Sham calculations for the Hugoniot can be accelerated by up to two orders of magnitude, while retaining ab initio accuracy. We apply this framework to calculate the Hugoniots of 14 materials in the FPEOS database, comprising 9 single elements and 5 compounds, between temperatures of 10 kK and 2 MK. We find good agreement with first principles results in the literature while providing tighter error bars. In addition, we confirm that the inter-element interaction in compounds decreases with temperature.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Global compilation of soil methane uptake measurements from 1984 to 2018

This data package contains a global compilation of soil methane uptake measurements collected from published field studies between 1989 and 2022. The dataset was developed to support machine learning (ML) estimation of the global terrestrial methane soil sink and includes monthly methane uptake rates, measurement dates, site coordinates, and associated ecosystem information from different ecosystems. Data were compiled from 164 peer-reviewed publications across approximately 260 study sites, resulting in ~12,000 monthly observations after quality control screening and removal of manipulated experimental treatments. The database was further processed to generate site-averaged methane uptake estimates for comparison between process-based (PB) and ML models.

earth science↗

Trust Not Verify? The Critical Need for Data Curation Standards in Materials Informatics

The importance of data curation has been recognized in multiple areas of research; however, the discussion of this important issue is only beginning to emerge in materials science. In this Perspective, we highlight the benefits of using the standardized data curation protocols in materials science and discuss current gaps in accurate and reproducible data reporting using case studies drawn from high-impact materials science papers and well-known databases such as the Crystallography Open Database (COD) and the Cambridge Structural Database (CSD). We argue that both experimental and computational materials scientists need to embrace a culture of rigorous data curation as part of modern research data management. We propose a sample data curation pipeline for materials chemistry and illustrate its use by creating two new materials chemistry databases. Here, we hope that this perspective will serve to catalyze further discussion and promote the continuous development of rigorous data curation practices within the materials science research community. We posit that adherence to best practices of data curation will promote and enhance the reliability, reproducibility, and integrity of materials research and enable the development of reliable AI and machine learning models that critically depend on the use of quality data.

Chemical structure↗

Deployment of Traditional and Hybrid Machine Learning for Critical Heat Flux Prediction in the CTF Thermal-Hydraulics Code

Critical heat flux (CHF) marks the transition from nucleate to film boiling, where heat transfer to the working fluid can rapidly deteriorate. Accurate CHF prediction is essential for efficiency, safety, and preventing equipment damage, particularly in nuclear reactors. Although widely used, empirical correlations frequently exhibit discrepancies when compared to experimental data, limiting their reliability in diverse operational conditions. Traditional machine learning (ML) approaches have demonstrated potential for CHF prediction but often suffer from limited interpretability, data scarcity, and insufficient knowledge of physical principles. Hybrid model approaches, which combine data-driven ML with base models, mitigate these concerns by incorporating prior knowledge of the domain. This study integrates an externally trained purely data-driven ML model and two hybrid models (using the Biasi and Bowring CHF correlations) within the CTF subchannel code via a custom Fortran framework. Performance was evaluated using two validation cases: a subset of the Nuclear Regulatory Commission (NRC) CHF database and the Bennett dryout experiments. In both cases, the hybrid models demonstrated significantly lower error metrics compared to conventional empirical correlations, with the best models often reducing relative error by about 5 percentage points. The pure ML model achieved comparable accuracy, outperforming the hybrid Biasi model in the NRC test case (3.3% versus 5.5% relative error) but exhibiting slightly higher error against the hybrid Bowring model in the Bennett test case (7.7% versus 6.1%). Trend analysis of error parity indicated that ML-based models reduced the tendency for CHF overprediction, improving overall accuracy. These results demonstrate that ML-based CHF models can be effectively integrated into subchannel codes and could potentially increase performance compared to conventional methods.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Low Activity Waste Glass Optimization with Property Models from Machine Learning, Part 2: Experimental Validation and Active Learning

The United States Department of Energy is responsible for managing legacy nuclear waste stored in underground tanks at the Hanford Site. To treat the waste, it is planned as the current baseline to separately vitrify low-activity waste (LAW) and high-level waste fractions. Previously, machine learning (ML) based glass property models (e.g., chemical durability, viscosity, electrical conductivity and SO3 solubility) were developed with prediction uncertainties. A waste glass optimization approach was then established to enable the capability of using these ML models in LAW glass formulation. In this study, the previous ML models were first experimentally validated, and the results were incorporated back into the database to update the ML models. The updated models and formulations showed increased waste loading while reducing the failure rate, demonstrating improved predictive accuracy, reduced uncertainties, and the effectiveness of active learning in guiding high-dimensional, nonlinear LAW glass design. This represents the first experimental validation of ML based LAW glass formulation, with practical benefits such as higher waste loading, shorter mission duration, and lower operational risk.

Lu, Xiaonan (ORCID:0000000179708148)↗

Machine learning enhanced predictions of ICRF heating: Overcoming numerical limitations via data curation

In this work, we present the development of robust surrogate models for Ion Cyclotron Range of Frequencies (ICRF) and High-Harmonic Fast Wave (HHFW) heating predictions in fusion plasmas. Building upon our previous efforts to achieve real-time capable models, we identify the cause of the outliers found using TORIC in certain HHFW heating scenarios. The outliers are observed to be spurious ion Bernstein wave (IBW)-like modes caused by a wavelength control algorithm designed to address challenging scenarios with high perpendicular wavenumbers. The effect arises from the modulation in the perpendicular susceptibility, which can induce sign reversal and IBW-like propagation for scenarios featuring normalized ion Larmor radius λ i ≫ 1. We use TORIC with this algorithm disabled to generate a novel HHFW-NSTX database that is free of outliers. Surrogate models trained on this database, including Random Forest Regressor (RFR), Multi-Layer Perceptrons, and Gaussian Process Regressors (GPR), demonstrate the ability to accurately predict HHFW heating profiles, with regression scores of R 2 ∈[0.93−0.99]. Additionally we demonstrate that it is possible to generalize predictions beyond training data by the use of both RFR and GPR models, enabling the prediction of scenarios previously limited to the original model. GPR models also provide uncertainty quantification, offering insights into model confidence. This work introduces a comprehensive Verification, Validation, and Uncertainty Quantification methodology for surrogate modeling, applicable not only to ICRF heating but also to other RF heating challenges and fusion physics problems. Beyond accelerated inference, these models show effective extrapolation capabilities, providing an alternative for addressing numerical challenges.

Artificial neural networks↗

One-shot gas detection with transformer paired neural networks in Mako collected longwave infrared hyperspectral imagery

To date, careful data treatment workflows and statistical detectors are used to perform hyperspectral image (HSI) detection of any gas contained in a spectral library, which is often expanded with physics models to incorporate different spectral characteristics. In general, surrounding evidence or known gas-release parameters are used to provide confidence in or confirm detection capability, respectively. This makes quantifying detection performance difficult as it is nearly impossible to develop an absolute ground truth for gas target pixel presence in collected HSI. Consequently, developing and comparing new detection methods, especially machine learning (ML)-based methods, is susceptible to subjectivity in derived detection map quality. Here, in this work, we demonstrate the first use of transformer-based paired neural networks (PNNs) for one-shot gas target detection for multiple gases while providing quantitative classification and detection metrics for their use on labeled data. Terabytes of training data are generated from a database of long-wave infrared HSI obtained from historical Mako sensor campaigns over Los Angeles. By incorporating labels, singular signature representations, and a model development pipeline, we can tune and select PNNs to detect multiple gas targets that are not seen in training on a quantitative basis. We additionally assess our test set detections using interpretability techniques widely employed with ML-based predictors, but less common with detection methods relying on learned latent spaces.

Hyperspectral imaging↗

Addressing the Split Incentive Challenge for Enhanced Solar Adoption in Multifamily Rental Properties [Abstract]

The split incentive problem is particularly pronounced in rental markets, where landlords prioritize investments that directly increase property value or rental income. Since energy savings from solar photovoltaic (PV) systems primarily benefit tenants, landlords may perceive little return on investment unless mechanisms exist to recapture some of the financial gains. The primary objective of this project is to develop a publicly available, web-based tool to analyze the U.S. Department of Energy’s ResStock database, which models the U.S. residential building stock. The tool allows users to filter buildings by location, type, HVAC system, square footage, and other characteristics, and outputs typical electric load profiles. By leveraging location-specific electric load data, Fram Energy aims to advance business strategies that address the split incentive barrier and promote the adoption of solar PV installations in rental properties. In addition, a machine learning model will be developed to weigh the marginal contribution of building features across the dataset in predicting electricity demand, supporting guided decision making in forecasting electric load profiles. Lastly, based on each building’s location, load profile, and utility’s electricity rate, an optimized solar photovoltaic array and battery energy storage system will be sized to provide energy arbitrage opportunities.

14 SOLAR ENERGY↗

Grain boundary segregation in BCC vanadium-based alloys: Quantum-accurate computed segregation spectra and targeted experimental validations

Grain boundaries are critically important to the material performance of fusion reactor materials such as vanadium, particularly mechanical properties and irradiation resistance. A key challenge to the design and control of grain boundaries in vanadium alloys is the lack of quantitative data on grain boundary segregation. In this study, we combine computational and experimental methods to address this gap. Furthermore, using a machine learning-accelerated quantum mechanics/molecular mechanics approach, we calculated the segregation spectra for 28 transition metal elements in polycrystalline vanadium, and validated these predictions experimentally for a subset of solutes that sample a range of segregation behavior, specifically zirconium, titanium, and tungsten, using analytical transmission electron microscopy. Furthermore, the agreement between experiment and theory highlights the predictive capability of our approach. Critically, this work provides a comprehensive database of quantum-accurate solute segregation enthalpies in vanadium, enabling the development of advanced alloys for fusion reactors applications.

Fusion materials↗

Computationally evaluating high-yield metabolites for sustainable aviation fuel (SAF) using machine learning

The computational tool described in this report helps identify promising biological pathways that produce SAF platform molecules (either a drop-in SAF, or a precursor that can be easily converted to a drop-in SAF). The workflow the computational tool follows first identifies possible biological pathways from a user-defined metabolite. These pathways may, or may not lead to a SAF platform molecule, thus the second step involves insilico testing of the end product of each pathway to assess whether it is, or is not, a SAF platform molecule. The identification of biological pathways performed in the first step is facilitated by linking the metabolite to a biological reaction database. Pathways are found by identifying pathways in the reaction database that include the metabolite. The computational tool includes an alternative way to find pathways. The alternative way develops a Flux Balanced Analysis (FBA), and modifying the FBA to include reactions that transform the metabolite. These modifications serve as a basis for understanding, in a semi-quantitative way, if there is an increase in the flux to desirable products. The second step, in silico testing of the end-products, is accomplished by estimating key physical properties relevant to SAF. When good models are available, we have integrated those models into the computational tool. In a few instances, we have developed our own models. In all instances, we have validated the models against available measured data. Finally, we have evaluated the effectiveness of our computational tool by genetically engineering Rhodosporidium toruloides. Validation occurred without the use of a FBA, and further validation is required.

09 BIOMASS FUELS↗

Transformation rate maps of dissolved organic carbon in the contiguous US

Riverine dissolved organic carbon (DOC) plays a vital role in regional and global carbon cycles. However, the processes of DOC conversion from soil organic carbon (SOC) and leaching into rivers are insufficiently understood, inconsistently represented, and poorly parameterized, particularly in land surface and Earth system models. As a first attempt to fill this gap, we propose a generic formula that directly connects SOC concentration with DOC concentration in headwater streams, where a single parameter, the transformation rate from SOC in the soil to DOC leaching flux (P r ), accounts for the overall processes governing SOC conversion to DOC and leaching from soils (along with runoff) into headwater streams. We then derive high-resolution P r maps over the contiguous US (CONUS) using SOC data from two different sources: the Harmonized World Soil Database v1.2 (HWSD) and SoilGrids 2.0. Both maps are developed following the same five major steps: (1) selecting independent catchments where observed riverine DOC data are available with reasonable quality; (2) estimating catchment-average SOC for the independent catchments; (3) estimating the P r values for these catchments based on the generic formula and catchment-average SOC; (4) developing a predictive model of P r with machine learning (ML) techniques and catchment-scale climate, hydrology, geology, and other attributes; and (5) deriving a national map of P r based on the ML model. For evaluation, we compare the DOC concentration derived using the P r map and the observed DOC concentration values at evaluation catchments. The resulting mean absolute scaled error and coefficient of determination are 0.73 and 0.47 for the HWSD-based model and 0.58 and 0.72 for the SoilGrids-based model, respectively, suggesting the effectiveness of the overall methodology. Efforts to constrain uncertainty and evaluate sensitivity of P r to different factors are discussed. To illustrate the use of such maps, we derive a riverine DOC concentration reanalysis dataset over CONUS. The two P r maps, robustly derived and empirically validated, lay a critical cornerstone for better simulating the terrestrial carbon cycle in land surface and Earth system models. Our findings not only set a foundation for improving our predictive understanding of the terrestrial carbon cycle at the regional and global scales, but also hold promises for informing policy decisions related to decarbonization and climate change mitigation. The data presented in this study are publicly available at https://doi.org/10.5281/zenodo.14563816 (Li et al., 2024).

54 ENVIRONMENTAL SCIENCES↗

An AI-accelerated pathway for reproducible and stable halide perovskites

Halide perovskites (HPs) have remarkable optoelectronic properties, and in the last decade their photovoltaic power conversion efficiency and light-emitting diode efficiency have skyrocketed. Despite the surge in research on these burgeoning materials, two key challenges in the field remain: material irreproducibility and instability. Their behavior is especially dynamic in response to environmental stressors, due to complex interactions with the perovskite crystal lattice. Here, in this review, we survey the latest achievements in HP materials research accomplished with the assistance of artificial intelligence (AI), through the implementation of automated experimentation and machine learning (ML) data analysis. Automated synthesis and characterization tackle problems with material irreproducibility by systematically controlling parameters with very high precision, creating massive datasets, and allowing methodical comparisons from which unbiased conclusions can be drawn. AI can reveal otherwise unnoticed trends, inform future experiments with the highest potential information gain, and forecast future performance. The review concludes with a forward viewpoint of how human-assisted closed-loop laboratories and shared databases allow halide perovskite materials’ processing, properties, and performance to be potentially optimized with AI, accelerating the development of highly reproducible and stable optoelectronic devices.

Hering, Abigail R. [Univ. of California, Davis, CA↗

OEDI—Solar Grid Integration Data and Analytics Library

As a part of the Open Energy Data Initiative, this effort aims to develop and demonstrate novel distribution state estimation, control optimization, and transient analysis as well as provide access to data, data integration, and mapping information. More specifically, the focus of the effort will be on physics-based distribution system state estimation, hybrid (physics-based and machine learning) distribution optimal power flow, and event detection/analysis for solar integration and analytics. This work will enable reproducible, robust, replicable, and generalizable R&D in simulation and emulation of solar system integration. These test models and datasets will provide an integrated library for developing and testing power system operation technologies. To make the library user-friendly, this project will provide data curation tools such as data translators, mapping scripts and APIs, database schemas and metadata, interfaces and user dashboard, source code for the reference algorithms, description of the use-cases/scenarios, and comprehensive information on all the assumptions.

14 SOLAR ENERGY↗

Spotlight: efficient automated global optimization in rietveld analysis of diffraction data

Performing reliable Rietveld analysis on tens or hundreds of powder diffraction datasets from parametric or time-resolved experiments often poses a bottleneck in extracting meaningful results from the data. While automated analysis of data has recently been demonstrated, high temperature annealing studies, during which phase transformations occur and lattice parameters may change due to repartitioning of elements, are prime examples where automation by a simple phase identification from a database of room temperature structures or automation by sequential refinements is likely to fail. To enable reliable, efficient, automated Rietveld analysis, we present a Python package named Spotlight , building on established Rietveld packages such as MAUD, GSAS , or GSAS-II , which extends the refinement of best fit parameters to a global optimization using an ensemble of optimizers leveraging hierarchical parallel execution on high-performance computing clusters. Spotlight further enables the efficient design of refinement plans through the iterative automated machine-learning of a surrogate for the refinement on which the global optimizations are performed until results from the surrogate converge to the response surface data. We demonstrate Spotlight with the analysis of uranium molybdenum and Ti–6Al–4V datasets, as well as in two open-source tutorials analyzing aluminium oxide and lead sulphate.

36 MATERIALS SCIENCE↗