Search NASA⌕ Search

SEARCH · Search NASA

Results for “generative machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

Machine Learning Application to Atmospheric Chemistry Modeling

Atmospheric chemistry models are a central tool to study the impact of chemical constituents on the environment, vegetation and human health. These models split the atmosphere in a large number of grid-boxes and consider the emission of compounds into these boxes and their subsequent transport, deposition, and chemical processing. The chemistry is represented through a series of simultaneous ordinary differential equations, one for each compound. Given the difference in life-times between the chemical compounds (milli-seconds for O (sup 1) D (Deuterium) to years for CH4) these equations are numerically stiff and solving them consists of a significant fraction of the computational burden of a chemistry model. We have investigated a machine learning approach to emulate the chemistry instead of solving the differential equations numerically. From a one-month simulation of the GEOS-Chem model we have produced a training dataset consisting of the concentration of compounds before and after the differential equations are solved, together with some key physical parameters for every grid-box and time-step. From this dataset we have trained a machine learning algorithm (regression forest) to be able to predict the concentration of the compounds after the integration step based on the concentrations and physical state at the beginning of the time step. We have then included this algorithm back into the GEOS-Chem model, bypassing the need to integrate the chemistry. This machine learning approach shows many of the characteristics of the full simulation and has the potential to be substantially faster. There are a wide range of application for such an approach - generating boundary conditions, for use in air quality forecasts, chemical data assimilation systems, etc. We discuss speed and accuracy of our approach, and highlight some potential future directions for improving it.

Keller, Christoph A.↗

Atmospheric Chemistry Modeling and Air Quality Forecasting Using Machine Learning

Atmospheric chemistry models are a central tool to study the impact of chemical constituents on the environment, vegetation and human health. These models split the atmosphere in a large number of grid-boxes and consider the emission of compounds into these boxes and their subsequent transport, deposition, and chemical processing. The chemistry is represented through a series of simultaneous ordinary differential equations, one for each compound. Given the difference in life-times between the chemical compounds (milli-seconds for O1D to years for CH4) these equations are numerically stiff and solving them consists of a significant fraction of the computational burden of a chemistry model.We have investigated a machine learning approach to emulate the chemistry instead of solving the differential equations numerically. From a one-month simulation of the GEOS-Chem model we have produced a training dataset consisting of the concentration of compounds before and after the differential equations are solved, together with some key physical parameters for every grid-box and time-step. From this dataset we have trained a machine learning algorithm (regression forest) to be able to predict the concentration of the compounds after the integration step based on the concentrations and physical state at the beginning of the time step. We have then included this algorithm back into the GEOS-Chem model, bypassing the need to integrate the chemistry.This machine learning approach shows many of the characteristics of the full simulation and has the potential to be substantially faster. There are a wide range of application for such an approach - generating boundary conditions, for use in air quality forecasts, chemical data assimilation systems, etc. We discuss speed and accuracy of our approach, and highlight some potential future directions for improving it.

Keller, Christoph A.↗

Toward machine learning interatomic potentials for modeling uranium mononitride

Uranium mononitride (UN) is a promising accident-tolerant fuel because of its high fissile density and high thermal conductivity. In this study, we developed the first machine learning interatomic potentials for reliable atomic-scale modeling of UN at finite temperatures. We constructed a training set using density functional theory (DFT) calculations that was enriched through an active learning procedure, and two neural network potentials were generated. Both potentials successfully reproduce key thermophysical properties of interest, such as temperature-dependent lattice parameter, specific heat capacity, and bulk modulus. We also evaluated the energy of stoichiometric defect reactions and defect migration barriers and found close agreement with DFT predictions, demonstrating that our potentials can be used for modeling defects in UN. Additional tests provide evidence that our potentials are reliable for simulating diffusion, noble gas impurities, and radiation damage.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Developing a digital twin framework for remotely monitoring nuclear reactor facilities

A digital twin must seek to represent all applicable functional components of the system of interest. Different expertise is required for understanding the physical system being modeled than the skills needed for transforming those models into a functional digital twin through physics modeling, machine learning analysis, and visualization. The diversity of knowledge requires a multi-disciplinary team to ensure all system details are captured. Team members also need a method to verify that the data they generate within their domain can be effectively communicated to professionals in other fields. To address this challenge, this work provides an approach for developing a digital twin framework to remotely monitoring nuclear facilities. Through this, general knowledge of the framework is presented along with two examples to solidify the process. The AGN-201 digital twin and microreactor digital twins provide varying levels of complexity in a potential nuclear facility, where common threads are identified and lessons learned are provided. The goal of this research is to aid future researchers by providing a formula for a successful digital twin and in turn reducing the development time of nuclear system digital twins, specifically for remote monitoring.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

RU Net for Automatic Characterization of TRISO Fuel Cross Sections

TRistructural ISOtropic (TRISO) particle fuel is a type of nuclear fuel known for its high-temperature and high-burnup performance. Each sub-millimeter diameter TRISO particle consists of uranium-oxycarbide (UCO) or UO2 fuel kernel, coated with buffer, inner pyrolytic carbon (IPyC), silicon carbide (SiC), and outer pyrolytic carbon (OPyC) layers. The SiC layer acts as the main containment barrier for the TRISO particle to retain the fission products, while the IPyC and OPyC layers provide additional barriers to the release of fission products, especially fission gases. During irradiation, phenomena like kernel swelling, buffer densification, and IPyC fracture may impact fuel performance. Post-irradiation microscopy on entire compact cross sections or samples of individual particles deconsolidated from compacts is often used to identify these irradiation-induced changes in morphology. However, each fuel compact generally contains thousands of TRISO particles. To get statistical information on these phenomena, it is cumbersome work if done manually. For example, to get information about swelling/densification behaviors of different layers or kernels after irradiation, researchers previously manually measured the perimeter of each TRISO layer in hundreds of particles after four rounds of iterative grinding and polishing encompassing more than 2000 cross-section images for a total of four fuel compacts. To attempt to reduce the subjectivity inherent in that process and accelerate data analysis, we conducted a study on the automatic TRISO layer segmentation on cross-sectional microscopic images using Convolutional Neural Networks (CNNs). CNNs are a class of machine learning algorithms specifically designed for processing structured grid data that have gained popularity in recent years due to their remarkable performance in various computer vision tasks, including image classification, object detection, and image segmentation. In this research, we have generated the large irradiated TRISO layer dataset with more than 2000 cross-section TRISO microscopic images and the corresponding annotated images. Based on these annotated images, we have employed different CNNs for automatic segmentation of different TRISO layers. These include RU-Net (developed in this study), as well as three existing architectures: U-Net, Residual Network (ResNet), and Attention U-Net. The preliminary results show that the model based on RU-Net has the best performance in terms of intersection-over-union (IoU). Through the aid of these CNN models, we can expedite the analysis of TRISO particle cross-sections, significantly reducing the manual labor involved and improving the objectivity of the segmentation results.

Convolutional Neural Networks↗

Multi-fidelity equations of state and transport coefficient datasets for pulsed-power applications

Reliably simulating experiments relevant to the National Nuclear Security Administration (NNSA) requires a detailed description of material properties across a wide range of conditions. Such properties include the equations of state, charged-particle transport coefficients, and optical properties like the opacity. Together, these properties make up the material models used in radiation-magnetohydrodynamic simulations of nuclear fusion experiments. Many of these models do not incorporate uncertainties in the data used to produce them. It is unknown whether these uncertainties significantly impact the interpretation of simulation results and diagnostics. The purpose of this work is to quantify how such uncertainties impact simulations of pulsed-power experiments. We accomplished this task by first assessing discrepancies between approaches used to generate the data. This included bringing together members of the high-energy-density community spanning the three NNSA laboratories and multiple universities. Then, using these data, we developed a general framework that systematically incorporates physical uncertainties within the material models suitable for uncertainty quantification analyses. The framework utilizes machine learning, Bayesian inference, and incorporates multi-fidelity datasets. We demonstrated the framework by quantifying the impact that material model uncertainties have on simulations of pulsed-power experiments underway on Z at Sandia National Laboratories. As a result of this work, we discovered that modest uncertainties in material models (roughly 20%) correspond to significant uncertainties in the outputs from simulations. Our framework has enabled rapid construction of material models through an automated procedure and allows for the generation of material models of interest to the NNSA.

36 MATERIALS SCIENCE↗

Inverse design of photonic surfaces via multi fidelity ensemble framework and femtosecond laser processing

We demonstrate a multi-fidelity (MF) machine learning ensemble framework for the inverse design of photonic surfaces, trained on a dataset of 11,759 samples that we fabricate using high throughput femtosecond laser processing. The MF ensemble combines an initial low fidelity model for generating design solutions, with a high fidelity model that refines these solutions through local optimization. The combined MF ensemble can generate multiple disparate sets of laser-processing parameters that can each produce the same target input spectral emissivity with high accuracy (root mean squared errors < 2%). SHapley Additive exPlanations analysis shows transparent model interpretability of the complex relationship between laser parameters and spectral emissivity. Finally, the MF ensemble is experimentally validated by fabricating and evaluating photonic surface designs that it generates for improved efficiency energy harvesting devices. Our approach provides a powerful tool for advancing the inverse design of photonic surfaces in energy harvesting applications.

97 MATHEMATICS AND COMPUTING↗

Automated annotation of scientific texts for ML-based keyphrase extraction and validation

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lack the essential metadata required for researchers to find, curate, and search them effectively. The lack of metadata poses a significant challenge in the utilization of these data sets. Machine learning (ML)–based metadata extraction techniques have emerged as a potentially viable approach to automatically annotating scientific data sets with the metadata necessary for enabling effective search. Text labeling, usually performed manually, plays a crucial role in validating machine-extracted metadata. However, manual labeling is time-consuming and not always feasible; thus, there is a need to develop automated text labeling techniques in order to accelerate the process of scientific innovation. This need is particularly urgent in fields such as environmental genomics and microbiome science, which have historically received less attention in terms of metadata curation and creation of gold-standard text mining data sets. In this paper, we present two novel automated text labeling approaches for the validation of ML-generated metadata for unlabeled texts, with specific applications in environmental genomics. Our techniques show the potential of two new ways to leverage existing information that is only available for select documents within a corpus to validate ML models, which can then be used to describe the remaining documents in the corpus. The first technique exploits relationships between different types of data sources related to the same research study, such as publications and proposals. The second technique takes advantage of domain-specific controlled vocabularies or ontologies. In this paper, we detail applying these approaches in the context of environmental genomics research for ML-generated metadata validation. Our results show that the proposed label assignment approaches can generate both generic and highly specific text labels for the unlabeled texts, with up to 44% of the labels matching with those suggested by a ML keyword extraction algorithm.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Decentralized Microgrid Protection Through Relative Fault Direction Classification: Preprint

Protection in inverter-based resources (IBRs) dominated microgrids generally face significant challenges due to the low fault current and inconsistent fault behaviors from IBRs. Recently, machine learning-based approaches have attracted considerable attention to address these challenges. This paper introduces a novel decentralized protection strategy for microgrids. The proposed method decomposes the protection challenge into several distributed learning tasks, enabling individual relays to autonomously determine the direction of faults using a binary classification framework based on support vector machine (SVM) algorithms. Following the distributed fault direction estimation, classifier outcomes are shared among neighboring relays, facilitating a local decision-making process to ascertain the presence of faults within the neighborhood. Finally, a tripping signal is generated based on the classifier results of each relay to operate the circuit breaker. To test and validate this approach, a 100% renewable microgrid model is simulated in MATLAB/Simulink. In the numerical analysis, the application of SVM classifiers in our approach yields impressive results: an average relay classification accuracy of 98%, and a 96% accuracy in circuit breaker control. These findings highlight the potential of machine-learning-based approaches in enhancing the efficiency and reliability of microgrid protection systems.

decentralized algorithm↗

The role of AI in detecting and mitigating human errors in safety-critical industries: A review

For safety-critical industries, human error (HE) presents continual risks to system productivity, reliability and safety. Artificial intelligence (AI) and machine learning (ML) methods have emerged as promising approaches to understand, categorize and mitigate the risk of HE in safety-critical industries. Furthermore, this review offers an examination of the current landscape regarding the utilization of AI/ML with regards to HE in safety-critical industries, categorizing literature into descriptive modeling, predictive modeling, prescriptive modeling, and generative modeling techniques. Additionally, the review aims to provide insights regarding themes in literature, challenges, and future research directions. Findings of the review suggest that AI/ML methods can prove useful in addressing the HE problem across safety-critical industries.

42 ENGINEERING↗

Modeling and observations of North Atlantic cyclones: Implications for U.S. Offshore wind energy

To meet the Biden-Harris administration's goal of deploying 30 GW of offshore wind power by 2030 and 110 GW by 2050, expansion of wind energy into U.S. territorial waters prone to tropical cyclones (TCs) and extratropical cyclones (ETCs) is essential. This requires a deeper understanding of cyclone-related risks and the development of robust, resilient offshore wind energy systems. Here, this paper provides a comprehensive review of state-of-the-science measurement and modeling capabilities for studying TCs and ETCs, and their impacts across various spatial and temporal scales. We explore measurement capabilities for environments influenced by TCs and ETCs, including near-surface and vertical profiles of critical variables that characterize these cyclones. The capabilities and limitations of Earth system and mesoscale models are assessed for their effectiveness in capturing atmosphere–ocean–wave interactions that influence TC/ETC-induced risks under a changing climate. Additionally, we discuss microscale modeling capabilities designed to bridge scale gaps from the weather scale (a few kilometers) to the turbine scale (dozens to a few meters). We also review machine learning (ML)-based, data-driven models for simulating TC/ETC events at both weather and wind turbine scales. Special attention is given to extreme metocean conditions like extreme wind gusts, rapid wind direction changes, and high waves, which pose threats to offshore wind energy infrastructure. Finally, the paper outlines the research challenges and future directions needed to enhance the resilience and design of next-generation offshore wind turbines against extreme weather conditions.

17 WIND ENERGY↗

Graph Identification of Proteins in Tomograms (GRIP-Tomo) 2.0: Topologically aware classification for proteins

Cryo-electron tomography (cryo-ET) enables structural characterization of biomolecules under near-native conditions. Existing approaches for interpreting the resulting three-dimensional volumes are computationally expensive and have difficulty interpreting density associated with small proteins/complexes. To explore alternate approaches for identifying proteins in cryo-ET data we pursued a Graph Network and topologically invariant approach. Here, we report on a fast algorithm that classifies particles by searching for nuances of evolutionarily conversed motifs and the geometrical characteristics of protein structure. GRIP-Tomo 2.0 is a machine-learning pipeline that extracts interpretable topological features of protein structures within noisy experimental backgrounds. Compared to version 1.0, the new pipeline includes three upgrades that significantly improve performance including synthetic tomogram generation simulating realistic noise, graph-based persistent feature extraction as protein fingerprints, and high-performance computing acceleration. GRIP-Tomo 2.0 achieves over 90% accuracy in classifying between proteins and noise using both real and synthetic datasets which represents a foundational step toward advancing cryo-ET workflows and empowering automated visual proteomics.

Li, Chengxuan↗

Method to simultaneously facilitate all jet physics tasks

Machine learning has become an essential tool in jet physics. Due to their complex, high-dimensional nature, jets can be explored holistically by neural networks in ways that are not possible manually. However, innovations in all areas of jet physics are proceeding in parallel. We show that specially constructed machine learning models trained for a specific jet classification task can improve the accuracy, precision, or speed of all other jet physics tasks. This is demonstrated by training on a particular multiclass generation and classification task and then using the learned representation for different generation and classification tasks, for datasets with a different (full) detector simulation, for jets from a different collision system ($pp$ versus $ep$), for generative models, for likelihood ratio estimation, and for anomaly detection. We consider our omnilearn approach thus as a jet-physics foundation model. It is made publicly available for use in any area where state-of-the-art precision is required for analyses involving jets and their substructure.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A User-Friendly GUI Tool for Automated Microstructural Analysis of Fiber-Reinforced Composites and Porous Structures

Understanding and quantifying microstructural features such as fiber orientation and porosity is critical for predicting the mechanical behavior and performance of fiber-reinforced polymer composites. Traditional manual analysis is time-consuming, subjective, and unsuitable for high-throughput datasets. We present a graphical user interface (GUI) application that automates the analysis of microscopy images to extract key microstructural metrics, including fiber orientation tensors, fiber orientation distribution, porosity and pore size distribution. The app integrates multiple image segmentation techniques including global and local thresholding, clustering, and region-based approaches, offering flexibility for different types of image qualities and features. Users can load microstructural images, select regions of interest and segmentation techniques tailored to their image dataset. It also addresses a critical challenge in fiber orientation analysis: the ambiguities caused by touching, overlapping, or partially cut fibers. It supports autorun examples for standardized workflows, enabling reproducible analysis and facilitating training and benchmarking. This tool significantly reduces manual intervention, enhances consistency, and accelerates data generation for structure–property modeling, process optimization, and digital materials research. The tool is intended for use by materials scientists, engineers, and researchers engaged in composite characterization, quality control, and machine learning-based microstructural studies.

Chawla, Komal [ORNL] (ORCID:0000000190327565)↗

Storage and retrieval of mass spectral information

Computer handling of mass spectra serves two main purposes: the interpretation of the occasional, problematic mass spectrum, and the identification of the large number of spectra generated in the gas-chromatographic-mass spectrometric (GC-MS) analysis of complex natural and synthetic mixtures. Methods available fall into the three categories of library search, artificial intelligence, and learning machine. Optional procedures for coding, abbreviating and filtering a library of spectra minimize time and storage requirements. Newer techniques make increasing use of probability and information theory in accessing files of mass spectral information.

Hohn, M. E.↗

Decoding the effects of synonymous variants

Synonymous single nucleotide variants (sSNVs) are common in the human genome but are often overlooked. However, sSNVs can have significant biological impact and may lead to disease. Existing computational methods for evaluating the effect of sSNVs suffer from the lack of gold-standard training/evaluation data and exhibit over-reliance on sequence conservation signals. We developed synVep (synonymous Variant effect predictor), a machine learning-based method that overcomes both of these limitations. Our training data was a combination of variants reported by gnomAD (observed) and those unreported, but possible in the human genome (generated). We used positive-unlabeled learning to purify the generated variant set of any likely unobservable variants. We then trained two sequential extreme gradient boosting models to identify subsets of the remaining variants putatively enriched and depleted in effect. Our method attained 90% precision/recall on a previously unseen set of variants. Furthermore, although synVep does not explicitly use conservation, its scores correlated with evolutionary distances between orthologs in cross-species variation analysis. synVep was also able to differentiate pathogenic vs. benign variants, as well as splice-site disrupting variants (SDV) vs. non-SDVs. Thus, synVep provides an important improvement in annotation of sSNVs, allowing users to focus on variants that most likely harbor effects.

Zishuo Zeng↗

Probabilistic Calibration of Expensive Models using Efficiently Trained Surrogates

Calibration of computational models in the presence of uncertainty is often cast as a Bayesian inference problem and solved via sampling methods, e.g., Markov chain Monte Carlo. When the computational model is expensive, this task becomes intractable due to the large number of samples required to accurately estimate the posterior distribution of the calibration parameters. A popular solution to this problem is to use machine learning to develop a faster-to-evaluate, lower-fidelity substitute for the original model to serve as a surrogate while solving the inference problem. Although considered an offline cost, generating training data to construct this surrogate model can still be an expensive task in practice. An active learning algorithm is presented that focuses training on improving surrogate accuracy specifically in and around the bulk of the posterior distribution, as this is where the model is exercised during calibration. Candidate samples are drawn from families of distributions related to an approximation of the posterior. The sample maximizing predictive variance is then selected for evaluation by the original computational model, yielding a label for the training point. Iterating this approach increases efficiency relative to space filling designs (e.g., Latin hypercube sampling) by avoiding low probability points. Practical considerations are discussed, including the benefits of using a sequential Monte Carlo sampling approach, convergence heuristics, and the importance of both exploration and exploitation given that the true posterior is unknown a priori.

uncertainty quantification↗

SUBTASK 1.6 – BASIN ELECTRIC CARBON STORAGE RESEARCH PROJECT: NOVEL MONITORING TECHNIQUES

The Energy & Environmental Research Center (EERC) conducted baseline activities associated with an applied research project at Basin Electric Power Cooperative’s (Basin’s) carbon capture and storage (CCS) site in Beulah, North Dakota, to establish novel carbon storage-monitoring techniques as commercial methods under Cooperative Agreement No. DE-FE0024233, Subtask 1.6. The following report summarizes the baseline activities performed and briefly describes the subsequent (operational monitoring) activities that have been proposed to the U.S. Department of Energy (DOE) as part of the overall project to develop and demonstrate novel monitoring techniques at North America’s largest permitted CCS operation. Dakota Gasification Company (DGC), a wholly owned subsidiary of Basin, owns and operates the Great Plains Synfuels Plant (GPSP) approximately 5 miles northwest of the town of Beulah, North Dakota (Figure 1). In 2023, DGC received approval from the North Dakota Industrial Commission (NDIC) to develop a storage facility on-site for injecting a stream of carbon dioxide (CO2) captured from GPSP. DGC will transport the captured CO2 stream with approximately 6.8 miles of transmission lines that extend north of GPSP and inject >1 million tonnes (MMt) of CO2 annually (>1 MMt/yr) over a 12-year period with up to six underground injection control (UIC) Class VI-compliant injection wells completed in the Broom Creek Formation, a predominantly sandstone reservoir and saline aquifer underlying GPSP. The Broom Creek Formation lies approximately 5900 feet (ft) below ground surface (bgs) at GPSP. The commercial scale (i.e., >1 MMt/yr) of DGC’s permitted carbon storage project is ideal for developing and testing the novel monitoring techniques included within Subtask 1.6. The goals of this project are to demonstrate 1) the cost-effectiveness of novel monitoring technologies included as part of this research, 2) technology capability for tracking the CO2 plume and/or associated pressure response in the subsurface and monitoring out-of-zone migration, and 3) compliance with UIC Class VI program requirements. The research activities proposed for the overall project include 1) design of an automated, integrated, modular (AIM) monitoring station; 2) time-lapse electromagnetic (EM) field surveys; 3) drone-based surveillance studies; 4) time-lapse monitoring with seismic methods; 5) advanced wellbore-monitoring methods; 6) deployment of an AIM monitoring network; 7) EM monitoring of CO2 with real-time data processing; 8) continued seasonal drone-based surveillance studies; 9) seismic monitoring with passive and active surveys; and 10) wellbore monitoring with nuclear magnetic resonance (NMR) for near-surface characterization. Completion of Activities 1.0–5.0 (baseline activities) are described in this report. Upon authorization of funding by DOE, the EERC will initiate Activities 6.0– 10.0 (operational monitoring activities). Current state-of-the-art (SOA) carbon storage-monitoring techniques require countless labor hours dedicated to the acquisition of data. Once data are gathered, these SOA techniques often rely on commercial facilities to process raw data from the field. However, it is anticipated that next-generation monitoring techniques, such as those being demonstrated, will lower acquisition footprints, be less operationally intensive, and improve data acquisition efficiencies. These new techniques are more conducive to the application of machine learning, artificial intelligence, and automation, thus providing a pathway for integration into active control systems, informing site operability, and improving the integration of data for future CCS projects across the United States. Additionally, reclaimed and active mining lands are present within the project site, creating a unique opportunity to demonstrate the effectiveness of remote sensing and surface-based geophysics monitoring techniques at similar project sites that may include disturbed, unconsolidated, or actively excavated near-surface environments. The efforts included in the overall project will produce necessary designs, learnings, and data acquired during the baseline and operational monitoring periods that are necessary for time-lapse demonstration and validation of the described monitoring techniques. In addition, it is anticipated that the monitoring technologies included in this study will be compliant with UIC Class VI requirements to enable the potential for implementation at other CCS sites across the United States.

42 ENGINEERING↗