Search NASASearch

SEARCH · Search NASA

Results for “AI/ML training”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Molecular simulation and artificial intelligence for the circular economy of bioenergy and bioproducts

The concept of the circular bioeconomy is a carbon neutral, sustainable system with zero waste. One vision for such an economy is based upon lignocellulosic biomass. This lignocellulosic circular bioeconomy requires CO 2 absorption from biomass growth and the efficient deconstruction of recalcitrant biomass into solubilized and fractionated biopolymers, which are then used as precursors for the sustainable production of high-quality liquid fuels, chemical bioproducts, and bio-based materials. Here, in this study, we summarize the roles that molecular dynamics (MD) simulations and machine learning (ML) are playing in overcoming several fundamental challenges hindering the adoption of a circular bioeconomy. Specifically, we discuss the role of MD and ML/AI in overcoming lignocellulose recalcitrance by designing biomass pretreatment methods to efficiently produce solubilized cellulose/lignin/hemicellulose and of that in improving energy-intensive manufacturing of biomass-based materials and their structural and mechanical properties. Quantum mechanical methods and MD simulations, in addition to offering a mechanistic understanding of biomass deconstruction and biomaterials design, can provide meaningful structural, energetics, and physiochemical properties as inputs to train AI/ML models. The ML models can guide the experimental prioritization of materials/solvents and process parameters that significantly accelerate the development of biofuel and biomaterial components of the circular bioeconomy.

Smith, Jeremy C. [Oak Ridge National Laboratory (O

Data From Experiments on Bubbling Fluidization of Zeolite in a Rectangular Bubbling Fluidized Bed

Fluidization experiments were conducted in a lab-scale rectangular bubbling fluidized bed with the objective of generating a high-quality dataset for model validation and artificial intelligence/machine learning (AI/ML) training. Zeolite was chosen as the bed material, and the fluidizing medium was air as supplied by a compressor. Three different flow rates at the inlet were chosen such that the particles were fluidized but not elutriated from the system. The test matrix involved randomization and replicates to provide uncertainty estimates as well as four different batches of zeolite as the bed material. The quantities of interest obtained from this study were statistics of differential pressures, interface heights, and particle velocities. Considering all the components of the elaborate test plan, the results obtained were consistent and reproducible. Characterization tests were performed to estimate particle properties including size, density, coefficient of friction, coefficient of restitution, and minimum fluidization velocity. In addition, the angle of repose from granular discharge experiments has been reported to account for rolling friction, though its effect on the overall process is expected to be negligible.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

From sequence to protein structure and conformational dynamics with artificial intelligence/machine learning

The 2024 Nobel Prize in Chemistry was awarded in part for de novo protein structure prediction using AlphaFold2, an artificial intelligence/machine learning (AI/ML) model trained on vast amounts of sequence and three-dimensional structure data. AlphaFold2 and related models, including RoseTTAFold and ESMFold, employ specialized neural network architectures driven by attention mechanisms to infer relationships between sequence and structure. At a fundamental level, these AI/ML models operate on the long-standing hypothesis that the structure of a protein is determined by its amino acid sequence. More recently, AlphaFold2 has been adapted for the prediction of multiple protein conformations by subsampling multiple sequence alignments. Herein, we provide an overview of the deterministic relationship between sequence and structure, which was hypothesized over half a century ago with profound implications for the biological sciences ever since. We postulate that protein conformational dynamics are also determined, at least in part, by amino acid sequence and that this relationship may be leveraged for construction of AI/ML models dedicated to predicting protein conformational ensembles. Accordingly, we describe a conceptual model architecture, which may be trained on sequence data in combination with conformationally sensitive structural information, coming primarily from nuclear magnetic resonance (NMR) spectroscopy. Notwithstanding certain limitations in this context, NMR offers abundant structural heterogeneity conducive to conformational ensemble prediction. As NMR and other data continue to accumulate, sequence-informed prediction of protein structural dynamics with AI/ML has the potential to emerge as a transformative capability across the biological sciences.

Artificial intelligence

Cloverleaf Data Artifacts for ArtIMis LDRD

This report summarizes the use of the open-source CloverLeaf/CloverLeaf3D mini-apps to generate synthetic data sets to train foundation models for the ArtIMis LDRD DI. These data artifacts are intended to be used by LANL collaborators and shared externally with our university and institutional partners. Note that CloverLeaf/CloverLeaf3D is not a LANL simulation code.

97 MATHEMATICS AND COMPUTING

The Aerosol Model Benchmarking Repository: A toolkit for model intercomparison

The Aerosol Model Benchmarking Repository and Standards (AMBRS) project was initiated to provide tools and to establish community standards for benchmarking aerosol models. This report describes a set of open-source tools for building, running, and analyzing aerosol box model simulations in a standardized framework. The framework consists of three core components: AMBuilder, a CMake-based build system that compiles supported models consistently; AMBRS, a Python module that defines unified numerical experiments and executes them with aligned inputs; and PyParticle, an aerosol analysis package that standardizes output, computes diagnostics, and visualizes simulation results. Together, these tools enable reproducible intercomparison of aerosol schemes and support process-level evaluation of how model simplifications affect predictions of size distributions, cloud condensation nuclei activity, and other relevant properties relevant for the Earth-Energy system. Beyond its role in benchmarking, AMBRS provides a platform for studying aerosol processes across scales and can be used to generate training data for AI/ML applications in support of a broader hierarchical aerosol modeling strategy.

54 ENVIRONMENTAL SCIENCES

Smart culture medium optimization for recombinant protein production: Experimental, modeling, and AI/ML-driven strategies

Recombinant protein production (RPP) is central to biotechnology, where recombinant proteins are used as either end products or catalysts in the synthesis of chemicals, fuels, and materials. Among the major cost drivers, culture medium plays a pivotal role in determining protein yield and quality. This review presents a comprehensive perspective on the critical stages of “smart” culture medium optimization: planning, screening, modeling, optimization, and validation. In the planning stage, we examine the nutritional and energetic roles of medium components, including carbon, nitrogen, amino acids, salts, and trace metals, and their impacts on culture parameters such as pH, oxidative state, and osmolality. We highlight the variability in trace metal content due to water sources, culture vessels, and raw materials, which can substantially influence RPP. The screening stage covers Design of Experiments (DoE) approaches, assessing their theoretical basis, implementation, and limitations. For modeling, we describe methods that integrate experimental data to develop predictive models for smart medium formulation. Model-based optimization strategies can then be employed to select optimal media compositions for a given application. The validation stage aims to evaluate model predictions and provide feedback for model training and refinement. Finally, we survey mechanistic and artificial intelligence/machine learning (AI/ML)-driven models as integrated, transformational tools for predictive modeling of bioprocess conditions, nutrient availability, cellular metabolism, and protein quality, with the goal of optimizing culture media to enhance protein yields while reducing costs and environmental impact. We conclude by addressing the challenges of translating laboratory-scale medium optimization to industrial-scale settings and exploring future AI/ML-driven approaches that may overcome current bottlenecks and accelerate medium design for RPP. Overall, this review provides a unified framework for advancing smart medium design in RPP.

Artificial Intelligence/Machine Learning (AI/ML)

Anomaly Detection In DUNE Using AI/ML

We designed and built an AI/ML model to detect anomalies in the DUNE far detector data. The model has been trained on simulated radiological background (rbkg) data, which is the major background for supernova burst neutrinos. The trained model was evaluated on both new samples of radiological backgrounds and supernova burst neutrino events in the elastic scattering and charged current interaction channels. We found that the trained model can successfully identify supernova burst neutrino events as anomalies while identifying radiological backgrounds as nominal events.

Novello, Eric [Unlisted, US]

Anomaly Detection in DUNE FD LArTPC Readouts

We designed and built an AI/ML model to detect anomalies in the DUNE far detector data. The model has been trained on simulated radiological background (rbkg) data, which is the major background for supernova burst neutrinos that we want to detect as anomaly in thios work. The trained model was evaluated on both new samples of radiological backgrounds and supernova burst neutrino events in the elastic scattering and charged current interaction channels. We found that the trained model can successfully identify supernova burst neutrino events as anomalies while identifying radiological backgrounds as nominal events.

Novello, Eric [Fermilab]

Virtual Reality for Shoot/No-Shoot Decision Training in Law Enforcement: A Literature Review and Research Agenda

Virtual reality (VR) can materially improve “shoot / no-shoot” (SNS) training by giving officers realistic, repeatable practice making high-stakes decisions under pressure. Traditional tools—live-fire ranges and video simulators—build basics, but they cannot adapt to each officer in real time or fully mirror the complexity of the field. VR closes that gap by creating immersive scenarios that are safer, more flexible, easier to scale across units, and able to capture objective performance data. SNS decisions are not just about marksmanship; they rely on perception, judgment, memory, and the ability to hold fire when a threat is uncertain. Effective training therefore needs realism, decision complexity, and branching outcomes that reflect the true consequences of choices. These elements strengthen recognition of hostile intent while reducing false positives and building the self-control required in ambiguous situations. VR brings specific advantages: dynamic environments, full-body interaction, and the ability to measure performance with precision—enabling targeted feedback and better transfer of learning to the street. At the same time, responsible deployment must address scenario quality (credible environments and behaviors), lawful decision models, and user wellbeing (appropriate stress levels, comfort, and safety). Sandia’s VIPER Lab is positioned to lead this work. The team combines human-performance science, AI/ML, and VR/AR development with a deep equipment bench (e.g., omnidirectional treadmill, eye-tracking, haptics, multiple HMDs). This ecosystem supports building and validating next-generation SNS training that is immersive, measurable, and trustworthy. Bottom line: Investment in VR-enabled SNS training that blends evidence-based design with careful validation and legal safeguards is expected to pay off in safer, more consistent decision-making and improved community trust, delivered through training that is practical to deploy at scale.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF

Predicting U 3 O 8 powder processing conditions: An AI/ML approach analyzing deep learning embeddings of SEM micrographs

High-resolution SEM images of uranium-oxide powders encode micro- and nanoscale clues to their synthesis route and calcination temperature. We trained a ResNet-50 model on 11 commercial-scale U₃O₈ classes, ammonium diuranate (ADU) or uranyl peroxide (H₂O₂) precursors calcined at temperatures ranging from 400 to 750 °C and added a 256-D projection head before the classifier to analyze the learned representation. The best of eight seeds reached 92.4 % accuracy on reserved testing data, but our focus is the structure of the embedding space rather than the accuracy and labels. We quantify class relatedness in the original 256-D space using centroid similarity and distributional distances, and we use Uniform Manifold Approximation Projection (UMAP) for visualization. ‘Unknown’ images from different preparation methods, SEM operators, and from the literature localized near the expected classes under a nearest-centroid analysis without retraining, as well as clustered in similar UMAP space. In conclusion, this embedding-centered workflow complements black-box classification by providing quantitative, similarity-based comparisons of U₃O₈ morphologies and reduces storage space by up to 98 % for image data used in millisecond vector search comparisons.

36 MATERIALS SCIENCE

2019 Budget Request for the DOE Computational Science Graduate Fellowship (CSGF) Grant

The Department of Energy Computational Science Graduate Fellowship (DOE CSGF) is necessary to meet the continual challenging national workforce needs that arise as computational science and engineering problems continue to grow in scope and complexity. Computational science and engineering (CSE) is a multidisciplinary approach that uses scientific computing to solve practical problems methods and to supply technical tools across the scientific discovery spectrum. In particular, the DOE CSGF emphasizes high-performance computing (HPC) that enables CSE that advances science and engineering in directions important to the DOE and the economy in general. Over the past half-century, HPC has been an essential tool for DOE’s success. During this period, important missions, such as nuclear stockpile stewardship, have turned to HPC as an essential technology. Entire science disciplines, such as biology and cosmology, have been transformed through the augmentation of scientific observation via HPC. At government laboratories and in industry, DOE CSGF alumni are helping push traditional HPC boundaries while contributing to discoveries in high-energy physics, renewable energy, fusion-reactor design, additive manufacturing, nanomaterials for next-generation batteries and transistors, and turbine and advanced nuclear reactor modeling. In addition, HPC is used to address national health needs that will eventually point to cures both by helping cancer researchers manage and analyze huge troves of data, by simulating biological mechanisms, and by accelerating drug development — including continuing to rise to the challenge of pandemic-related research. A 2023 report from the ASCAC Subcommittee on American Competitiveness and Innovation to the ASCR office, “Can the United States Maintain Its Leadership in High-Performance Computing?” says of the Program, “The CSGF program provides a barometer for disciplines that will be of interest to future DOE computing.” An explosion in scientific and technological data has driven the need for increasingly sophisticated HPC to transform those data into scientific understanding. With access to more and more data and the proliferation of HPC, Machine Learning and Artificial Intelligence are experiencing a renaissance, complementing the now well-established use of computational simulation. Indeed, in its September 2020 subcommittee report on “AI/ML, Data Intensive Science and High-Performance Computing”, the DOE Advanced Scientific Computing Advisory Committee (ASCAC) explicitly called for a fellowship program to train computational and data scientists to tackle exascale and data-intensive computing challenges. This collaboration of empirical and theory-based modeling will increasingly inform federal policymakers whose decisions affect American society and future generations, and it requires highly skilled and intellectually agile computational scientists who can support the fast-moving DOE National Laboratory research environment. In fact, the DOE CSGF program has explicitly and consistently addressed this need.

97 MATHEMATICS AND COMPUTING

2020 Budget Request for the DOE Computational Science Graduate Fellowship (CSGF) Grant

The Department of Energy Computational Science Graduate Fellowship (DOE CSGF) is essential for addressing the increasingly complex national workforce demands stemming from the growth of computational science and engineering challenges. Computational science and engineering (CSE) takes a multidisciplinary approach that utilizes scientific computing to tackle practical problems and provide technical tools across the spectrum of scientific discovery. The DOE CSGF specifically highlights high-performance computing (HPC) as a critical enabling technology in CSE, driving advancements in science and engineering that are vital to both the DOE and the broader economy. Over the past half-century, HPC has been an essential tool for DOE’s success. During this period, important missions, such as nuclear stockpile stewardship, have turned to HPC as an essential technology. Entire science disciplines have been transformed through the augmentation of scientific observation via HPC. At government laboratories, academic institutions, and in industry, DOE CSGF alumni are helping push traditional HPC boundaries while contributing to discoveries in high-energy physics, quantum information systems, fusion-reactor design, machine learning, additive manufacturing, nano materials for next-generation batteries and transistors, and advanced nuclear reactor modeling. In addition, HPC is used to address national health needs that will eventually point to cures both by helping cancer researchers manage and analyze huge troves of data, by simulating biological mechanisms, and by accelerating drug development. A 2023 report from the ASCAC Subcommittee on American Competitiveness and Innovation to the ASCR office, “Can the United States Maintain Its Leadership in High-Performance Computing?” says of the Program, “The CSGF program provides a barometer for disciplines that will be of interest to future DOE computing. Computational biology, machine learning, and quantum computing are among the subjects that began to swell in the ranks of CSGF applicants before the labs were hiring as high a percentage of employees in these categories.” The explosion of scientific and technological data has heightened the demand for advanced high-performance computing (HPC) to transform these data into meaningful scientific insights. As access to vast amounts of data increases, the fields of Machine Learning and Artificial Intelligence are experiencing a resurgence, enhancing the established practices of computational modeling and simulation. In its September 2020 subcommittee report on "AI/ML, Data Intensive Science, and High-Performance Computing," the DOE Advanced Scientific Computing Advisory Committee (ASCAC) specifically called for a fellowship program to train computational and data scientists to address exascale and data-intensive computing challenges. This integration of empirical and theoretical modeling will increasingly guide federal policymakers in making decisions that impact American society and future generations. It demands a workforce of highly skilled and intellectually agile computational scientists capable of navigating the rapid advancements in scientific computing within the DOE National Laboratory research environment. The DOE CSGF program has consistently addressed this critical need.

97 MATHEMATICS AND COMPUTING

TCR-H: explainable machine learning prediction of T-cell receptor epitope binding on unseen datasets

Artificial-intelligence and machine-learning (AI/ML) approaches to predicting T-cell receptor (TCR)-epitope specificity achieve high performance metrics on test datasets which include sequences that are also part of the training set but fail to generalize to test sets consisting of epitopes and TCRs that are absent from the training set, i.e., are ‘unseen’ during training of the ML model. We present TCR-H, a supervised classification Support Vector Machines model using physicochemical features trained on the largest dataset available to date using only experimentally validated non-binders as negative datapoints. TCR-H exhibits an area under the curve of the receiver-operator characteristic (AUC of ROC) of 0.87 for epitope ‘hard splitting’ (i.e., on test sets with all epitopes unseen during ML training), 0.92 for TCR hard splitting and 0.89 for ‘strict splitting’ in which neither the epitopes nor the TCRs in the test set are seen in the training data. Furthermore, we employ the SHAP (Shapley additive explanations) eXplainable AI (XAI) method for post hoc interrogation to interpret the models trained with different hard splits, shedding light on the key physiochemical features driving model predictions. TCR-H thus represents a significant step towards general applicability and explainability of epitope:TCR specificity prediction.

60 APPLIED LIFE SCIENCES

LLMs for Mfg.—On the State of Large Language Models and Applications to Manufacturing

Additive Manufacturing (AM), referred to as 3D printing, has emerged as a key pillar of Industry 4.0 enabling layer-by-layer fabrication of intricate geometries from CAD models. In parallel, Large Language Models (LLMs), deep learning models for natural language generation trained on vast text corpora, have demonstrated unprecedented capabilities in understanding and generating human-like text. The convergence of these trends opens new opportunities at the intersection of AM and AI/ML, where LLMs can assist engineers and researchers in design, manufacture planning, and knowledge discovery. Recent academic work has begun to explore LLM applications in AM and adjacent fields, such as material science, mechanical engineering, and design for additive manufacturing. This exploration ranges from intelligent process planning to domain-specific knowledge retrieval. This survey provides a comprehensive review of current developments, focusing on peer-reviewed literature contributions that apply, adapt, and advance LLMs in general and domain-specific domains. We analyze state-of-the-art (SOTA) techniques, such as fine-tuning foundational models for specific domains, retrieval-augmented generation (RAG) pipelines, knowledge graph integration, and delve into the architectures and evaluation methods employed. The goal of this survey is to inform researchers and practitioners of the current capabilities and limitations of LLMs in general and in domain-specific applications, and to outline how these models are being tailored to meet the requirements of these applications.

36 MATERIALS SCIENCE

Advancing AI-Driven Analysis in X-ray Absorption Spectroscopy: Spectral Domain Mapping and Universal Models

In recent years, rapid progress has been made in developing artificial intelligence (AI) and machine learning (ML) methods for X-ray absorption spectroscopy (XAS) analysis. Compared to traditional XAS analysis methods, AI/ML approaches offer dramatic improvements in efficiency and help eliminate human bias. To advance this field, we advocate an AI-driven XAS analysis pipeline that features several interconnected key building blocks: benchmarks, workflows, databases, and AI/ML models. Specifically, we present two case studies for XAS ML. In the first study, we demonstrate the importance of reconciling the discrepancies between simulation and experiment using spectral domain mapping (SDM). Our ML model, which is trained solely on simulated spectra, predicts an incorrect oxidation state trend for Ti atoms in a combinatorial zinc titanate film. After transforming the experimental spectra into a simulation-like representation using SDM, the same model successfully recovers the correct oxidation state trend. In the second study, we explore the development of universal XAS ML models that are trained on the entire periodic table, which enables them to leverage common trends across elements. Looking ahead, we envision that an AI-driven pipeline can unlock the potential of real-time XAS analysis to accelerate scientific discovery.

36 MATERIALS SCIENCE

Generalized parton distributions from symbolic regression

AI/ML informed Symbolic Regression is the next stage of scientific modeling. We utilize a highly customizable symbolic regression package “PySR” to model the x and t dependence of the flavor isovector combination Hu-d(x, t, ) at = 0. These PySR models were trained on GPD results provided by both Lattice QCD and phenomenological sources GGL, GK, and VGG. We demonstrate, for the first time, the consistency and systematic convergence of Symbolic Regression by quantifying the disparate models through their Taylor expansion coefficients. In addition to PySR penalizing models with higher complexity and mean-squared error, we implement schemes that test specific physics hypotheses, including force-factorized x and t dependence and Regge behavior in PySR GPDs. We show that PySR can identify factorizing GPD sources based on their response to the Force-Factorized model. Here, knowing the precise behavior of the GPDs, and their uncertainties in a wide range in x and t, crucially impacts our ability to concretely and quantitatively predict hadronic spatial distributions and their derived quantities.

Dotson, Andrew [New Mexico State University, Las C

A Fast, Accurate Prediction for System-Wide Damage Due to Dynamic Wind Loading

The complex relationship between photovoltaic (PV) hardware configurations, overall system dynamics, and turbulent aerodynamic phenomena generates highly unsteady, non-uniform loads that can lead to damaging instabilities. These effects may result in glass breakage, cell cracking, and structural failures in frames and mounting systems, even under moderate wind conditions. Addressing industry concerns about premature system failures in field conditions deemed survivable, our research aims to develop a fast and accurate predictive model for system damage. This model integrates configurable hardware choices with advanced simulation tools to represent the overall system-specific dynamics effectively. Using this model, we predict responses under varying weather conditions and hardware setups, translating these predictions into pre-trained surrogate models capable of accurately identifying failure risks and rapidly testing new system hardening measures. In this presentation, we will showcase preliminary results in capturing system dynamics through our customizable library of PV hardware configurations. Additionally, we will highlight how these new tools build upon PVade's established wind load modeling capabilities and foster the development of advanced AI/ML surrogates for improving system robustness.

97 MATHEMATICS AND COMPUTING

FAIR Surrogate Benchmarks Supporting AI and Simulation Research (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia (UVA). SBI repositories include data, code, and all relevant collateral artifacts that the science and engineering community need to use and reuse these data sets and surrogates. SBI repositories generate active research from both the participants in SBI and the broad community of AI and domain scientists. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and captures them as surrogate benchmarks with a rich set of metadata covering: Data; Model; Metrics specification; Machine specification; and Science, Speed, and Power Results. We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non-Surrogate benchmarks that have many common features and similar issues as regards FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, Benchmarks have datasets, models, and metadata and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING