Search NASA⌕ Search

SEARCH · Search NASA

Results for “generative machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES↗

Developing Scenario‐Based Strategies for Health, Climate, and Environmental Preparedness: The One Health, One Earth Approach

Climate change amplifies many threats to human health. Despite advances in understanding climate change dynamics and impacts, there remains a critical gap in translating scientific knowledge into equitable, and community-driven health interventions. The inaugural One Earth, One Health workshop sought to explore this gap through human-centered design exercises involving interdisciplinary researchers from climate and Earth sciences, engineering, epidemiology, microbiology, and environmental health. Although participants did not co-develop solutions with affected communities, they used stakeholder role-playing to guide ideation and lay groundwork for actionable plans. Through these methods, participants identified community needs and proposed prototype solutions to alleviate health threats exacerbated by global environmental change. Prototypes were organized around infectious diseases, extreme weather, and air quality, as illustrative themes rather than an exhaustive set of risks. Key solutions included strategies for anticipatory systems and early warning (e.g., integrating environmental signals with health data), inclusive communication and infrastructure needs for responding to extreme weather events, and integrated platforms visualizing air quality trends to support tailored, context-aware guidance beyond one-size-fits-all alerts. The workshop highlighted opportunities such as leveraging machine learning, Earth observation, and real-time surveillance to protect communities, but also noted barriers including data quality, technological redundancy, privacy, and governance challenges. Additionally, participants emphasized the need for interdisciplinary teams capable of collaborating across sectors, breaking down silos and addressing gaps in training and education. Overall, the workshop illustrates how process-driven, human-centered approaches can help surface user needs and generate testable prototype concepts, while underscoring the importance of direct community partnership for implementation.

Abadi, Azar M. [University of Alabama, Birmingham,↗

Particle hit clustering and identification using point set transformers in liquid argon time projection chambers

Liquid argon time projection chambers are often used in neutrino physics and dark-matter searches because of their high spatial resolution. The images generated by these detectors are extremely sparse, as the energy values detected by most of the detector are equal to 0, meaning that despite their high resolution, most of the detector is unused in a particular interaction. Instead of representing all of the empty detections, the interaction is usually stored as a sparse matrix, a list of detection locations paired with their energy values. Traditional machine learning methods that have been applied to particle reconstruction such as convolutional neural networks (CNNs), however, cannot operate over data stored in this way and therefore must have the matrix fully instantiated as a dense matrix. Operating on dense matrices requires a lot of memory and computation time, in contrast to directly operating on the sparse matrix. We propose a machine learning model using a point set neural network that operates over a sparse matrix, greatly improving both processing speed and accuracy over methods that instantiate the dense matrix, as well as over other methods that operate over sparse matrices. Compared to competing state-of-the-art methods, our method improves classification performance by 14%, segmentation performance by more than 22%, while taking 80% less time and using 66% less memory. Compared to state-of-the-art CNN methods, our method improves classification performance by more than 86%, segmentation performance by more than 71%, while reducing runtime by 91% and reducing memory usage by 61%.

calibration and fitting methods↗

Electric Grid Security (EGS) FY24 Annual Report

Sandia’s Electric Grid Security program advances a national vision of a secure, resilient, and affordable electric system for all users. Our achievements reflect a strategic approach combining technology development; modeling, simulation, and data analytics; and partnered demonstrations and outreach to further the adoption of advanced grid and storage technologies. Our FY24 efforts leverage the strengths of our partnerships—spanning Sandia’s core science and technology competencies as well as external technology leaders—to develop the solutions today which enable the grid of tomorrow. Key accomplishments in this report that support our strategy span our technical program areas and include: • The advancement of energy storage technologies, including creation of a national Long Duration Energy Storage Consortium; • Applications of artificial intelligence and machine learning to enhanced grid operations and planning; • Development of solid-state power conversion technologies and a new medium-voltage research lab; • New technologies to assess wildfire vulnerabilities and mitigate potential impacts; • Advanced applications of new cybersecurity technologies with industry partners; • Contributions to understanding the impacts of electromagnetic pulses and geomagnetic disturbances on grid components; and • Digital twin development for hybrid microgrids with multiple generators, storage, and loads.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Improving Text Classification with Large Language Model-Based Data Augmentation

Large Language Models (LLMs) such as ChatGPT possess advanced capabilities in understanding and generating text. These capabilities enable ChatGPT to create text based on specific instructions, which can serve as augmented data for text classification tasks. Previous studies have approached data augmentation (DA) by either rewriting the existing dataset with ChatGPT or generating entirely new data from scratch. However, it is unclear which method is better without comparing their effectiveness. This study investigates the application of both methods to two datasets: a general-topic dataset (Reuters news data) and a domain-specific dataset (Mitigation dataset). Our findings indicate that: 1. ChatGPT generated new data consistently enhanced model’s classification results for both datasets. 2. Generating new data generally outperforms rewriting existing data, though crafting the prompts carefully is crucial to extract the most valuable information from ChatGPT, particularly for domain-specific data. 3. The augmentation data size affects the effectiveness of DA; however, we observed a plateau after incorporating 10 samples. 4. Combining the rewritten sample with new generated sample can potentially further improve the model’s performance.

97 MATHEMATICS AND COMPUTING↗

PDF Entity Annotation Tool (PEAT)

While different text mining approaches – including the use of Artificial Intelligence (AI) and other machine based methods - continue to expand at a rapid pace, the tools used by researchers to create the labeled datasets required for training, modeling, and evaluation remain rudimentary. Labeled datasets contain the target attributes the machine is going to learn; for example, training an algorithm to delineate between images of a car or truck would generally require a set of images with a quantitative description of the underlying features of each vehicle type. Development of labeled textual data that can be used to build natural language machine learning models for scientific literature is not currently integrated into existing manual workflows used by domain experts. Published literature is rich with important information, such as different types of embedded text, plots, and tables that can all be used as inputs to train ML/natural language processing (NLP) models, when extracted and prepared in machine readable formats. Currently, both normalized data extraction of use to domain experts and extraction to support development of ML/NLP models are labor intensive and cumbersome manual processes. Automatic extraction of data and information from formats such as PDFs that are optimized for layout and human readability, not machine readability. The PDF (Portable Document Format) Entity Annotation Tool (PEAT) was developed with the goal of allowing users to annotate publications within their current print format, while also allowing those annotations to be captured in a machine-readable format. One of the main issues with traditional annotation tools is that they require transforming the PDF into plain text to facilitate the annotation process. While doing so lessens the technical challenges of annotating data, the user loses all structure and provenance that was inherent in the underlying PDF. Also, textual data extraction from PDFs can be an error prone process. Challenges include identifying sequential blocks of text and a multitude of document formats (multiple columns, font encodings, etc.). As a result of these challenges, using existing tools for development of NLP/ML models directly from PDFs is difficult because the generated outputs are not interoperable. We created a system that allows annotations to be completed on the original PDF document structure, with no plain text extraction. The result is an application that allows for easier and more accurate annotations. In addition, by including a feature that grants the user the ability to easily create a schema, we have developed a system that can be used to annotate text for different domain-centric schemas of relevance to subject matter experts. Different knowledge domains require distinct schemas and annotation tags to support machine learning.

97 MATHEMATICS AND COMPUTING↗

Predicting Atomistic Transitions with Transformers

Accurate knowledge of the atomistic transition pathways in materials and material surfaces is crucial for many material science problems. However, conventional simulation techniques used to find these transitions are extremely computationally intensive. Even with large-scale, accelerated material simulations, the computational cost constrains the applicable domain in practice. Machine learning models, with the potential to learn the complex emergent behaviors governing atomistic transitions as a fast surrogate model, have great promise to predict transitions with a vastly reduced computational cost. Here, we demonstrate how transformers can be trained to predict atomistic transitions in nano-clusters. We show how we evaluate physical validity of the predictions and how a multitude of additional, different microstates can be generated by slightly varying the data provided to the model.

36 MATERIALS SCIENCE↗

Design and optimization of a modular hydrogen-based integrated energy system to maximize revenue via nuclear-renewable sources

Here, this paper demonstrates a novel modular distributed framework that uses optimal energy-dispatching strategies to enable greater flexibility and profitability in nuclear-renewable integrated energy systems (NR-IES). Hydrogen is used as a commodity in this framework since its production can improve grid stability and system operational flexibility, decarbonize heavy industry, and create an additional revenue stream for electricity generators, particularly nuclear power plants with high operational expenses. The proposed solution addresses the challenges associated with merging multiple software and services from various domains by using functional mock-up units (FMU) to co-simulate diverse subsystems designed in various platforms. The tightly coupled integrated energy system (IES) is optimized to maximize revenue by utilizing the deep reinforcement learning (DRL) technique to make smart dispatching decisions based on variable electricity prices and the availability of renewable energy. Proximal policy optimization (PPO) algorithm is used in training and testing the DRL agent. Over a period of 120 days, the proposed hydrogen-based IES framework showed about 10% revenue boost compared to a non-hydrogen generating baseline IES while also providing an easily-adoptable framework which can help to improve the flexibility of future generation nuclear power plants.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Monitoring Fracture Hydromechanical Evolution in the Lab and Field Using Unsupervised Metric Learning

Fractures evolve in time through thermal‐hydraulic‐mechanical‐chemical (THMC) processes that alter their long‐range hydraulic transport properties and modify subsurface behavior and activities. The location of subsurface fractures makes it necessary to use remote sensing techniques such as passive or active seismic monitoring for fracture characterization. In this paper, we develop a machine learning approach to monitor the evolution of fracture properties using passive seismic sources in a laboratory setting and using active seismic monitoring from the Sanford Underground Research Facility in Lead, South Dakota, at a depth of 1.25 km in amphibolite rock during stimulation of natural fractures as well as during induced fracturing. The unsupervised metric learning technique applies tandem neural networks (twin (Siamese) or triplet) with contrastive loss and adaptive margins to track slowly varying systems for which class or similarity labels are not available. The approach adopts locality‐sensitive hashing to divide time‐ordered contiguous data into an arbitrary number of pseudo‐classes. Contrastive‐loss training with many hash bins generates an evolving latent‐space trajectory. This approach enables unsupervised metric learning for seismic data stacks under the condition of contiguous state sampling and slowly varying fracture properties. The displacement discontinuity theory provides a mechanistic foundation for the fracture‐dependent trajectories that are related to relaxation of fractures with time‐dependent specific stiffness responding to changes in stress or fluid saturation.

02 PETROLEUM↗

Thermal property characterization of phase change materials in building applications: A systematic review of fundamentals, recent progress, and future directions

Phase change materials (PCMs) can reduce building peak loads and enable demand-responsive thermal energy storage (TES), but their deployment depends on reliable measurement and interpretation of thermal properties across laboratory, intermediate, and application scales. Here, this review systematically examines characterization methods, testing protocols, and recent advances for neat PCMs and PCM composites, emphasizing thermal conductivity, enthalpy-related properties (phase change temperature, latent heat, specific heat), and cycling stability. For thermal conductivity, we compare steady-state and transient techniques and note limitations when phase transition and contact resistance affect measurements. For enthalpy–temperature characterization, we discuss differential scanning calorimetry together with intermediate- and bulk-scale methods, including T-history, heat flow meter testing, and three-layer calorimetry (3LC), to generate application-relevant enthalpy–temperature profiles. Cycling stability is organized into four experimental families: thermoelectric–air, fully thermoelectric, water-bath, and in situ chamber approaches, with attention to separating reversible supercooling from true degradation such as phase segregation. We highlight emerging noncontact diagnostics, including infrared thermography and embedded sensing, for spatially resolved validation and multiscale interpretation. Finally, we review the growing use of AI and machine learning for property prediction, inverse characterization from experimental signals, and real-time state estimation in building-integrated TES. Key needs include harmonized protocols, interlaboratory benchmarking, uncertainty reporting, and metadata-rich datasets to accelerate reproducible PCM qualification for grid-flexible buildings.

AI↗

Expediting field-effect transistor chemical sensor design with neuromorphic spiking graph neural networks

Improving the sensitive and selective detection of analytes in a variety of applications requires accelerating the rational design of field-effect transistor (FET) chemical sensors. Achieving high-performance detection relies on identifying optimal probe materials that can effectively interact with target analytes, a process traditionally driven by chemical intuition and time-consuming trial-and-error methods. To address the difficulties in probe screening for FET sensor development, this work presents a methodology that combines neuromorphic machine learning (ML) architectures, specifically a hybrid spiking graph neural network (SGNN), with an enriched dataset of physicochemical properties through semi-automated data extraction using large language models. Achieving a classification accuracy of 0.89 in predicting sensor sensitivity categories, the SGNN model outperformed traditional ML techniques by leveraging its ability to capture both global physicochemical properties and sparse topological features through a hybrid modeling framework. Next-generation sensor design was informed by the actionable insights into the connections between material properties and sensing performance offered by the SGNN framework. Through virtual screening for the detection of per- and polyfluoroalkyl substances (PFAS) as a use case, the effectiveness of the SGNN model was further validated. Density functional theory simulations confirmed graphene as a promising active material for PFAS detection as suggested by the SGNN framework. By bridging gaps in predictive modeling and data availability, this integrated approach provides a strong foundation for accelerating advancements in FET sensor design and innovation.

Ferreira, Rodrigo Pires [Univ. of Chicago, IL (Uni↗

An atomic cluster expansion (ACE) potential for water under extreme conditions

We present a machine learning interatomic potential for water designed to capture its complex multiphase behavior, including both molecular and superionic ice phases. The potential is based on the atomic cluster expansion (ACE) formulation and has been parameterized to enable high-fidelity molecular dynamics simulations of water under extreme conditions, for pressures up to 100 GPa and for temperatures between 500 and 6000 K. A diverse range of configurations was generated through ab initio molecular dynamics (AI-MD) simulations, covering insulating and superionic ice phases, liquid water, and dissociated plasma phase. We demonstrate that the H 2 O ACE potential accurately reproduces experimental and DFT predicted isotherms and Hugoniots. Crucially, the potential is able to capture the intricate phase behavior of water, including the transition from molecular fluid to the appropriate solid ice phases, and the superionic ice phases. This work provides a robust interatomic potential that can be used for large-scale, accurate simulations of water under extreme thermodynamic conditions.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

2025 Workshop on Envisioning Frontiers in AI and Computing for Biological Research: Position Papers

This workshop aims to identify key research directions for transforming biology using artificial intelligence (AI), machine learning (ML) and computational methods to facilitate the discovery of new behaviors, mechanisms, and designs of biological processes relevant to DOE missions, underpinning a broader U.S. bioeconomy. By developing novel AI/ML technologies to analyze and interpret complex biological data, researchers can organize and simulate biological processes at various scales as well as advance predictive understanding and manipulation of biological systems. This integration of computation, experimentation, and next-generation experimental technologies can lead to discoveries in new biological behaviors and mechanisms relevant to DOE missions. The focus is on how advanced computational and mathematical methods can impact this mission by exploring digital twins, foundation models, automated laboratory experiments, modeling of complex living systems, and data-driven approaches for the biodesign of plants and microbial systems. While data management is important, it is not the primary focus of this workshop, which will assess the current state, trends, and AI/ML challenges at the interface between biology and computational science to identify opportunities for high-impact research at their intersection. The goal is to define research needs and opportunities that align with biological sciences, computational sciences, and applied mathematics research.

59 BASIC BIOLOGICAL SCIENCES↗

Evaluating Acoustic vs. AI-Based Satellite Leak Detection in Aging US Water Infrastructure: A Cost and Energy Savings Analysis

The aging water distribution system in the United States, constructed mainly during the 1970s with some pipes dating back 125 years, is experiencing significant deterioration leading to substantial water losses. Along with the potential for water loss savings, improvements in the distribution system by using leak detection technologies can create net energy and cost savings. In this work, a new framework has been presented to calculate the economic level of leakage within water supply and distribution systems for two primary leak detection technologies (acoustic vs. satellite). In this work, a new framework is presented to calculate the economic level of leakage (ELL) within water supply and distribution systems to support smart infrastructure in smart cities. A case study focused using water audit data from Atlanta, Georgia, compared the costs of two leak mitigation technologies: conventional acoustic leak detection and artificial intelligence–assisted satellite leak detection technology, which employs machine learning algorithms to identify potential leak signatures from satellite imagery. The ELL results revealed that conducting one survey would be optimum for an acoustic survey, whereas the method suggested that it would be expensive to utilize satellite-based leak detection technology. However, results for cumulative financial analysis over a 3-year period for both technologies revealed both to be economically favorable with conventional acoustic leak detection technology generating higher net economic benefits of USD 2.4 million, surpassing satellite detection by 50%. A broader national analysis was conducted to explore the potential benefits of US water infrastructure mirroring the exemplary conditions of Germany and The Netherlands. Achieving similar infrastructure leakage index (ILI) values could result in annual cost savings of $\$4$–$\$4.8$ billion and primary energy savings of 1.6–1.9 TWh. These results demonstrate the value of combining economic modeling with advanced leak detection technologies to support sustainable, cost-efficient water infrastructure strategies in urban environments, contributing to more sustainable smart living outcomes.

acoustic leak detection↗

How to Build a Quantum Supercomputer: Scaling from Hundreds to Millions of Qubits

In the span of four decades, quantum computation has evolved from an intellectual curiosity to a potentially realizable technology. Today, small-scale demonstrations have become possible for quantum algorithmic primitives on hundreds of physical qubits and proof-of-principle error-correction on a single logical qubit. Nevertheless, despite significant progress and excitement, the path toward a full-stack scalable technology is largely unknown. There are significant outstanding quantum hardware, fabrication, software architecture, and algorithmic challenges that are either unresolved or overlooked. These issues could seriously undermine the arrival of utility-scale quantum computers for the foreseeable future. Here, we provide a comprehensive review of these scaling challenges. We show how the road to scaling could be paved by adopting existing semiconductor technology to build much higher-quality qubits, employing system engineering approaches, and performing distributed quantum computation within heterogeneous high-performance computing infrastructures. These opportunities for research and development could unlock certain promising applications, in particular, efficient quantum simulation/learning of quantum data generated by natural or engineered quantum systems. To estimate the true cost of such promises, we provide a detailed resource and sensitivity analysis for classically hard quantum chemistry calculations on surface-code error-corrected quantum computers given current, target, and desired hardware specifications based on superconducting qubits, accounting for a realistic distribution of errors. Furthermore, we argue that, to tackle industry-scale classical optimization and machine learning problems in a cost-effective manner, heterogeneous quantum-probabilistic computing with custom-designed accelerators should be considered as a complementary path toward scalability.

Mohseni, Masoud↗

Automating Anomaly Detection for Target systems at Spallation Neutron Source

The Spallation Neutron Source (SNS) at Oak Ridge National Laboratory, produces the world’s most intense pulse neutrons beams. An accelerated proton beam is directed into a mercury target to generate neutrons via spallation. The target system accounted for over 40% of the overall downtime of the facility in 2022. Thus, early detection in anomalies in the target systems can enable taking corrective actions to avoid failures and reduce downtime. Fault prognostics and anomaly detection in accelerators, both at SNS and outside, has largely focused on the beam side. This paper presents one the first studies exploring leveraging machine learning to automate the detection of anomalies in the target system. The target system consists of over 30 different interconnected subsystems, and the present work focuses on the mercury process system as a use case. Analyzing data from 28 process variables from 2022 and 2023, tree-based and reconstruction-based algorithms are employed to detect anomalies in archived data. The algorithms detected previously unreported anomalies, several of which were deemed alert worthy by human experts, particularly those found by reconstruction-based algorithms. Using data from each production run in the accelerator increased the generalizability of the models in time. Efforts are now underway to implement a workflow for incorporating human feedback to update the models and evaluating performance on unseen data. The models will eventually be integrated into the existing System Tracking and Reliability system with a web interface for automated anomaly detection and reporting along with a pathway for incorporating human feedback for model updates.

Raj, Anant [ORNL] (ORCID:0000000306711244)↗

Comparative modeling reveals the molecular determinants of aneuploidy fitness cost in a wild yeast model

Although implicated as deleterious in many organisms, aneuploidy can underlie rapid phenotypic evolution. However, aneuploidy will be maintained only if the benefit outweighs the cost, which remains incompletely understood. To quantify this cost and the molecular determinants behind it, we generated a panel of chromosome duplications in Saccharomyces cerevisiae and applied comparative modeling and molecular validation to understand aneuploidy toxicity. We show that 74%–94% of the variance in aneuploid strains’ growth rates is explained by the cumulative cost of genes on each chromosome, measured for single-gene duplications using a genomic library, along with the deleterious contribution of small nucleolar RNAs (snoRNAs) and beneficial effects of tRNAs. Machine learning to identify properties of detrimental gene duplicates provided no support for the balance hypothesis of aneuploidy toxicity and instead identified gene length as the best predictor of toxicity. Our results present a generalized framework for the cost of aneuploidy with implications for disease biology and evolution.

genic load↗

ChemEcho v1.0

ChemEcho is a tool that converts tandem mass spectra into embeddings used to build machine learning (ML) models with fully explainable predictions. It provides an API for transforming raw tandem mass spectral data into embeddings, along with functions for training and validating ML models. Additionally, it includes utilities for retrieving and cleaning training data. ChemEcho is broadly applicable in ML pipelines that use tandem mass spectra for a variety of tasks, such as chemical classification or bioactivity mining. While there are existing methods to generate embeddings from fragmentation data, ChemEcho's approach ensures that predictions remain interpretable, enabling experts to evaluate results and generate hypotheses about the underlying data.

Harwood, Thomas [Lawrence Berkeley National Labora↗