Search NASA⌕ Search

SEARCH · Search NASA

Results for “Generative Machine Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35

A Quantum-Assisted Algorithm for Sampling Applications in Machine Learning

An increase in the efficiency of sampling from Boltzmann distributions would have a significant impact in deep learning and other machine learning applications. Recently, quantum annealers have been proposed as a potential candidate to speed up this task, but several limitations still bar these state-of-the-art technologies from being used effectively. One of the main limitations is that, while the device may indeed sample from a Boltzmann-like distribution, quantum dynamical arguments suggests it will do so with an instance-dependent effective temperature, different from the physical temperature of the device. Unless this unknown temperature can be unveiled, it might not be possible to effectively use a quantum annealer for Boltzmann sampling. In this talk, we present a strategy to overcome this challenge with a simple effective-temperature estimation algorithm. We provide a systematic study assessing the impact of the effective temperatures in the learning of a kind of restricted Boltzmann machine embedded on quantum hardware, which can serve as a building block for deep learning architectures. We also provide a comparison to k-step contrastive divergence (CD-k) with k up to 100. Although assuming a suitable fixed effective temperature also allows to outperform one step contrastive divergence (CD-1), only when using an instance-dependent effective temperature we find a performance close to that of CD-100 for the case studied here. We discuss generalizations of the algorithm to other more expressive generative models, beyond restricted Boltzmann machines.

Perdomo-Ortiz, Alejandro↗

Scalable Hybrid Learning Techniques for Scientific Data Compression

Data compression is becoming critical for storing scientific data because many scientific applications need to store large amounts of data and post process this data for scientific discovery. Unlike image and video compression algorithms that limit errors to primary data (PD), scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Here, this article presents a physics-informed compression technique implemented as an end-to-end, scalable, GPU-based pipeline for data compression that addresses this requirement. Our hybrid compression technique combines machine learning techniques and standard compression methods. Specifically, we combine an autoencoder, an error-bounded lossy compressor to provide guarantees on raw data error, and a constraint satisfaction post-processing step to preserve the QoIs within a minimal error (generally less than floating point error). The effectiveness of the data compression pipeline is demonstrated by compressing nuclear fusion simulation data generated by a large-scale fusion code, XGC, which produces hundreds of terabytes of data in a single day. Our approach works within the ADIOS framework and results in compression by a factor of more than 150 while requiring only a few percent of the computational resources necessary for generating the data, making the overall approach highly effective for practical scenarios.

ITER↗

Mapping wall-to-wall fractional cover of Arctic tundra plant functional types in Alaska using 20-m spatial resolution satellite imagery and harmonized plot observations

Estimates of fractional cover (fCover) across given land surfaces are used to assess, and often model, vegetation composition and diversity, which are crucial for understanding the health and functioning of terrestrial ecosystems. Remote sensing provides a useful means for scaling local, plot-measured fCover estimates to regional scales. Leveraging a recently synthesized and harmonized plot database, this study generated wall-to-wall maps of fCover for six Alaskan-Arctic plant functional types (PFT), including non-vascular plants, forbs, graminoids, and deciduous and evergreen shrubs, using 20-m satellite data (Sentinel-1, Sentinel-2, ArcticDEM) using a machine learning regression approach, specifically the random forest (RF) algorithm, which is well-suited for handling nonlinear relationships and high-dimensional satellite datasets. This study additionally addressed the spatio-temporal inconsistencies e.g., sampling scale, plot size, and collection year in plot measured fCover by adopting a multivariate outlier detection approach—Cook’s distance—to identify high-quality plots for model training and validation. Our approach achieves high accuracy (R 2 = 0.59–0.93, root mean squared errors = 0.02–0.10 for all PFTs) between plot-observed and satellite-derived fCover when using high-quality plot samples. The mapped fCover characterizes the spatial patterns of different PFTs across the tundra biome at a 20-m resolution, providing key information needed for improved representation of Arctic tundra vegetation in terrestrial biosphere models to better understand climate-vegetation feedback across the Arctic tundra.

Arctic tundra↗

Simultaneous Unbinned Differential Cross-Section Measurement of Twenty-Four Z+jets Kinematic Observables with the ATLAS Detector

Z boson events at the Large Hadron Collider can be selected with high purity and are sensitive to a diverse range of QCD phenomena. As a result, these events are often used to probe the nature of the strong force, improve Monte Carlo event generators, and search for deviations from standard model predictions. All previous measurements of Z boson production characterize the event properties using a small number of observables and present the results as differential cross sections in predetermined bins. In this analysis, a machine learning method called omnifold is used to produce a simultaneous measurement of twenty-four Z +jets observables using 139 fb -1 of proton-proton collisions at √s =13 TeV collected with the ATLAS detector. Unlike any previous fiducial differential cross-section measurement, this result is presented unbinned as a dataset of particle-level events, allowing for flexible reuse in a variety of contexts and for new observables to be constructed from the twenty-four measured observables.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

LeWRON: Learning ElectroWeak phase tRansitiON with agentic architecture

An agent to analyze electroweak phase transition. LeWRON turns a BSM model description or a reproduction target into a structured run: symbolic setup, effective-potential artifacts, finite-temperature machinery, generated model code, a scientific report, and an interactive exploration session. It keeps both machine-readable artifacts and human-readable notes, so a run can be resumed, audited, revised, and shared.

Wang, Isaac [Fermi National Accelerator Laboratory↗

Quantum annealing-assisted lattice optimization

High Entropy Alloys (HEAs) have drawn great interest due to their exceptional properties compared to conventional materials. The configuration of HEA system is considered a key to their superior properties, but exhausting all possible configurations of atom coordinates and species to find the ground energy state is extremely challenging. In this work, we proposed a quantum annealing-assisted lattice optimization (QALO) algorithm, which is an active learning framework that integrates the Field-aware Factorization Machine (FFM) as the surrogate model for lattice energy prediction, Quantum Annealing (QA) as an optimizer and Machine Learning Potential (MLP) for ground truth energy calculation. By applying our algorithm to the NbMoTaW alloy, we reproduced the Nb depletion and W enrichment observed in bulk HEA. We found our optimized HEAs to have superior mechanical properties compared to the randomly generated alloy configurations. Our algorithm highlights the potential of quantum computing in materials design and discovery, laying a foundation for further exploring and optimizing structure-property relationships.

36 MATERIALS SCIENCE↗

On the universality of S n -equivariant k -body gates

The importance of symmetries has recently been recognized in quantum machine learning from the simple motto: if a task exhibits a symmetry (given by a group $\mathfrak{G}$), the learning model should respect said symmetry. This can be instantiated via $\mathfrak{G}$-equivariant quantum neural networks (QNNs), i.e. parametrized quantum circuits whose gates are generated by operators commuting with a given representation of $\mathfrak{G}$. In practice, however, there might be additional restrictions to the types of gates one can use, such as being able to act on at most k qubits. In this work we study how the interplay between symmetry and k-bodyness in the QNN generators affect its expressiveness for the special case of $\mathfrak{G}=S_n$, the symmetric group. Our results show that if the QNN is generated by one- and two-body Sn-equivariant gates, the QNN is semi-universal but not universal. That is, the QNN can generate any arbitrary special unitary matrix in the invariant subspaces, but has no control over the relative phases between them. Then, we show that in order to reach universality one needs to include n-body generators (if n is even) or ($n-1$)-body generators (if n is odd). As such, our results brings us a step closer to better understanding the capabilities and limitations of equivariant QNNs.

97 MATHEMATICS AND COMPUTING↗

Iterative ML and Experiments for Emerging VOCs

SAND2026-17074O Iterative ML and Experiments for Emerging VOCs is a tool that analyzes and predicts the behaviors of SARS-CoV-2 variants. It processes experimental data on ACE2 (the receptor for the SARS-CoV-2 virus that allows it to infect the cell) and antibody binding using machine learning models, including neural networks, to forecast ACE2 interactions and variant expression. The tool employs transfer learning and global epistasis modeling, integrating public datasets with proprietary data to enhance prediction accuracy. Additionally, it fits concentration-response curves to determine dissociation constants and generates visualizations to support research findings, thereby aiding in the identification of new antibodies for emerging variants of concern. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Sheffield, Thomas [Sandia National Lab. (SNL-NM), ↗

Intelligent Triggers for Rare Event Detection in Liquid Argon Detectors

Next-generation neutrino experiments like SBND and DUNE rely on Liquid Argon Time Projection Chambers (LArTPCs), which produce exceptionally detailed data at high volume. Capturing rare or unexpected events in real-time is a major challenge. Our project explores the use of machine learning, specifically autoencoder-based anomaly detection, to identify unusual activity directly from raw detector signals. Inspired by successes at the CMS experiment, we demonstrate that such methods can be adapted to LArTPCs and show promising results in both simulated studies and early steps toward real-time hardware deployment. This approach could open new avenues for detecting signals from physics beyond the Standard Model.

Chung, Seokju [Columbia U. (main)]↗

Physics-Informed Machine Learning Model for Ceramic Matrix Composite Creep

A physics-informed recurrent neural network (RNN) based surrogate model is developed to emulate the nonlinear, time-dependent constitutive behavior of ceramic matrix composites (CMCs) driven by matrix damage and constituent creep at the microscale. Physics-informed constraints are introduced into the surrogate model through regularization to ground the prediction in physics and improve its predictive capabilities. Training data is generated using the high-fidelity generalized method of cells (HFGMC) approach which calls appropriate creep and damage models for each of the constituents. This coupling permits simulating the nonlinear behavior of CMCs based on constituent response at the microscale along with microstructural features such as fiber and porosity volume fraction and fiber radius. The microscale repeating unit cell is loaded under creep fatigue conditions to replicate the material loading experienced in a turbine engine. Therefore, the RNN-based surrogate model is tasked with predicting, as a function of variable input stress sequence, temperature, and microstructural features, the resulting strain history response while satisfying physical constraints related to creep rate, isochoric inelastic deformation, and strain energy density. The trained surrogate model is shown to effectively match the strain history over quantified distributions of microstructural features and relevant loading regimes and temperatures. Neural network based surrogate models can offer efficient alternatives to running computationally intensive multiscale material models to simulate the nonlinear response of large structural models. Therefore, the presented work provides evidence towards the feasibility of developing, training, and running such models for CMCs with complex microstructures, nonlinear time-dependent material response, and under non-monotonic loading conditions.

ceramic matrix composites↗

Treyson Ricks - Intern Showcase Poster

Quinone-based sorbents offer a tunable, energy-efficient route to electrochemical CO2 capture, but systematic guidance for molecular design is lacking. Here, we report a high-throughput computational workflow that combines density functional theory (DFT) screening with machine-learning (ML) modeling to evaluate CO2 binding thermodynamics across several quinone derivatives, spanning benzoquinones, naphthoquinones, and anthraquinones. In addition to using solvents to stabilize the quinone anion and dianion, we studied the effect of ion-pairing on the reduction potentials and the CO2 binding energy. Automated Python scripts handled geometry optimizations and adduct-formation energies on an HPC cluster, reducing manual effort significantly. This integrated platform can uncover structure–property relationships and enables rapid in silico evaluation of untested candidates. We present one example from our workflow to showcase the capability of using quinones with ion-pairing to effectively capture CO2. Our approach paves the way for the rational selection of optimal quinone sorbents and can be extended with experimental thermochemical and kinetic data, alternative redox cycles, and stability assessments to accelerate development of next-generation electrochemical CO2 capture materials.

37 - INORGANIC, ORGANIC, PHYSICAL AND ANALYTICAL C↗

Machine learning assisted prediction of tungsten heavy alloy plasma facing component performance for fusion energy applications

Tungsten and tungsten heavy alloys (WHAs), known for their remarkably high hardness, durability, and corrosion resistance, play a critical role in the thriving development of nuclear fusion reactors in recent years. However, the exploration in tungsten alloys for the nuclear-related applications has been limited by the difficulty of manufacturing and the complexity of experiments to reproduce the environment of nuclear reaction. Therefore, this project aims to utilize nanoscale simulation methods such as density functional theory (DFT) and molecular dynamics (MD) with the help of machine learning techniques to not only understand the mechanisms of tungsten alloys but also allow us to computationally predict their mechanical behaviors under extreme environments. One critical problem of the application of WHAs in nuclear reactors is the surface melting. In the current design of the SPARC reactor, the WHA, W97Ni2.1Fe0.9 or W97NiFe, is chosen to be the first wall components to confine the plasma where the particles are fiercely moving and colliding into each other to create nuclear fusion reaction. This process will generate extremely high heat flux onto these WHA tiles, leaving high surface temperature that could possibly melt the surface of the WHA tiles, As illustrated in Fig. 1(a). a laser experiment previously done illustrates that a rough surface damage would be made after the surface melting where the matrix area mainly composed of nickel and iron as shown in Fig. 1(b), will first melt and then leave vacancies between these tungsten grains. Unfortunately, these kinds of roughness on the first-wall components could be deadly to the plasma inside a Tokmak reactor because the heat that is supposed to dissipate at a designed ratio through the tiles may in turn be excessively absorbed and accumulated on any uneven area of the surface, which will eventually make the whole nuclear reaction fail. In this project, we will introduce a machine learning potential, Allegro, based on DFT calculation and then build a MD model for W-Ni-Fe alloys.

36 MATERIALS SCIENCE↗

Failure Analysis–Informed Risk Assessment Framework for Geological Carbon Storage Using Numerical Simulation and Machine Learning

Geological carbon storage (GCS) is recognized as a critical technology for achieving large-scale reductions in anthropogenic carbon dioxide (CO 2 ) emissions. Ensuring long-term containment and safety requires robust risk assessment frameworks that account for geological uncertainty and identify potential failure scenarios. Among various indicators, the area of review (AoR) serves as a key metric for evaluating storage performance, regulatory compliance, and monitoring design, as it delineates the spatial extent impacted by pressure buildup and plume migration. However, conventional AoR-based risk assessments typically perturb parameters within narrow uncertainty bounds, potentially overlooking rare but high-impact events arising from extreme geological conditions. In this study, we present a failure analysis–informed risk assessment framework for large-scale GCS projects to improve site prescreening and monitoring design. A suite of 300 numerical simulations was generated using stochastic geological models that vary five key parameters: net-to-gross ratio, anisotropy azimuth, porosity multiplier, permeability multiplier, and vertical-to-horizontal permeability ratio. Among these, 200 realizations represent normal geological uncertainty, while 100 additional cases explore extreme yet plausible conditions for failure-case analysis. The AoR was simulated and computed from pressure and CO 2 saturation fields, where the baseline AoR boundary, representing the extent predicted under typical geological uncertainty, was defined as the union of 200 normal-range simulations, and failure was identified when extreme-range cases exceeded this baseline. Results show that incorporating broader parameter uncertainty produces significantly larger AoR extents, underscoring the potential underestimation of risk under conventional uncertainty ranges. Furthermore, spatial probability maps derived from failure-induced AoR exceedance identify regions requiring enhanced monitoring attention. Various machine learning (ML)–based classifiers were developed to predict failure occurrence from geological parameters, with the random forest model achieving the highest performance (F1-score of 0.986). Consistent findings from correlation coefficient, feature importance, and Sobol sensitivity analyses reveal that low net-to-gross ratios and permeability multipliers are the dominant risk drivers, reflecting reduced reservoir connectivity and limited pressure dissipation. Altogether, these results provide a novel framework for risk-informed site prescreening and monitoring design that explicitly considers rare but high-impact geological scenarios in GCS projects.

25 ENERGY STORAGE↗

RCSB protein data Bank: Next‐generation advanced search for exploration of experimental structures and computed structure models

Abstract The Protein Data Bank (PDB), established in 1971, is the primary global, open‐access archive for experimentally determined 3D macromolecular structures (proteins, RNA, DNA). The research‐focused RCSB.org web‐portal provides access to these data alongside more than one million machine‐learning‐predicted structure models, greatly expanding the available structural landscape. Rapid growth of both experimental and computational structures has increased the need for powerful yet accessible search tools that serve a broad and diverse scientific community. Herein, we describe a redesigned RCSB Protein Data Bank RCSB.org Advanced Search capability that supports intuitive discovery of 3D structures through a unified interface. This interface integrates annotation‐, sequence‐, and 3D structure‐based searches, embeds an interactive 3D viewer, and incorporates curated biological knowledge, such as catalytic site definitions from Mechanism and Catalytic Site Atlas and ligand‐guided structural motifs, for constructing geometry‐driven queries. A new Chemical Search tool allows definition of chemical queries via an integrated drawing tool or standard identifiers, seamlessly combining them with annotation filters. By allowing query definition directly within spatial and chemical contexts, these search interfaces reduce the need for detailed knowledge of residue numbering, chain identifiers, or external cheminformatics software. This capability enables efficient exploration of structures, chemical diversity, and structure–function relationships across all life domains. The redesigned interfaces can be accessed directly at rcsb.org/search/advanced for Advanced Search and rcsb.org/search/chemical for Chemical Search.

Rose, Yana [Research Collaboratory for Structural ↗

Seismicity-constrained fault detection and characterization with a multitask machine learning model

Geological fault detection and characterization are crucial for understanding subsurface dynamics across scales. While methods for fault delineation based on either seismicity location analysis or seismic image reflector discontinuity are well-established, a systematic approach that integrates both data types remains absent. We develop a novel machine learning model that unifies seismic reflector images and seismicity location information to automatically identify geological faults and characterize their geometrical properties. The model encodes a seismic image and a seismicity location image separately, and fuses the encoded features with a spatial-channel attention fusion module to improve the learning of important features in both inputs. We design an automated strategy to generate high-quality synthetic training data and labels. To improve the realism of the seismicity location image, we include random seismicity noise and missing seismicity location associated with some of the faults. We validate the model’s efficacy and accuracy using synthetic data examples and two field data examples. Moreover, we show that fine-tuning the trained model with a small, domain-specific dataset enhances its fidelity for field data applications. The results demonstrate that integrating seismicity location and seismic images into a unified framework allows the end-to-end neural network to achieve higher fidelity and accuracy in delineating subsurface faults and their geometrical properties compared with image-only fault detection methods. Our approach offers an adaptive data-driven tool for geological fault characterization and seismic hazard mitigation, bridging the gap between seismicity location and image-based fault detection methods.

58 GEOSCIENCES↗

A knowledge-informed large language model framework for U.S. nuclear power plant shutdown initiating event classification for probabilistic risk assessment

Identifying and classifying shutdown initiating events (SDIEs) is critical for developing shutdown probabilistic risk assessment for nuclear power plants. Existing computational approaches cannot achieve satisfactory performance due to the challenges of unavailable large, labeled datasets, imbalanced event types, and label noise. To address these challenges, we propose a hybrid pipeline that integrates a knowledge-informed machine learning model to prescreen non-SDIEs and a large language model (LLM) to classify SDIEs into four types. In the prescreening stage, we proposed a set of 44 SDIE text patterns that consist of the most salient keywords and phrases from six SDIE types. Text vectorization based on the SDIE patterns generates feature vectors that are highly separable by using a simple binary classifier. The second stage builds Bidirectional Encoder Representations from Transformers (BERT)-based LLM, which learns generic English language representations from self-supervised pretraining on a large dataset and adapts to SDIE classification by fine-tuning it on an SDIE dataset. The proposed approaches are evaluated on a dataset with 10,928 events using precision, recall ratio, F 1 score, and average accuracy. In conclusion, the results demonstrate that the prescreening stage can exclude more than 97% non-SDIEs, and the LLM achieves an average accuracy of 95.1% for SDIE classification.

99 - GENERAL AND MISCELLANEOUS↗

A general mechanistic framework for cross-scale understanding of hot spots and hot moments in carbon and water fluxes

Semi-arid ecosystems, like those in the American Southwest, exert a massive impact on the interannual variability of carbon and water cycling. Unfortunately, these carbon and water fluxes are notoriously difficult to predict due to their high spatial and temporal variability, which is poorly captured by the current generation of vegetation models. Indeed, this region is exemplified by the ‘hot spots and hot moments’ concept, which states that small areas in space (‘hot spots’) and transient moments in time (‘hot moments’) exert an outsized influence on biogeochemical cycling. However, the factors that regulate these pulses in biogeochemical activity are unknown, as is their variability across space and time. These uncertainties severely limit efforts to better represent hot spots and hot moments in models. Here, we seek to develop a generalized method for detecting and quantifying the importance of hot spots and hot moments from individual plant to regional scales. Underpinning this method is our recently developed statistical approach for identifying hot spots and hot moments. By applying this method to semi-continuous measurements of plant water status, a depth profile of soil water potential, and ecosystem fluxes via eddy covariance, we will track the fate of water through the soil-plant-atmosphere continuum and identify the mechanistic drivers of these transient pulses in biogeochemical activity. Then, we will expand this approach across a broad network of Ameriflux towers, and apply a machine learning approach that will allow us to upscale measurements of hot spots and hot moments across the American Southwest and quantify their impact on carbon and water cycles. These products will allow us to identify hot spots and hot moments across spatio-temporal scales and will serve as crucial data sources for validating a new generation of models that can better capture highly dynamic carbon and water fluxes. The proposed method will be easily transferable across biomes and will serve as a framework for future research on hot spots and hot moments across the plant ecophysiology, biometeorology, and vegetation modeling communities.

54 ENVIRONMENTAL SCIENCES↗

A novel approach for large-scale wind energy potential assessment

Increasing wind energy generation is central to grid decarbonization, yet methods to estimate wind energy potential are not standardized, leading to inconsistencies and even skewed results. This study aims to improve the fidelity of wind energy potential estimates through an approach that integrates geospatial analysis and machine learning (i.e., Gaussian process regression). We demonstrate this approach to assess the spatial distribution of wind energy capacity potential in the Contiguous United States (CONUS). We find that the capacity-based power density ranges from 1.70 MW/km2 (25th percentile) to 3.88 MW/km2 (75th percentile) for existing wind farms in the CONUS. The value is lower in agricultural areas (2.73 ± 0.02 MW/km2, mean ± 95 % confidence interval) and higher in other land cover types (3.30 ± 0.03 MW/km2). Notably, advancements in turbine manufacturing could reduce power density in areas with lower wind speeds by adopting low specific-power turbines, but improve power density in areas with higher wind speeds (>8.35 m/s at 120m above the ground), highlighting opportunities for repowering existing wind farms. Wind energy potential is shaped by wind resource quality and is regionally characterized by land cover and physical conditions, revealing significant capacity potential in the Great Plains and Upper Texas. The results indicate that areas previously identified as hot spots using existing approaches (e.g., the west of the Rocky Mountains) may have a limited capacity potential due to low wind resource quality. Improvements in methodology and capacity potential estimates in this study could serve as a new basis for future energy systems analysis and planning.

Dai, Tao↗