Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Data Summarization and Inference at Scale

This is the final report for the DOE ASCR grant SC-0022260, Data Summarization and Inference at Scale, PI: Alex Pothen, Purdue University. The goal of the project was to solve data-intensive and compute-intensive problems in the physical sciences, engineering, information science, data science, etc. by designing and implementing new algorithms that could work with a subset of the data. The four subgoals were: (a) The solution of problems where the data is too large to be stored in the memory of a computer. In this streaming model of computation, the data arrives as a stream of elements to the computer, each element is processed as it arrives, and a decision is made to discard the data or to store it; only a small subset of the data proportional to the size of the output solution is stored, and when all the data has been streamed, a solution to the problem is computed from the stored subset. (b) The use of machine learning methods to compute solutions to data-intensive problems. The use of GPUs is critical to obtain high performance on machine learning tasks, but their memory sizes are smaller relative to that of CPUs. For large-scale problems, the data is sampled many times, and small samples are used with repetition, for robustness, to compute solutions to inference tasks. This sampling reduces the memory required to solve the problem, but attention is needed to avoid slow convergence to the solutions, and reduced accuracy of inference. We propose submodular optimization, Large Language Models, and physics-informed neural networks to enable GPU computations here. (c) Modeling and visualization of high-dimensional data using interpretable features. Clinical proteomic data sets from immunology for the detection of cancer and other diseases are temporal and high-dimensional, and algorithms for visualizing these data sets using clinically interpretable features are lacking. We propose methods that compute distances based on the optimal transportation problem and graph edit distances to address this problem. We also propose the use of optimal transport-based distances, spatial statistics, and network structure to classify image data sets, We apply these algorithms to electron micrographs of the peripheral nervous system in the digestive tract. (d) The design of data-intensive algorithms on emerging architectures, specifically, noisy, intermediate-scale quantum (NISQ) devices. Quantum computers offer the possibility of exploring large solution spaces due to the principle of superposition, but current quantum computers are limited by few qubits, short coherence times due to noise, poor interconections among the qubits, etc. We propose the use of the divide and conquer paradigm to solve large-scale problems, wherein collections of small subproblems are solved on the quantum devices, and the solutions to the subproblems are integrated into a solution for the original problem on a classical computer.

97 MATHEMATICS AND COMPUTING↗

Precise Modeling of a Complex Solenoidal Magnetic Field Using a Combination of Analytic Functions and a PINN

We demonstrate an iterative approach to modeling a sparsely measured magnetic field in a large-bore solenoid. This approach uses a hybrid of traditional and machine learning techniques. The traditional technique is a linear least-squares fit using a series solution to Laplace's equation, while the machine learning technique involves the training of a physics-informed neural network (PINN) on the least-squares fit residuals. We use a newly defined activation function "DELTAsnake," a modification to the snake activation function proposed by Ziyin et al. that allows for stronger curvature and non-monotonicity. The combined model approximately obeys Maxwell's equations to a level sufficient for producing high quality physics simulations and analysis. Our approach is applied to a highly realistic calculation of the expected magnetic field in the Mu2e experiment's Detector Solenoid which includes a simple model for the expected statistical measurement uncertainties. Using ten toy measurement simulations, we demonstrate the capabilities of our model in comparison to the least-squares method alone; the least-squares method alone results in a reduced chi-squared statistic of ${2.15 \pm 0.01}$, while our approach improves the reduced chi-square to ${1.034 \pm 0.005}$. Furthermore, for an average toy simulation, we show that the range of the RMS of the three field component residuals reduces from ${0.07-0.37}$ Gauss to ${0.05-0.07}$ Gauss. We find that this novel method is robust against a realistic systematic uncertainty deriving from Hall probe calibration bias and can be used to significantly reduce the number of measurements required to achieve an accurate model.

Kampa, Cole [Caltech] (ORCID:0000000192972920)↗

Federated Learning for Efficient Condition Monitoring and Anomaly Detection in Industrial Cyber-Physical Systems

Detecting and localizing anomalies in cyber-physical systems (CPS) has become increasingly challenging as systems grow in complexity, particularly due to varying sensor reliability and node failures in distributed environments. While federated learning (FL) offers a foundation for distributed model training, existing approaches lack mechanisms to handle these CPS-specific challenges. This paper presents an enhanced FL framework that introduces three key innovations: adaptive model aggregation based on sensor reliability, dynamic node selection for resource optimization, and Weibull-based checkpointing for fault tolerance. Our framework enables reliable condition monitoring while addressing the computational and reliability challenges of industrial CPS deployments. Experiments on NASA Bearing and Hydraulic System Datasets demonstrate superior performance over state-of-the-art FL methods, achieving 99.5% AUC-ROC in anomaly detection and maintaining accuracy under node failures. Statistical validation using Mann-Whitney (U) test confirms significant improvements (p < 0.05) in both detection accuracy and computational efficiency across diverse operational scenarios.1

Marfo, William [University of Texas at El Paso,Dep↗

RU Net for Automatic Characterization of TRISO Fuel Cross Sections

TRistructural ISOtropic (TRISO) particle fuel is a type of nuclear fuel known for its high-temperature and high-burnup performance. Each sub-millimeter diameter TRISO particle consists of uranium-oxycarbide (UCO) or UO2 fuel kernel, coated with buffer, inner pyrolytic carbon (IPyC), silicon carbide (SiC), and outer pyrolytic carbon (OPyC) layers. The SiC layer acts as the main containment barrier for the TRISO particle to retain the fission products, while the IPyC and OPyC layers provide additional barriers to the release of fission products, especially fission gases. During irradiation, phenomena like kernel swelling, buffer densification, and IPyC fracture may impact fuel performance. Post-irradiation microscopy on entire compact cross sections or samples of individual particles deconsolidated from compacts is often used to identify these irradiation-induced changes in morphology. However, each fuel compact generally contains thousands of TRISO particles. To get statistical information on these phenomena, it is cumbersome work if done manually. For example, to get information about swelling/densification behaviors of different layers or kernels after irradiation, researchers previously manually measured the perimeter of each TRISO layer in hundreds of particles after four rounds of iterative grinding and polishing encompassing more than 2000 cross-section images for a total of four fuel compacts. To attempt to reduce the subjectivity inherent in that process and accelerate data analysis, we conducted a study on the automatic TRISO layer segmentation on cross-sectional microscopic images using Convolutional Neural Networks (CNNs). CNNs are a class of machine learning algorithms specifically designed for processing structured grid data that have gained popularity in recent years due to their remarkable performance in various computer vision tasks, including image classification, object detection, and image segmentation. In this research, we have generated the large irradiated TRISO layer dataset with more than 2000 cross-section TRISO microscopic images and the corresponding annotated images. Based on these annotated images, we have employed different CNNs for automatic segmentation of different TRISO layers. These include RU-Net (developed in this study), as well as three existing architectures: U-Net, Residual Network (ResNet), and Attention U-Net. The preliminary results show that the model based on RU-Net has the best performance in terms of intersection-over-union (IoU). Through the aid of these CNN models, we can expedite the analysis of TRISO particle cross-sections, significantly reducing the manual labor involved and improving the objectivity of the segmentation results.

Convolutional Neural Networks↗

Real-Time event reconstruction for Nuclear Physics Experiments using Artificial Intelligence

Charged track reconstruction is a critical task in nuclear physics experiments, enabling the identification and analysis of particles produced in high-energy collisions. Machine learning (ML) has emerged as a powerful tool for this purpose, addressing the challenges posed by complex detector geometries, high event multiplicities, and noisy data. Traditional methods rely on pattern recognition algorithms like the Kalman filter, but ML techniques, such as neural networks, graph neural networks (GNNs), and recurrent neural networks (RNNs), offer improved accuracy and scalability. By learning from simulated and real detector data, ML models can identify and classify tracks, predict trajectories, and handle ambiguities caused by overlapping or missing hits. Moreover, ML-based approaches can process data in near-real-time, enhancing the efficiency of experiments at large-scale facilities like the Large Hadron Collider (LHC) and Jefferson Lab (JLAB). As detector technologies and computational resources evolve, ML-driven charged track reconstruction continues to push the boundaries of precision and discovery in nuclear physics. In these proceedings, we highlight advancements in charged track identification leveraging Artificial Intelligence within the CLAS12 detector, achieving a notable enhancement in experimental statistics compared to traditional methods. Additionally, we showcase real-time event reconstruction capabilities, including the inference of charged particle properties, such as momentum, direction, and species identification, at speeds matching data acquisition rates. These innovations enable the extraction of physics observables directly from the experiment in real-time.

Gavalian, Gagik (ORCID:0000000267385457)↗

A Machine Learning Framework for Modeling Ensemble Properties of Atomically Disordered Materials

Atomic disorder can strongly influence material properties such as charge transport, optical response, and catalytic activity. However, efficiently modeling these disorder effects remains challenging for first-principles methods due to the cost of sampling large configurational spaces and computing complex physical quantities. Recent advances of machine learning techniques, particularly graph neural networks (GNNs), has enabled the efficient and accurate predictions of complex material properties, offering promising tools for studying disordered systems. In this work, we present a general machine-learning-assisted computational framework that integrates equivariant GNNs with Monte Carlo simulations to compute the thermodynamic and ensemble-averaged functional properties of disordered materials. Using the surface-termination-disordered MXene monolayer Ti 3 C 2 T 2–x as a representative system, we find that electrical conductivity exhibits an emergent peak near the order–disorder phase transition temperature due to the interplay between electron scattering and doping. In contrast, optical conductivity remains largely insensitive to local atomic disorder and reflects the global surface chemical composition. These results highlight the role of atomic disorder in affecting material properties and demonstrate the potential of our approach for statistically modeling disorder effects in a wide range of materials such as high-entropy alloys and spin liquids.

MXene↗

Pathway-based analyses of gene expression profiles at low doses of ionizing radiation

Radiation exposure poses a significant threat to human health. Emerging research indicates that even low-dose radiation once believed to be safe, may have harmful effects. This perception has spurred a growing interest in investigating the potential risks associated with low-dose radiation exposure across various scenarios. To comprehensively explore the health consequences of low-dose radiation, our study employs a robust statistical framework that examines whether specific groups of genes, belonging to known pathways, exhibit coordinated expression patterns that align with the radiation levels. Notably, our findings reveal the existence of intricate yet consistent signatures that reflect the molecular response to radiation exposure, distinguishing between low-dose and high-dose radiation. Moreover, we leverage a pathway-constrained variational autoencoder to capture the nonlinear interactions within gene expression data. By comparing these two analytical approaches, our study aims to gain valuable insights into the impact of low-dose radiation on gene expression patterns, identify pathways that are differentially affected, and harness the potential of machine learning to uncover hidden activity within biological networks. This comparative analysis contributes to a deeper understanding of the molecular consequences of low-dose radiation exposure.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Large-scale tearing-mode hazard function analysis with standard matched equilibrium reconstructions

The association between features from standard tokamak equilibrium reconstructions and the onset of n = 1 tearing modes (TMs) is analyzed at scale. The TM onset rate is directly modeled with a ‘hazard’ function which gives the expected number of onsets (per unit time spent) in a given equilibrium parameter region. In particular the different statistical modeling performance achieved for magnetics-only reconstructions and motional Stark effect (MSE) enhanced reconstructions is studied. It is observed that a better hazard model for the TM onset rate can be built with the MSE-enhanced equilibria compared to the matched magnetics-only situation. This advantage disappears if internal profile details are withheld from the matched analysis. Plausibility of the hazard function is further demonstrated with visualizations of global trends in the operational space, and time-traces from specific tokamak discharges. As a result, TMs typically degrade tokamak plasma performance and may lead to plasma termination, motivating this statistical study.

equilibrium↗

Exploring Saccharomycotina Yeast Ecology Through an Ecological Ontology Framework

Yeasts in the subphylum Saccharomycotina are found across the globe in disparate ecosystems. A major aim of yeast research is to understand the diversity and evolution of ecological traits, such as carbon metabolic breadth, insect association, and cactophily. This includes studying aspects of ecological traits like genetic architecture or association with other phenotypic traits. Genomic resources in the Saccharomycotina have grown rapidly. Ecological data, however, are still limited for many species, especially those only known from species descriptions where usually only a limited number of strains are studied. Moreover, ecological information is recorded in natural language format limiting high throughput computational analysis. To address these limitations, we developed an ontological framework for the analysis of yeast ecology. A total of 1,088 yeast strains were added to the Ontology of Yeast Environments (OYE) and analyzed in a machine-learning framework to connect genotype to ecology. This framework is flexible and can be extended to additional isolates, species, or environmental sequencing data. Widespread adoption of OYE would greatly aid the study of macroecology in the Saccharomycotina subphylum.

59 BASIC BIOLOGICAL SCIENCES↗

Neural network based emulation of galaxy power spectrum covariances: A reanalysis of BOSS DR12 data

We train neural networks to quickly generate redshift-space galaxy power spectrum covariances from a given parameter set (cosmology and galaxy bias). This covariance emulator utilizes a combination of traditional fully connected network layers and transformer architecture to accurately predict covariance matrices for the high redshift, north galactic cap sample of the BOSS DR12 galaxy catalog. We run simulated likelihood analyses with emulated and brute-force computed covariances, and we quantify the network’s performance via two different metrics: (1) difference in Χ 2 and (2) likelihood contours for simulated BOSS DR 12 analyses. We find that the emulator returns excellent results over a large parameter range. We then use our emulator to perform a reanalysis of the BOSS HighZ NGC galaxy power spectrum, and find that varying covariance with cosmology along with the model vector produces Ω m = $0.27⁢6$$^{+0.013}_{–0.015}$, H 0 = 70.2 ± 1.9 km/s/Mpc, and σ 8 = $0.67⁢4$$^{+0.058}_{–0.077}$. These constraints represent an average 0.46⁢σ shift in best-fit values and a 5% increase in constraining power compared to fixing the covariance matrix (Ω m = 0.293 ± 0.017, H 0 = 70.3 ± 2.0 km/s/Mpc, σ 8 = $0.70⁢2$$^{+0.063}_{–0.075}$). As a result, this work demonstrates that emulators for more complex cosmological quantities than second-order statistics can be trained over a wide parameter range at sufficiently high accuracy to be implemented in realistic likelihood analyses.

79 ASTRONOMY AND ASTROPHYSICS↗

High-Resolution South American Wind Resource Data Downscaled with Generative Machine Learning Conditioned on Near-Surface Observations

High-resolution historical wind data was developed for the entirety of South America using the innovative Super-Resolution for Renewable Resource Data (sup3r) machine learning framework. The publicly available Sup3rWind South America dataset represents a significant advancement in wind resource data generation, leveraging generative machine learning conditioned on near-surface observations from the Meteorological Assimilation Data Ingest System (MADIS) to efficiently and accurately downscale coarse reanalysis data from the European Centre for Medium-Range Weather Forecasts (ERA5). This approach produces fine-scale, spatially and temporally coherent wind and meteorological fields hundreds of times more computationally efficient than traditional numerical weather modeling methods, enabling access to high-fidelity wind information across both continental and offshore regions. Sup3rWind South America builds on the earlier Sup3rWind Ukraine dataset through improvements in model architecture and outputs conditioned on near-surface observation inputs. As with the Ukraine data release, this dataset includes wind speed, wind direction, temperature, relative humidity, and pressure at a horizontal resolution of ~2 km, representing a 15x spatial enhancement relative to the 31 km ERA5 grid. Wind speed and direction are provided at 5-minute resolution, a 12x temporal refinement compared to the hourly ERA5 data, while temperature, relative humidity, and pressure remain at hourly resolution. The data covers all years from 2005 to 2024. Before downscaling, ERA5 inputs were bias-corrected using long-term monthly means and a limited number of quality-controlled observations to align large-scale statistics with regional conditions. The resulting dataset is the first publicly available high-resolution timeseries wind record that provides full spatial coverage of South America. Model validation demonstrates strong agreement with observations across several statistical metrics, consistent with other state-of-the-art high-resolution wind resource datasets. The potential applications of Sup3rWind South America span renewable energy resource assessment, energy system modeling, and grid resilience analysis. The 20-year record and high spatial and temporal resolution support accurate estimation of long-term energy yield and the economic feasibility of potential wind development sites. Continuous coverage across both continental and offshore regions enables comprehensive site prospecting within exclusive economic zones. The 2 km, 5-minute resolution data provide the spatial and temporal variability required for power system simulation, operational planning, and regional risk assessments.

17 WIND ENERGY↗

Combining Machine Learning and Comparative Effectiveness Methodology to Study Primary Care Pharmacotherapy Pathways for Veterans With Depression

Our objective is to demonstrate an innovative method combining machine learning with comparative effectiveness research techniques and to investigate a hitherto unstudied question about the effectiveness of common prescribing patterns. For Operation Enduring Freedom/Operation Iraqi Freedom veterans with major depressive disorder, we generate pharmacotherapy pathways (of antidepressants) using process mining and machine learning. We select the medication episodes that were started at subtherapeutic doses by the first assigned primary care physician and observe the paths that those medication episodes follow. Using 2-stage least squares, we test the effectiveness of starting at a low dose and staying low for longer versus ramping up fast while balancing observable and unobservable characteristics of patients and providers through instrumental variables. We leverage predetermined provider practice patterns as instruments. We collected outpatient pharmacy data for selective serotonin reuptake inhibitors and selective norepinephrine reuptake inhibitors, patient and provider characteristics (as control variables), and the instruments for our cohort. All data were extracted for the period between 2006 and 2020. There is a statistically significant positive effect (0.68, 95% CI 0.11–1.25) of “ramping up fast” on engagement in care. When we examine the effect of “ramping up slow”, we see an insignificant negative impact on engagement in care (−0.82, 95% CI −1.89 to 0.25). As expected, the probability of drop-out also seems to have a negative effect on engagement in care (−0.39, 95% CI −0.94 to 0.17). We further validate these results by testing with medication possession ratios calculated periodically as an alternative engagement in care metric. Our findings contradict the “Start low, go slow” adage, indicating that ramping up the dose of an antidepressant faster has a significantly positive effect on engagement in care for our population.

60 APPLIED LIFE SCIENCES↗

Forest aboveground biomass estimation through integration of sentinel-2 and PALSAR-2 time series: assessing models trained on GEDI and field inventory benchmarks

Accurate and spatially explicit forest Aboveground Biomass (AGB) mapping through remote sensing is critical for quantifying terrestrial carbon stocks and informing effective forest management strategies. However, AGB estimation in dense forests with complex terrain remains challenging due to satellite sensor signal saturation problem (saturation issue occurs in high biomass forests), structural complexity, and limited ground truth for calibration. This study presents a novel framework that integrates multi-temporal Sentinel-2 optical imagery, ALOS PALSAR-2 Synthetic Aperture Radar (SAR) data, and topographic variables with explainable Machine Learning to map AGB across mountainous forests within subtropical and temperate oceanic climate zones of Mexico. We evaluate the effects of temporal granularity and sensor synergy by comparing multiple temporal inputs and sensor configurations (Sentinel-2, PALSAR-2, and their fusion), and assess model performance using two reference datasets: NASA GEDI LiDAR-derived biomass and Mexico’s National Forest and Soil Inventory (INFyS). Our results showed that models trained on INFyS consistently outperformed those trained on GEDI, highlighting limitations in GEDI’s reliability in biomass estimates within this study region. Furthermore, the integration of Sentinel-2 and PALSAR-2 provided improved predictions compared to single-sensor models, particularly when combined with temporally explicit yearly statistics. The best-performing model, which was trained on INFyS data, and considered both Sentinel-2 and PALSAR-2 yearly statistics, as well as topographic variables, achieved an R2 of 0.64, RMSE of 51.10 Mg/ha, and relative RMSE (rRMSE) of 58.69%. Explainable ML analysis identified Sentinel-2 spectral indices and topographic features as key predictors, while PALSAR-2 metrics provided complementary information, partially mitigating saturation effects in high-biomass areas. Specifically, integrating both sensors substantially improved AGB estimation in high biomass forest (≥200 Mg/ha), yielding 98% gains over optical-only model, with resulting estimates exceeding GEDI L4B by 29% and ESA-CCI-BIOMASS by 174%. Terrain-stratified analysis indicated close agreement with GEDI in low-slope areas, with increasing divergence as slope steepness increased, while estimates remained consistently higher than ESA-CCI-BIOMASS across all slope classes. The proposed approach advances multi-sensor fusion and temporal feature engineering for AGB mapping using open-access satellite datasets, providing a scalable and reproducible framework for annual biomass monitoring in topographically complex mountainous forests. The resulting 25 m resolution biomass product has the potential to provide spatially detailed information for forest monitoring and may support applications in carbon accounting and forest management.

54 ENVIRONMENTAL SCIENCES↗

A Multi-Sensor Approach for Measuring Bird and Bat Collisions with Offshore Wind Turbines (Final Technical Report)

Collision of birds and bats with wind turbines is a conservation concern for both land-based and offshore wind projects. The fatality rates of birds and bats at land-based turbines are well documented. The measurement strategies on land focus on finding carcasses following collision, estimating the number of carcasses missed through searcher efficiency, carcass persistence trials and carcass fall distributions, and modeling statistically robust fatality rates. Few technologies have been developed to monitor offshore bird and bat collisions, and many that have been developed focused on detecting collisions with large birds. The few studies that have attempted to document collisions at offshore turbines do not account for smaller bodied animals or for collisions that might be missed, which prevents the calculation of statistically robust fatality rates. The overall goal of this report, A Multi-Sensor Approach for Measuring Bird and Bat Collisions with Offshore Wind Turbines (Project), was to develop an effective multi-sensor system for quantifying bird and bat collision rates, specifically for offshore wind facilities. The Project goal and resulting automated collision detection system was achieved through two major technological advancements: 1) refining The Netherlands Organisation for Applied Scientific Research’s (TNO’s) existing WT-Bird® vibration sensing system, that had successfully detected large bird collisions during daytime, to allow for improved detection of smaller birds and bats during both daytime and nighttime hours and 2) improving image processing systems and developing and integrating machine learning algorithms to automatically detect and classify small and large bird and bat collisions with offshore turbines. This final technical report (FTR) summarizes Methods , Results , Conclusions , and Lessons Learned during each of the five Tasks identified for this research and development effort. This FTR includes summaries of the following: Task 1. Initial Engineering Tests to Improve WT-Bird® Task 2. Installation of WT‐Bird® on a Utility-scale Turbine at the National Wind Technology Center – National Renewable Energy Laboratory Task 3. Field Tests and Refinement of the Object Detection System Task 4. Validation of WT-Bird® on a Land-based Turbine Task 5. Preparation for the Implementation of WT-Bird® on an Offshore Turbine. This research and development effort documented successful improvement of the WT Bird® collision detection system to detect small birds and bats, and WT-Bird® is the first collision detection system to validate results compared to land-based post-construction monitoring. The collision trials provide estimates of missed targets that can be used to estimate fatality rates, a significant improvement relative to other offshore collision monitoring systems. Advances were made in developing an edge-processing solution to reduce data storage requirements, which is important if the system is deployed for long periods of time at offshore turbines. The improved WT-Bird® system also provides an important option for wind operators on land or offshore who need to document specific details about when collisions occur, particularly efforts to further research on bat impact minimization, or when standard fatality searches are impractical (e.g. offshore) or inadequate (e.g. challenging locations on land).

17 WIND ENERGY↗

Modeling Protein–Protein and Protein–Ligand Interactions by the ClusPro Team in CASP16

ABSTRACT In the CASP16 experiment, our team employed hybrid computational strategies to predict both protein–protein and protein–ligand complex structures. For protein–protein docking, we combined physics‐based sampling—using ClusPro FFT docking and molecular dynamics—with AlphaFold (AF)‐based sampling, followed by AF‐based refinement. Our method produced numerous high‐accuracy complex models, including cases where AF alone failed, underscoring the critical role of physics‐based sampling alongside deep learning‐based refinement. For protein–ligand docking, we integrated the ClusPro LigTBM template‐based approach with a machine learning‐based confidence model for rescoring. The method preserves conserved interaction fragments derived from homologous complexes, followed by local resampling using physics‐based sampling and a diffusion model. Our template‐based strategy achieved a mean lDDT‐PLI of 0.69 across 233 targets, which was highly competitive. These results demonstrate that combining physics‐based modeling with AI‐driven refinement can significantly enhance the accuracy of both protein–protein and protein–ligand structure predictions.

Ashizawa, Ryota [Department of Applied Mathematics↗

Machine Learning Approaches to Predicting Induced Seismicity and Imaging Geothermal Reservoir Properties

This project developed machine learning (ML) methods, lab data sets, and field data to advance geothermal exploration and geothermal energy production. The work had three focus areas. One involved the development of ML methods to use microearthquakes (MEQs) for imaging geothermal reservoir properties and improving subsurface characterization – most importantly the evolution of permeability within the evolving reservoir. This part of the work included development of ML approaches for automated MEQ location, focal mechanism determination and identification of earthquake precursors. The second area focused on using MEQ signals generated by geothermal exploration and production to predict the relationship between fluid injection and seismicity. Here, we extended to reservoir scale our success in using ML to predict laboratory earthquakes and fault zone stress state. The third focus area was on lab experiments. Here, we developed new ML models for lab earthquake prediction and identification of precursors to failure to improve earthquake forecasting and early warning in geothermal settings. Major outcomes of our work include ML models that learn from MEQ signals during geothermal exploration and production to predict induced seismicity. MEQs occur naturally in connection with drilling and energy production. We developed ML methods to use the seismic waves from these events to characterize the elastic, hydraulic and poromechanical properties of reservoirs. Our work illuminated fracture geometry and the evolution of fracture permeability by incorporating seismic coda wave analysis and ML methods to relate fluid injection and seismicity. We significantly expanded laboratory earthquake prediction to include methods that use both passive measurements of microearthquakes within the lab fault zones and also active source acoustic measurements of fault zone elastic properties. These methods can now predict fault zone stress state, time to failure and the magnitude of lab earthquakes. Our work showed that repetitive stick- slip failure events during frictional sliding (the lab equivalent of earthquakes) are preceded by a cascade of micro-failure events that radiate energy in a manner that foretells unstable failure – manifest as laboratory MEQs. We documented a mapping between fracture properties and statistical attributes of elastic radiation. We extended existing works to geothermal reservoir scale and developed ML methods to determine reservoir permeability, fracture properties, and their evolution during geothermal energy production. An attractive feature of ML algorithms is their ability to handle big datasets and reveal patterns and correlations that may remain invisible to conventional analyses. Our work connected data from field, laboratory and intermediate scales to study permeability, stress, strength, fracture stiffness and geometry. At the field scale we used data from the Newberry Volcano field site, UtahFORGE, EGS Collab, and also the Bedretto underground research lab in Switzerland. These data sets are bridging the gap between the lab scale, theory, and reservoir scale. Our work produced plain language summaries to improve public understanding of DOE research. We also developed openly distributed ML and seismicity datasets for use by all researchers and we published connections between induced seismicity in geothermal areas and reservoir properties including permeability, fracture properties, and stress state. Our models are designed for the large data sets of induced seismicity typically associated with geothermal sites. We produced labeled event catalogs and used them on geothermal data to assess how ML can facilitate geothermal production and exploration. All datasets are available on the GDR Productivity: The project produced 32 publications in peer reviewed journals (two are in review). It supported the work of 6 PhD students, 40 conference presentations, 6 keynote talks at national meetings, and mentoring and professional development for 4 postdoctoral fellows.

15 GEOTHERMAL ENERGY↗

Improvement and generalization of ABCD method with Bayesian inference

To find New Physics or to refine our knowledge of the Standard Model at the LHC is an enterprise that involves many factors, such as the capabilities and the performance of the accelerator and detectors, the use and exploitation of the available information, the design of search strategies and observables, as well as the proposal of new models. We focus on the use of the information and pour our effort in re-thinking the usual data-driven ABCD method to improve it and to generalize it using Bayesian Machine Learning techniques and tools. We propose that a dataset consisting of a signal and many backgrounds is well described through a mixture model. Signal, backgrounds and their relative fractions in the sample can be well extracted by exploiting the prior knowledge and the dependence between the different observables at the event-by-event level with Bayesian tools. We show how, in contrast to the ABCD method, one can take advantage of understanding some properties of the different backgrounds and of having more than two independent observables to measure in each event. In addition, instead of regions defined through hard cuts, the Bayesian framework uses the information of continuous distribution to obtain soft-assignments of the events which are statistically more robust. To compare both methods we use a toy problem inspired by pp\to hh\to b\bar b b \bar b p p → h h → b b ‾ b b ‾ , selecting a reduced and simplified number of processes and analysing the flavor of the four jets and the invariant mass of the jet-pairs, modeled with simplified distributions. Taking advantage of all this information, and starting from a combination of biased and agnostic priors, leads us to a very good posterior once we use the Bayesian framework to exploit the data and the mutual information of the observables at the event-by-event level. We show how, in this simplified model, the Bayesian framework outperforms the ABCD method sensitivity in obtaining the signal fraction in scenarios with 1% and 0.5% true signal fractions in the dataset. We also show that the method is robust against the absence of signal. We discuss potential prospects for taking this Bayesian data-driven paradigm into more realistic scenarios.

Alvarez, Ezequiel↗

Basic Research Needs for Inverse Methods for Complex Systems under Uncertainty

Inverse problems, which aim to infer unknown properties of a system using experimental and observational data, are central to addressing many of the U.S. Department of Energy’s (DOE) most critical scientific and engineering challenges. Accurate, computationally efficient, and data-efficient solutions to inverse problems are essential for advancing DOE mission-critical science drivers, including analyzing data from large-scale experimental facilities, optimizing fusion reactor performance, accelerating materials discovery, enhancing geophysical imaging, improving wildfire predictions, and enabling autonomous systems and digital twins. However, these problems are becoming increasingly complex, often involving nonlinear, highdimensional, and interconnected systems and models that span multiple physics and scales, while relying on data with varying quantity, quality, and information content. Compounding these challenges is the uncertainty inherent in DOE-relevant systems, where errors in inputs, noise in data, incompleteness of data, and discrepancies between models and reality constrain the accuracy and precision of solutions. At the same time, the convergence of recent scientific computing trends—scientific machine learning, artificial intelligence, and computing advances such as exascale computing—is creating unprecedented opportunities for tackling these challenges. The cross-cutting nature of inverse problems, combined with their growing complexity and rapidly evolving data and algorithmic demands, strongly motivates the formulation of a prioritized research agenda to maximize their capabilities and impact. In response to this need, DOE’s Advanced Scientific Computing Research (ASCR) program in the Office of Science convened the Workshop on Basic Research Needs for Inverse Problems for Complex Systems Under Uncertainty in June 2025. This workshop brought together experts across disciplines to identify grand challenges and major opportunities in the field. Through collaborative discussions, the workshop defined transformative research directions aimed at addressing the mathematical, statistical, and computational challenges posed by inverse problems under uncertainty. As a result of these efforts, four priority research directions (PRDs) were identified to guide future research and development in this area. These PRDs, summarized below, represent a roadmap for advancing the foundational science and mathematics of inverse problems, enabling robust, scalable, and uncertainty-aware solutions that are critical for DOE applications.

97 MATHEMATICS AND COMPUTING↗