Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Knowledge-guided learning with curated prior genetic biomarkers for robust model interpretation

Abstract Motivation Knowledge-guided learning offers effective and robust model training strategies in data-scarce settings by incorporating established domain knowledge, thereby enhancing generalization, robustness, and interpretability. By contrast, conventional deep learning approaches rely purely on data-driven learning, which can limit robust model interpretability, particularly in high-dimensional settings with limited size samples. In computational biology, knowledge-guided learning has primarily leveraged network- and structural-based knowledge, leading to biologically interpretable representations and enhanced predictive performance compared to conventional approaches. However, curated biomarkers, one of the most accessible forms of biological knowledge, remain largely unexplored within knowledge-guided paradigms. Results In this study, we propose a model-agnostic training paradigm, Biomarker-driven Explainable Prior-guided Learning (BioExPL), that can be applied to any neural networks that incorporates curated prior knowledge. BioExPL enforces neural networks to reflect curated biomarker priors in their latent representations through a novel knowledge-alignment loss. BioExPL consistently demonstrated significantly improved predictive performance and enhanced model interpretability with minimized computational overhead in simulation studies and intensive experiments on multiple cancer datasets. BioExPL not only integrates prior curated knowledge into the model but also accurately identifies unknown associated signals additionally. BioExPL is model-agnostic and domain-independent, enabling its integration into diverse neural network architectures. Availability and implementation The open-source is publicly available at: https://github.com/datax-lab/BioExPL.

Baek, Beomsu [Department of Computer Science, Univ↗

Deep-learning-based canopy height model generation from sub-meter resolution panchromatic satellite imagery

Canopy height models (CHMs) with sufficient resolution to distinguish individual trees are useful for a variety of applications. However, standard techniques to acquire such data, such as airborne lidar surveying, are often prohibitively expensive. Deep learning techniques for generating CHMs from high-resolution imagery are an attractive option to reduce costs. To date, success with these methods has been demonstrated using multichannel aerial photography and specialized satellite data products derived from multiple sensors, neither of which is commonly available at temporal resolutions finer than one year. Here we demonstrate a method to generate sub-meter resolution CHMs in three forests in California using a more abundant data source: sub-meter resolution, panchromatic satellite imagery from a single sensor. We show that phenology and species composition play important roles in model transferability; when trained using imagery from a single conifer forest in autumn, the model performs well on autumn imagery from a second conifer forest several hundred kilometers distant with no re-training. With modest additions to the training dataset, the same model generates minimally biased estimates of canopy height in both conifer and deciduous forests during multiple seasons. Because the model operates on satellite data with global coverage and a relatively short return interval, we propose its suitability to extrapolate tree-level canopy height data to remote regions and conduct high-temporal resolution monitoring of forest structure. We furthermore demonstrate the workflow’s applicability to fire modeling by conducting simulations in forests populated by trees measured using both this approach and airborne lidar surveying. We find minimal differences in fire behavior relative to a baseline case in which only statistical distributions of tree height and crown area are known. This result underscores the value of forest structural information derived from our workflow for improving the fidelity of wildland fire simulations, among other ecological applications.

54 ENVIRONMENTAL SCIENCES↗

Wavelet flow for extragalactic foreground simulations

Extragalactic foregrounds in cosmic microwave background (CMB) observations are both a source of cosmological and astrophysical information and a nuisance to the CMB. Effective field-level modeling that captures their non-Gaussian statistical distributions is increasingly important for optimal information extraction, particularly given the low-noise observations from current and upcoming experiments. Here, we explore the use of Wavelet Flow (WF) models to tackle the novel task of modeling the field-level probability distributions of multi-component CMB secondaries and foregrounds. Specifically, we jointly train correlated CMB lensing convergence (κ) and cosmic infrared background (CIB) maps with a WF model and obtain a network that statistically recovers the input to high accuracy — the trained network generates samples of κ and CIB fields whose average power spectra are within a few percent of the inputs across all scales, and whose Minkowski functionals are similarly accurate compared to the inputs. Leveraging the multiscale architecture of these models, we fine-tune both the model parameters and the priors at each scale independently, optimizing performance across different resolutions. These results demonstrate that WF models can accurately simulate correlated components of CMB secondaries, supporting improved analysis of cosmological data. Our code and trained models can be found on this GitHub repo.

cosmological simulations↗

White paper on light sterile neutrino searches and related phenomenology

This white paper provides a comprehensive review of our present understanding of experimental neutrino anomalies that remain unresolved, charting the progress achieved over the last decade at the experimental and phenomenological level, and sets the stage for future programmatic prospects in addressing those anomalies. It is purposed to serve as a guiding and motivational "encyclopedic" reference, with emphasis on needs and options for future exploration that may lead to the ultimate resolution of the anomalies. We see the main experimental, analysis, and theory-driven thrusts that will be essential to achieving this goal being: 1) Cover all anomaly sectors -- given the unresolved nature of all four canonical anomalies, it is imperative to support all pillars of a diverse experimental portfolio, source, reactor, decay-at-rest, decay-in-flight, and other methods/sources, to provide complementary probes of and increased precision for new physics explanations; 2) Pursue diverse signatures -- it is imperative that experiments make design and analysis choices that maximize sensitivity to as broad an array of these potential new physics signatures as possible; 3) Deepen theoretical engagement -- priority in the theory community should be placed on development of standard and beyond standard models relevant to all four short-baseline anomalies and the development of tools for efficient tests of these models with existing and future experimental datasets; 4) Openly share data -- Fluid communication between the experimental and theory communities will be required, which implies that both experimental data releases and theoretical calculations should be publicly available; and 5) Apply robust analysis techniques -- Appropriate statistical treatment is crucial to assess the compatibility of data sets within the context of any given model.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Extended Fayans energy density functional: optimization and analysis

The Fayans energy density functional (EDF) has been very successful in describing global nuclear properties (binding energies, charge radii, and especially differences of radii) within nuclear density functional theory. In a recent study, supervised machine learning methods were used to calibrate the Fayans EDF. Building on this experience, in this work we explore the effect of adding isovector pairing terms, which are responsible for different proton and neutron pairing fields, by comparing a 13D model without the isovector pairing term against the extended 14D model. At the heart of the calibration is a carefully selected heterogeneous dataset of experimental observables representing ground-state properties of spherical even–even nuclei. To quantify the impact of the calibration dataset on model parameters and the importance of the new terms, we carry out advanced sensitivity and correlation analysis on both models. The extension to 14D improves the overall quality of the model by about 30%. The enhanced degrees of freedom of the 14D model reduce correlations between model parameters and enhance sensitivity.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Revealing the Hidden Third Dimension of Point Defects in Two-Dimensional MXenes

Point defects govern many important functional properties of two-dimensional (2D) materials. However, resolving the three-dimensional (3D) arrangement of these defects in multi-layer 2D materials remains a fundamental challenge, hindering rational defect engineering. Here, we overcome this limitation using an artificial intelligence-guided electron microscopy workflow to map the 3D topology and clustering of atomic vacancies in Ti3C2TX MXene. Our approach reconstructs the 3D coordinates of vacancies across hundreds of thousands of lattice sites, generating robust statistical insight into their distribution that can be correlated with specific synthesis pathways. This large-scale data enables us to classify a hierarchy of defect structures-from isolated vacancies to nanopores-revealing their preferred formation and interaction mechanisms, as corroborated by molecular dynamics simulations. This work provides a generalizable framework for understanding and ultimately controlling point defects across large volumes, paving the way for the rational design of defect-engineered functional 2D materials.

2D materials↗

The Dark Energy Survey: Cosmology Results with ∼1500 New High-redshift Type Ia Supernovae Using the Full 5 yr Data Set

Abstract We present cosmological constraints from the sample of Type Ia supernovae (SNe Ia) discovered and measured during the full 5 yr of the Dark Energy Survey (DES) SN program. In contrast to most previous cosmological samples, in which SNe are classified based on their spectra, we classify the DES SNe using a machine learning algorithm applied to their light curves in four photometric bands. Spectroscopic redshifts are acquired from a dedicated follow-up survey of the host galaxies. After accounting for the likelihood of each SN being an SN Ia, we find 1635 DES SNe in the redshift range 0.10 < z < 1.13 that pass quality selection criteria sufficient to constrain cosmological parameters. This quintuples the number of high-quality z > 0.5 SNe compared to the previous leading compilation of Pantheon+ and results in the tightest cosmological constraints achieved by any SN data set to date. To derive cosmological constraints, we combine the DES SN data with a high-quality external low-redshift sample consisting of 194 SNe Ia spanning 0.025 < z < 0.10. Using SN data alone and including systematic uncertainties, we find Ω M = 0.352 ± 0.017 in flat ΛCDM. SN data alone now require acceleration ( q 0 < 0 in ΛCDM) with over 5 σ confidence. We find ( Ω M , w ) = ( 0.264 − 0.096 + 0.074 , − 0.80 − 0.16 + 0.14 ) in flat w CDM. For flat w 0 w a CDM, we find ( Ω M , w 0 , w a ) = ( 0.495 − 0.043 + 0.033 , − 0.36 − 0.30 + 0.36 , − 8.8 − 4.5 + 3.7 ) , consistent with a constant equation of state to within ∼2 σ . Including Planck cosmic microwave background, Sloan Digital Sky Survey baryon acoustic oscillation, and DES 3 × 2pt data gives (Ω M , w ) = (0.321 ± 0.007, −0.941 ± 0.026). In all cases, dark energy is consistent with a cosmological constant to within ∼2 σ . Systematic errors on cosmological parameters are subdominant compared to statistical errors; these results thus pave the way for future photometrically classified SN analyses.

79 ASTRONOMY AND ASTROPHYSICS↗

Benchmarking large language models for materials synthesis: The case of atomic layer deposition

In this work, we introduce an open-ended question benchmark, ALDbench, to evaluate the performance of large language models (LLMs) in materials synthesis, and, in particular, in the field of atomic layer deposition, a thin film growth technique used in energy applications and microelectronics. Our benchmark comprises questions with a level of difficulty ranging from the graduate level to domain expert current with the state of the art in the field. Human experts reviewed the questions along the criteria of difficulty and specificity, and the model responses along four different criteria: overall quality, specificity, relevance, and accuracy. We ran this benchmark on an instance of OpenAI’s GPT-4o. The responses from the model received a composite quality score of 3.7 on a 1–5 scale, consistent with a passing grade. However, 36% of the questions received at least one below average score. An in-depth analysis of the responses identified at least five instances of suspected hallucination. Finally, we observed statistically significant correlations between the difficulty of the question and the quality of the response, the difficulty of the question and the relevance of the response, the specificity of the question, and the accuracy of the response as graded by the human experts. Furthermore, this emphasizes the need to evaluate LLMs across multiple criteria beyond difficulty or accuracy.

Artificial intelligence↗

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis - NLR Historical Wind

The U.S. Department of Energy and the National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis from variable sources, hydrogen compression and storage, and hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) initiative. This dataset represents part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure nationwide. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with various energy technologies. Future datasets will demonstrate how existing hydrogen fuel cell technologies can provide controllable, dispatchable, and variable power output for artificial intelligence (AI) data centers and other variable loads. This dataset entry describes hydrogen production by conducting a statistical analysis of historical wind data over a five-year period (2020-2025) from a single 1.5MW turbine manufactured by General Electric (GE) located at NLR’s Flatirons Campus, to generate an experimental test profile that was deployed on a 1.25-MW proton exchange membrane type MC250 electrolyzer system manufactured by Nel Hydrogen . [1] While the electrolyzer balance-of-plant supports up to 2.5 MW of electrolysis, NLR only has a single 1.25-MW electrolysis stack. The historical wind data provided several metrics, however, the analysis particularly focused on the measured power output by the wind turbine. The power output time series of data for each day was categorized by total energy generation and standard deviation, and the day that represented the highest combination of these two metrics was chosen – December 25th, 2022. This process was then repeated for a moving four-hour window within this day to identify the most statistically variable period. Finally, this four-hour period was scaled by 65% to match the 1.25 MW electrolyzer. The electrolysis system controls hydrogen production by varying DC current applied to the stack, from a maximum of 3000 A to a minimum safe operation of 300 A, or 10%. Because the current – voltage characteristic changes as the stack ages and efficiency degrades, the actual minimum safe operating power changes over time. The historical wind profiles were translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1 Hz frequency. For more details on the statistical analysis process, see the presentation labeled “ Public Reference Data for Megawatt-Scale Hydrogen Electrolysis” provided with each data entry. These datasets report relevant hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. Each .zip file represents a single wind turbine electrolysis experiment and is formatted as follows: {technology}_{scaling factor}-{electrolyzer ramp rate in amperes/second} For instance, “wind-GE1.5MW_0.65-400.zip” represents the hour-long experiment using historical data from the wind-GE1.5MW turbine, scaled to 65%, with the electrolyzer power supply set to a maximum ramp rate (gain and slew) of 400 A/s. Each .zip folder contains the following files: A .csv file containing raw data An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production, electrolysis power consumption, and wind power input. A PDF file detailing the historical wind data statistical analysis used to generate the wind profile. An experiment labeled “characterization_200.zip” demonstrates the MC250 electrolyzer steady-state response with 30-minute load steps for a total duration of 5 hours. Finally, a .csv file is provided with all simulated wind experiments combined into one dataset labeled "combined_historical_wind_experiments.csv". NLR also built an AI/machine-learning predictive model based on these datasets. The model ingests the electrolyzer current command in amperes, as well as various pressures and temperatures across the system, and predicts hydrogen output in kilograms per hour. The complete model can be found at https://huggingface.co/NatLabRockies/ptmelt-hydrogen-electrolysis [1] nelhydrogen.com/product/mc-series-electrolyser .

08 HYDROGEN↗

Using the optimal combined index weight ratio to improve the probability of anomaly detection in big area additive manufacturing

Big Area Additive Manufacturing (BAAM) of composites requires significant time, energy, and material, so it is critical to reduce production inefficiencies to make functional parts without multiple iterations. Statistical process control coupled with Principal Component Analysis (PCA) is a powerful technique that provides a quick, computationally inexpensive, and intuitive way for operators to detect defects that form in a manufacturing process without massive datasets. Recently, a combined index that is a weighted sum of the Hotelling's T 2 and squared residual error statistics has been proposed that can be monitored in one chart, improving interpretation accuracy and simplicity. However, the literature does not offer a formal method to optimise the weights. Here, we introduce two new approaches to the traditional weight selection approach using simulated and BAAM image data. Approach 1 uses a theoretically motivated optimum inspired by probabilistic principal component analysis. Approach 2 systematically varies the ratio of the weights to find the optimum. We show that approach 1 delivers optimal anomaly detection performance in select cases while approach 2 fares better in practice. Surprisingly, we also show that choosing a more complex PCA model has a minimal negative impact on anomaly detection performance compared to a more simplistic model.

3-dimensional printing↗

RADAI: A Large-Scale Realistic Dataset for Radiation Detection Algorithm Development

Open, realistic datasets are essential for developing and benchmarking radiation detection algorithms, yet they remain scarce. The Radiological Anomaly Detection and Identification (RADAI) project was develop to create datasets that meet the training and testing needs for sophisticated radiation detection algorithms. The RADAI dataset is a large-scale synthetic resource that integrates high-fidelity Monte Carlo simulations with realistic urban scenarios to capture both background variability and source signatures. RADAI models construction-material NORM, people and vehicles, urban clutter, and dynamic environmental effects such as cosmic-ray and rain-induced transients, and they provide list-mode detector data with motion and response modeling suitable for algorithm training and evaluation. The RADAI project resulted in three publicly-released complementary datasets together with an online scoring portal for standardized performance assessment and an open software toolkit that supports data access, augmentation, model development, and evaluation. These resources enable reproducible comparisons across methods and promote rigorous studies at the scale required by contemporary machine learning. By grounding algorithm development in realistic, well-documented conditions, RADAI supports progress toward more robust detection, identification, and localization in complex urban environments.

Ghawaly, James M. [Division of Computer Science an↗

A Field Guide to Corralling the Chaos: A Conceptual Framework for Using Models to Guide Opportunistic Field Studies of Natural Disturbances

Watersheds regulate biogeochemical processes and provide ecosystem services to human societies, but disturbances can fundamentally alter these processes across space and time. Determining when and where to sample to capture disturbance impacts in watersheds remains a central challenge. Manipulation studies and long-term monitoring are often constrained by scope, and opportunistic studies often lack pre-disturbance data needed to statistically determine disturbance impacts. We identify a persistent knowledge gap: the absence of a clear, transferable framework to guide opportunistic disturbance research where pre-disturbance data collection is not a feasible option. To address this gap, we present a conceptual framework that intentionally integrates modeling and empirical observation in an iterative, stepwise model–experiment workflow. We demonstrate its application through two contrasting case studies: wildfire impacts on headwater streams using a pre-disturbance preparedness approach, and saltwater flooding impacts on coastal forests using an ‘ex-post-facto’ approach. From these applications, we assess strengths, limitations, and the critical role of team science for transferability across disturbance types and study designs. Broadly, this framework offers a scalable path towards more rigorous, timely, and actionable disturbance science that can inform watershed management, hazard risk reduction, and ecosystem resilience.

Coastal Biogeochemistry↗

Explainable tokamak-agnostic forecasting of fusion plasma instability via megahertz turbulent fluctuations

Scientific applications of artificial intelligence (AI) often remain limited by device-specific training and unexplained “black-box” approaches, creating fundamental barriers to cross-system generalization. This challenge is critical for nuclear fusion, where future reactors will have limited operational data for AI training. Here, we demonstrate that our neural network, trained solely on megahertz-scale turbulence measurements from one machine (DIII-D), forecasts Type-I edge localized mode (ELM) onsets in a different tokamak (KSTAR) through zero-shot weight transfer following physics-consistent preprocessing without device-specific retraining. Through an explainable AI framework combining gradient-weighted class activation mapping with physics validation, we reveal that our network can internalize physics relationships governing the ELM instabilities rather than memorizing device-specific patterns. The network perceives spatiotemporal features that correlate consistently with independently calculated instability growth rates, magnetohydrodynamic stability limits, and pedestal structure dynamics. Statistical analyses of dimensionally-reduced saliency features reveal the identical triangular features between the saliency representations, instability growth rates, and prediction probability across tokamaks, providing evidence that our forecasting system can show tokamak-agnostic generalization. This work contributes to a foundation for explainable scientific AI systems, where cross-system developments are essential for transcending traditional domain-specific constraints.

AI↗

TPCpp-10M: Simulated proton-proton collisions in a time projection chamber for AI foundation models

Scientific foundation models hold great promise for advancing nuclear and particle physics by improving analysis precision and accelerating discovery. Yet, progress in this field is often limited by the lack of openly available large scale datasets, as well as standardized evaluation tasks and metrics. Furthermore, the specialized knowledge and software typically required to process particle physics data pose significant barriers to interdisciplinary collaboration with the broader machine learning community. This work introduces a large, openly accessible dataset of 10 million simulated proton-proton collisions, designed to support self-supervised training of foundation models. To facilitate ease of use, the dataset is provided in a common NumPy format. In addition, it includes 70,000 labeled examples spanning three well defined downstream tasks: track finding, particle identification, and noise tagging, to enable systematic evaluation of the foundation model's adaptability. The simulated data are generated using the Pythia Monte Carlo event generator at a center of mass energy of $\sqrt{s}$ = 200 GeV and processed with Geant4 to include realistic detector conditions and signal emulation in the sPHENIX Time Projection Chamber at the Relativistic Heavy Ion Collider, located at Brookhaven National Laboratory. This dataset resource establishes a common ground for interdisciplinary research, enabling machine learning scientists and physicists alike to explore scaling behaviors, assess transferability, and accelerate progress toward foundation models in nuclear and high energy physics. The complete simulation and reconstruction chain is reproducible with the sPHENIX software stack. All data and code locations are provided under Data Accessibility.

Data Analysis, Statistics and Probability (physics↗

Risk-Aware Framework Development for Disruption Prediction: Alcator C-Mod and DIII-D Survival Analysis

Abstract Survival regression models can achieve longer warning times at similar receiver operating characteristic performance than previously investigated models. Survival regression models are also shown to predict the time until a disruption will occur with lower error than other predictors. Time-to-event predictions from time-series data can be obtained with a survival analysis statistical framework, and there have been many tools developed for this task which we aim to apply to disruption prediction. Using the open-source Auton-Survival package we have implemented disruption predictors with the survival regression models Cox Proportional Hazards, Deep Cox Proportional Hazards, and Deep Survival Machines. To compare with previous work, we also include predictors using a Random Forest binary classifier, and a conditional Kaplan-Meier formalism. We benchmarked the performance of these five predictors using experimental data from the Alcator C-Mod and DIII-D tokamaks by simulating alarms on each individual shot. We find that developing machine-relevant metrics to evaluate models is an important area for future work. While this study finds cases where disruptive conditions are not predicted, there are instances where the desired outcome is produced. Giving the plasma control system the expected time-to-disruption will allow it to determine the optimal actuator response in real time to minimize risk of damage to the device.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Efficient and generalizable nested Fourier-DeepONet for three-dimensional geological carbon sequestration

Geological carbon sequestration (GCS) involves injecting CO2 into subsurface geological formationsfor permanent storage. Numerical simulations could guide decisions in GCS projects by predictingCO 2 migration pathways and the pressure distribution in storage formation. However, these simula-tions are often computationally expensive due to highly coupled physics and large spatial-temporalsimulation domains. Surrogate modelling with data-driven machine learning has become a promis-ing alternative to accelerate physics-based simulations. Among these, the Fourier neural operator(FNO) has been applied to three-dimensional synthetic subsurface models. Despite its good accuracyin simulating CO 2 plume migration, it requires large computational resources in training and alsolacks generalizability. Here, to further improve performance, we have developed a nested Fourier-DeepONet by combining the expressiveness of the FNO with the modularity of a deep operatornetwork (DeepONet). This new framework is twice as efficient as a nested FNO for training and has atleast 80% lower GPU memory requirement due to its flexibility to treat temporal coordinates sepa-rately. These performance improvements are achieved without compromising prediction accuracy.In addition, the generalization and extrapolation ability of nested Fourier-DeepONet beyond thetraining range has been thoroughly evaluated. Nested Fourier-DeepONet outperformed the nestedFNO for extrapolation in time with more than 50% reduced error. It also exhibited good extrapolationaccuracy beyond the training range in terms of reservoir properties, number of wells, and injectionrate.

Lee, Jonathan E. [Department of Chemical and Envir↗

Unsupervised anomaly detection in MeV ultrafast electron diffraction

MeV ultrafast electron diffraction (MUED) is a pump-probe technique used to study the dynamic structural evolution of materials. An ultrashort laser pulse triggers structural changes, which are then probed by an ultrashort relativistic electron beam. To overcome low signal-to-noise ratios, diffraction patterns are averaged over thousands of shots. However, shot-to-shot instabilities in the electron beam can distort individual patterns, introducing uncertainty. Improving MUED accuracy requires detecting and removing these anomalous patterns from large datasets. In this work, we developed a fully unsupervised methodology for the detection of anomalous diffraction patterns. Using a convolutional autoencoder, we calculate the reconstruction mean squared error of the diffraction patterns. Based on the statistical analysis of this error, we provide the user an estimation of the probability that the pattern is normal, which also allows a posterior visual inspection of the images that are difficult to classify. This method has been trained with only 100 diffraction patterns and tested on 1521 patterns, resulting in a false positive rate between 0.2% and 0.4%, with a training time of 10 s per image and a test time of about 1 s per image. Here, the proposed methodology can also be applied to other diffraction techniques in which large datasets are collected that include faulty images due to instrumental instabilities.

43 PARTICLE ACCELERATORS↗

Ripening of Rh Nanoparticle Catalysts in Reverse Water–Gas Shift via a Data-Driven Model Combining Physics, Theory, and Experiment

Degradation via sintering is an ongoing challenge that impedes the broad commercial success of supported metallic nanoparticle catalysts. To mitigate degradation via informed catalyst design and process operations, here we aim to disambiguate the underlying mechanisms of sintering by combining theory and experiment in a quantitative framework. While mechanistic sintering models exist, they only model a single sintering pathway, even though multiple sintering mechanisms can occur simultaneously or dominate at different stages of the process. Data-driven machine learning models have emerged as a means to represent complex processes through data regression. However, machine learning models have very large data needs and lack mechanistic insights due to their black-box encoding. To develop an interpretive model of catalyst degradation via sintering, we constructed a hybrid model combining mechanistic “physics-based” models and data-driven methods to obtain both reliable predictions and mechanistic insights regarding experimentally observed sintering phenomena. Focusing on nanoparticle sintering in the Rh–TiO 2 catalyst for the reverse water–gas shift (RWGS) reaction, the hybrid model couples a mechanistic term for Ostwald ripening with energy values calculated via density functional theory (DFT) with a parametric, data-driven discrepancy function term for unmodeled mechanisms. The hybrid model is trained using Bayesian inference with data collected from small-angle X-ray scattering (SAXS) in situ experiments wherein average nanoparticle diameter versus time was measured at three relevant operating temperatures. The calibrated hybrid model results show that an Ostwald ripening-only model parameterized with fixed DFT energies does not fully capture the time and temperature dependence of the SAXS-observed sintering kinetics, and that an additional functional contribution, or DFT energy calibration, is required to reconcile simulation and experiment. Analysis of the hybrid-model error confirms that the hybrid model outperforms both the purely mechanistic and purely data-driven alternatives in terms of expected predictive accuracy for time-evolving average particle sizes. Furthermore, the results support the hypothesis that the Ostwald ripening mechanism is less important for explaining the sintering phenomena as operating temperature increases under an assumed fixed DFT parameterization. This could be explained in one of two ways: either latent, unmodeled sintering mechanisms dominate at higher temperatures, or the DFT uncertainty increases with temperature. The proposed modeling approach directly links theory to experiments and simulations via a statistical hybrid modeling framework and can be extended to other catalytic systems to improve predictive models and mechanistic understanding.

Bayesian hybrid modeling↗