AI Surrogate Model for Distributed Computing Workloads
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Abstract This tutorial review provides a comprehensive overview of machine learning (ML)-based model predictive control (MPC) methods, covering both theoretical and practical aspects. It provides a theoretical analysis of closed-loop stability based on the generalization error of ML models and addresses practical challenges such as data scarcity, data quality, the curse of dimensionality, model uncertainty, computational efficiency, and safety from both modeling and control perspectives. The application of these methods is demonstrated using a nonlinear chemical process example, with open-source code available on GitHub. The paper concludes with a discussion on future research directions in ML-based MPC.
Plasma pyrolysis of methane is an effective route for zero-carbon hydrogen production. Yet, soot generated from pyrolysis of hydrocarbons is detrimental to the climate and human health. There is ample experimental and theoretical evidence that suggests polycyclic aromatic hydrocarbons (PAHs) are the molecular precursors to soot particles. The reaction pathways of PAH formation are intricately dependent on a multitude of process parameters, whose kinetic mechanisms are not well-understood in plasma pyrolysis. This project aims to leverage advances in the kinetic modeling of soot formation in combustion, as well as in surrogate modeling and active learning, to systematically investigate the effects of process parameter on the kinetics of PAH formation in plasma pyrolysis of methane. To this end, we propose to use the PAH formation kinetics model developed by the PPPL/PU group based on the well-established ABF and HACA mechanisms, coupled with low-temperature plasma models. We will develop an active learning (AL) framework based on Bayesian optimization to systematically and data-efficiently explore the complex and multivariable parameter space of plasma pyrolysis in order to quantify the effects of plasma and feed parameters on the ABF and HACA kinetic pathways. AL is the branch of machine learning concerned with systematically querying samples from a system (experimental or computational) to train a data-driven model that maps design parameters to a performance criterion. We will use the data generated via AL to perform global sensitivity analysis, combined with uncertainty quantification, to elucidate the impact of different reaction pathways on minimizing formation of soot precursors. This study will result in an improved understanding of kinetics of PAH formation in plasma pyrolysis and can pave the way for more advanced mechanistic studies (e.g., soot nucleation mechanisms). Additionally, the findings will be useful for establishing practical strategies for increasing the pyrolysis efficiency and producing high-grade carbon for synthesis of nanomaterials.
In current nuclear power plants (NPPs) a large amount of condition-based data which can be used to assess and monitor component health and performance. Assessing component health from such data can be performed with a large variety of methods. While the analysis of numeric data can be performed with several methods, the extraction of information from textual data remains a challenge. Currently employed natural language processing (NLP) methods do not really provide quantitative information that might be contained in IRs. In addition, the integration of numeric and textual data to identify possible causal relationships between data elements is still an unresolved challenge. This paper presents an approach to extract information from textual (e.g., incident or maintenance reports) and numeric data that relies on model based system engineer (MBSE) models. MBSE are diagrams designed to represent system and component dependencies (from both a form and functional point of view). In our approach, MBSE models emulate system engineer knowledge about component/system architecture. NLP methods are employed to perform syntactic and semantic analyses. Syntactic analysis analyzes the grammatical structure of a sentence while semantic analysis is designed to analyze the logic structure of a sentence. An innovative element of our approach is that semantic analysis uses MBSE models to identify links between textual elements. Similarly, numeric data is directly linked to elements of the MBSE models in order to map which functions are being monitored.
Surface precipitation measurements are essential for Earth system model (ESM) evaluation and understanding cloud processes. An ever-growing need for robust, temporally evolving, and easy-to-use statistical datasets provides motivation for a baseline ground-based precipitation properties data product. The U.S. Department of Energy Atmospheric Radiation Measurement (ARM) user facility operates an extensive suite of precipitation instruments with various sensitivities and operating mechanisms, which render the decision of which instrument to use based on one or more fixed thresholds challenging and prone to errors and bias. Using a long-term instrument inter-comparison from a unique per-precipitation event perspective, rather than instantaneous sample comparison, we demonstrate that ARM rainfall-measuring instruments are generally consistent with each other at the statistical level. Inter-instrument deviations at the single event level can be large, especially for specific rainfall event properties such as maximum precipitation rates. A machine-learning (ML) analysis using a random forest regressor indicates that in some cases, depending on instrument, local site climatology, and/or specific deployment configuration, certain atmospheric state variables influence the measured quantities in an unpredictable manner. Thus, a-priori weighting of different instruments does not necessarily lead to more accurate and less biased synthesis of instrument data. These results motivate the design of the ARM precipitation best-estimate (PrecipBE) value-added product, which incorporates all valid precipitation data while considering data quality and other instrument limitations. PrecipBE consists of time series and tabular statistics datasets in an easy-to-use and insightful per-precipitation event format. It provides a large set of precipitation event properties supplemented with ancillary data from ARM datasets that correspond to the detected precipitation events. We describe the PrecipBE algorithm and demonstrate its use via the examination of a single-day output as well as a long-term trend analysis of precipitation events at the ARM Southern Great Plains (SGP) site, covering more than 30 years of data. The trend analysis tentatively suggests a long-term temporal tendency for mainly shorter and less intense precipitation events at the SGP site, but a long-term increase in annual rainfall by more than 36 mm (5 %) per decade. This rainfall trend is catalyzed primarily by more extreme event properties of relatively rare, intense precipitation events, with event total and 1 min maximum precipitation rate at a 1 year timeframe increasing up to 5 mm and 9 mm h −1 (several percent) per decade, respectively. While the currently available PrecipBE datasets (at https://adc.arm.gov/discovery/, last access: 8 December 2025) cover rainfall from multiple ARM deployments up to March 2025, PrecipBE is planned to be expanded to include solid-phase precipitation and will soon become an operational product with a several-day lag from real-time. We invite the ARM user community to leverage this new product and welcome user feedback to further enhance the dataset.
NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “Fire Severity-Dependent Shifts in Vegetation Parameter Sensitivity: A Pre- and Post-Fire Analysis Using ELM-FATES and Explainable AI” submitted to Journal of Advances in Modeling Earth Systems (Zahura et al. 2026). The study examines vegetation physiological parameters controlling pre-fire and post-fire vegetation dynamics. To support this analysis, 73 vegetation parameters in Functionally Assembled Terrestrial Ecosystem Simulator (FATES) (Fisher et al., 2018) , which is coupled with E3SM (Energy Exascale Earth System Model) land model (ELM, ELM-FATES), were perturbed using a Sobol sequence to generate 1,024 ensemble members for two plant functional types: needleleaf evergreen extratropical trees (NEET) and C3 grass. Simulations were conducted for the pre-fire period (2016) and post-fire period (2018–2023). Burn severity was represented by modifying the Nesterov index in FATES to 75,000, 150,000, and 300,000 for low, moderate, and high severity, respectively. A no-fire scenario was also included. Simulations were performed for 16 grid cells in the American River Watershed across different burn severities and plant functional types. XGBoost (eXtreme Gradient Boosting) models were trained using the parameter ensembles and ELM-FATES-simulated outputs, including leaf area index (LAI), gross primary productivity (GPP), aboveground biomass, vegetation evaporation, transpiration, and soil evaporation. Models were trained separately for each year and burn severity, followed by SHAP (SHapley Additive exPlanations) analysis to identify changes in dominant parameters after fire disturbance. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. The data package contains the ELM-FATES simulation data. The scripts and data related to the analysis will be added later. The inputs and outputs from ELM-FATES are inside the “FATES” folder. “FATES_domain_surface” contains the domain and surface netcdfs that were used to run ELM-FATES in the study area. “FATES_parameters” contains the 1024 ensembles that were generated using Sobol sequence. “FATES_outputs” folder contains ELM-FATES simulated variables. All files are .csv and .nc (NetCDF).
Constructing a chemically diverse dataset while avoiding sampling bias is critical to training efficient and generalizable force fields. However, in computational chemistry and materials science, many common dataset generation techniques are prone to oversampling regions of the potential energy surface. Furthermore, these regions can be difficult to identify and isolate from each other or may not align well with human intuition, making it challenging to systematically remove bias in the dataset. While traditional clustering and pruning (down-sampling) approaches can be useful for this, they can often lead to information loss or a failure to properly identify distinct regions of the potential energy surface due to difficulties associated with the high dimensionality of atomic descriptors. In this work, we introduce the Multi-kernel Edge Attention-based Graph Autoencoder (MEAGraph) model, an unsupervised approach for analyzing atomic datasets. MEAGraph combines multiple linear kernel transformations with attention-based message passing to capture geometric sensitivity and enable effective dataset pruning without relying on labels or extensive training. Demonstrated applications on niobium, tantalum, and iron datasets show that MEAGraph efficiently groups similar atomic environments, allowing for the use of basic pruning techniques for removing sampling bias. This approach provides an effective method for representation learning and clustering that can be used for data analysis, outlier detection, and dataset optimization.
The Global Precipitation Measurement (GPM) mission Validation Network (VN) framework leverages over 118 ground-based polarimetric Doppler radars to validate a large subset of precipitation measurements and retrievals from the GPM Dual-frequency Precipitation Radar (DPR). Recently, GPM DPR reflectivity profiles within the VN have been classified according to their convective regime using unsupervised machine learning techniques. The archetypal regimes are stratiform, convective, mixed stratiform-convective (e.g., transition regions), and “other” (e.g., peripheral regions of light precipitation). Subcategories within these four primary regimes vary according to the characteristic depth of included reflectivity profiles, resulting in 12 main GPM DPR precipitation profile categories. Polarimetry of ground-based Doppler radars in the VN offers additional insights into the types of precipitation, while pairs of radars positioned near each other enable retrieval of vertical winds via dual-Doppler analysis. Geometrically matched to the DPR reflectivity profiles in the GPM VN, these ground-based data and retrievals contribute more detailed characterization of the distinct kinematic and microphysical structures associated with each of the 12 DPR precipitation regimes. DPR reflectivity profiles linked with wind in the VN are restricted to GPM overpasses of proximal radar pairs that allow dual-Doppler analysis. Although a limited subset of DPR profiles in the VN are matched with vertical motion, agreement between the reflectivity structures paired with wind data and those of the greater DPR dataset in the VN suggest that estimates of vertical motion may be inferred in regions without ground-based measurements. We present a climatology of the 12 convective regimes identified within the DPR VN dataset as well as early efforts to estimate the kinematic and microphysical structures of precipitation profiles within the greater GPM DPR dataset by applying machine learning techniques. Precipitation data paired with global estimates of vertical winds from these efforts offer early insight to and support upcoming missions to retrieve convective mass flux, including the Investigation of Convective Updrafts (INCUS) in the Tropics and the global Atmosphere Observing System (AOS).
The application of spatially resolved mass spectrometry (MS) and MS imaging approaches for studying biomolecular processes in the kidney is rapidly growing. These powerful methods, which enable label-free and multiplexed detection of many molecular classes across omics domains (including metabolites, drugs, proteins and protein post-translational modifications), are beginning to reveal new molecular insights related to kidney health and disease. Further, the complexity of the kidney often necessitates multiple scales of analysis for interrogating biofluids, whole organs, functional tissue units, single cells and subcellular compartments. Various MS methods can generate omics data across these spatial domains and facilitate both basic science and pathological assessment of the kidney. Optimal processes related to sample preparation and handling for different MS applications are rapidly evolving. Emerging technology and methods, improvement of spatial resolution, broader molecular characterization, multimodal and multiomics approaches and the use of machine learning and artificial intelligence approaches promise to make these applications even more valuable in the field of nephology. Overall, spatially resolved MS and MS imaging methods have the potential to fill much of the omics gap in systems biology analysis of the kidney and provide functional outputs that cannot be obtained using genomics and transcriptomic methods.
A challenge in population ecology studies is identifying how to best group individuals into populations, especially when individual origin is unknown. Machine learning has improved upon traditional methods of identifying population structure and is more efficient at handling large, complex datasets. We demonstrate the applicability of a machine learning method to identify hierarchical population structure in an emerging pathogen, Coccidioides spp., the causative agent of Valley fever. We compared the network clusters to structure identified by traditional tools as a validation of the network performance. We used publicly available whole-genome data for 48 C. immitis and 102 C. posadasii, resulting in 168,211 genome-wide SNPs among the two species. The network analysis grouped samples into populations comparable to the literature for these species but also identified fine-scale geographic structure and travel-associated cases not reported thus far. Exploring different resolutions in the network made it easy to identify unique genotypes specific to California and possibly Nevada, as well as Phoenix- and Tucson-acquired infections in non-endemic areas, regardless of reported travel history. The present study provides a promising example of how a ML-based network analysis can improve our ability to understand pathogen ecology, group cases into populations and infer travel-associated infections.
The poster at the 15th Wind Wildlife Research Meeting discusses how to leverage the potential of real-time thermal-imaging methodologies in quantifying nocturnal bat activities at wind turbines, using 3D computer vision techniques within a deep learning framework. This innovation enables the automatic detection and classification of bats, birds, and insects in thermal-imaging videos captured at wind turbine sites, facilitating efficient and accurate data analysis for enhanced understanding and mitigation of bat-wind turbine interactions.
The AIACHNE (AI/ML Informed cAlifornium CHi Nuclear data Experiment) project aims at designing an experiment for the 252 Cf Prompt Fission Neutron Spectrum (PFNS) that explores systematic biases in an experimental database retrieved from the EXFOR databases. To that end, machine learning (ML) methods were applied to pint-point measurement features likely related to bias. From that information, we selected a feature that should be explored by the AIACHNE experiment. Measurement features are metadata encapsulating all pertinent information about the physical measurement and analysis techniques. Examples are, for instance, what neutron and fission detectors were used for the physical metadata, and what background reduction techniques were employed for analysis techniques. Such metadata were retrieved both from EXFOR entries as well as the literature of data sets described in detail in Reference 2 (at the end of the article).
We developed a new machine learning-based tool for extracting information from interferometry measurements: MIDWAZE (Modular Interferometry Direct Waveform AnalyZEr). This paper showcases MIDWAZE’s ability to extract an object’s velocity information from Photonic Doppler Velocimetry (PDV) data at near-human accuracy with little to no human intervention. MIDWAZE can extract velocities roughly 350 times as fast as a human analyst "rushing" to complete their extractions, with similar extraction accuracy. MIDWAZE’s most outstanding feature is that it operates directly in waveform/temporal space, freeing analysis from certain limitations imposed by traditional spectrogram-based approaches and opening the way to "phase aware" PDV analysis. MIDWAZE also has limited ability to discriminate between different solid objects, which we develop as a first step towards automated discrimination of different kinds of objects such as ejecta clouds.
This dissertation explores the development and deployment of machine learning approaches to address critical challenges in anomaly detection across two distinct domains: neural network security in federated learning settings and cavity behavior analysis in particle accelerator operations at Jefferson Lab in Newport News, Virginia. Anomaly detection identifies deviations from expected patterns, safeguarding systems in cybersecurity, industry, and research against malicious activities and failures. This dissertation demonstrates how our machine learning approaches enhance detection accuracy and efficiency in both neural network security and industrial applications. First, we investigate vulnerabilities in deep neural networks deployed in federated learning. Although federated learning preserves user privacy by training models locally, it remains vulnerable to backdoor attacks, in which malicious participants embed hidden triggers that induce targeted misbehavior. We propose a self-supervised contrastive learning framework to detect and mitigate such backdoor attacks. In our experiments, this method achieves higher detection accuracy and lower false positive rates than existing defenses, while operating without access to local model updates or original training data and thus preserving the privacy guarantees of the federated setting. Second, we address the operational reliability of superconducting radio-frequency (SRF) cavities at the Continuous Electron Beam Accelerator Facility (CEBAF). Our research leverages an unsupervised learning approach, combined with Principal Component Analysis (PCA) and k-means clustering, to identify anomalous behaviors in SRF cavities. Our method detects subtle anomalous behavior by analyzing SRF signal data. This knowledge allows for the early detection and resolution of potential faults, significantly improving the efficiency and reliability of operations. Third, we extend these insights to time-series anomaly detection more broadly. We design a contrastive-learning based model tailored to increasingly dynamic environments and academic research. This model improves detection accuracy in settings that require real-time monitoring and predictive maintenance. Our research underscores the broader applicability and impact of advanced machine learning techniques in anomaly detection. By extracting meaningful patterns from complex data, machine learning can significantly enhance security in distributed neural networks and improve the efficiency of particle accelerator operations. This dissertation serves as a stepping stone for future investigations into the vast possibilities of anomaly detection, inspiring further exploration and development of machine learning techniques in this field.
In situ biogas upgrading achieves CO 2 conversion to CH 4 via hydrogenotrophic methanogenesis; however, gas-liquid mass transfer constraints limit the upgrading performance. Recognizing that optimization studies often underrepresent the effects of total solids (TS) and organic loading rate (OLR), this study undertook a holistic, statistics driven assessment of operating conditions for in situ H 2 assisted biogas upgrading, centering the analysis on TS and OLR. A dataset of 31 studies was compiled and comprised 99 observations. A rigorous analytical framework was employed, combining data standardization, fixed- and random-effects (REML) weighted regressions with cluster-robust errors, stratified analyses, and machine learning. Mixed-effects meta regression indicated that TS was the main factor explaining differences of methane fraction (CH 4 %) when considering the between studies heterogeneity. Focusing on a near-stoichiometric subset (H 2 /CO 2 ≈ 4:1), TS remained significant. Stratified results showed a stronger negative relationship between TS and CH 4 % in UASB reactors than in CSTRs, with a negative effect under mesophilic conditions and no significant effect under thermophilic conditions. A Random Forest model corroborated the statistical findings, consistently ranking H 2 /CO 2 ratio, OLR, TS, and hydrogen injection rate (HIR) as the most influential predictors. These findings delineate trends across increasing TS levels, particularly between 1% and 10%, and provide preliminary insights for TS above 15% in in situ biogas upgrading. They further provide insights for the influence of TS by reactor type and temperature, thereby advancing the evidence base for implementing biological CO 2 conversion to CH 4 in practice.
An aspiration for EIMO datascope is to realize artificial intelligence-enhanced solutions for analysis of crew health & performance data and to facilitate clinical decision support for autonomous medical operations. A vision proposed to the meeting participants was that of a “system of systems,” whereby EIMO will utilize AI-supported natural language processing and machine learning techniques to synthesize embedded reference databases and real-time data streams [input vectors] from multiple data sources to continuously and seamlessly assess crew health & performance. Constituent input vectors may include environmental controls, countermeasures data, behavioral data, physiologic wearables, point-of-care laboratory tests, personalized medical records, inventory trade space risk assessments, COTS medical databases, and ground support inputs. An ideal AI capability would possess trained fusion algorithms to cross reference input vectors with medical ‘knowledge’ [cultivated database] to stratify relevant data streams for predictive and actionable capabilities. In addition, EIMO will ideally have a degree of mobility, in that it can be accessed and can push/pull data within and between multiple vehicles/habitats.
Arcjet Computer Vision (arcjetCV) has been significantly upgraded to enhance accuracy and performance in tracking material recession and shock-material standoff in test videos. These improvements include integrating new machine learning models, developing a specialized edge detection class, and incorporating a more comprehensive training dataset. These upgrades have refined the software’s ability to automate time-resolved recession tracking, making it more precise and reliable for analyzing complex physical processes. In parallel, a new tool called STARscan (Spatial Targeting and Alignment Rig for Scanning) is being developed to capture detailed 3D surface data before and after testing. By comparing these pre- and post-test scans with arcjetCV’s automated video analysis results, users can achieve a more comprehensive assessment of material recession. This method enables cross-validation of results, improving confidence in the analysis of tested materials. The expanded capabilities of arcjetCV have been successfully demonstrated on videos from various facilities, including the NASA Ames arcjets, UIUC’s PlasmatronX, and the VKI Plasmatron. It has been adopted as a new standard for in-situ recession tracking by the Mars Sample Return Project and Orion. ArcjetCV’s improved efficiency and accuracy are critical for reducing testing uncertainties and validating heatshield material performance under extreme conditions. The software’s user-friendly graphical interface ensures ease of use, enabling seamless processing and precise analysis of arcjet videos, providing deeper insights into material behavior in hypersonic environments. ArcjetCV is now available on both PyPI and Conda, allowing easy installation via "pip install arcjetCV" or through the Conda package manager, ensuring broad accessibility and streamlined deployment for users across various platforms.
There is an increasing interest in applying methods based on Machine Learning Techniques (MLT) to problems in Air Traffic Management (ATM). The current interest is based on developments in Cloud Computing, the availability of open software and the success of MLT in automation, consumer behavior and finance involving large databases. This paper reviews the current-state-of-the art in applying MLT to aviation operations, its promises and challenges. Historically aviation operations have been analyzed using physics-based models and provide information for making operational decisions. Aviation operations involving many decision makers, multiple objectives, poor or unavailable physics-based models and a rich historical database are prime candidates for analysis using data-driven methods. The promises and challenges in applying MLT to ATM is traced through three examples based on the authors’ experience, each separated by a decade, to show the influence of data and feature selection in the successful application of MLT to ATM. As always, the best approach depends on the task, the physical understanding of the problem and the quality and quantity of the available data.