Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Inference”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27

ML–Enabled FPGA Framework for Fast Quantum State Discrimination in Mid-Circuit Measurement Regimes

Accurate and low-latency quantum state discrimination is essential for protocols involving mid-circuit measurement (MCM) and conditional feed-forward. In superconducting quantum systems, conventional readout pipelines transfer measurement data to host processors for post-processing, introducing millisecond-scale delays that far exceed qubit coherence times. To overcome this bottleneck, we present an in-situ machine learning (ML) inference engine implemented on an FPGA for real-time quantum state discrimination. Our design performs inference directly on digitized readout signals with 40 ns latency, supports both qubit and qutrit readout, and enables conditional operations without host-side intervention. This capability is critical for MCM and for feedback-driven protocols such as quantum error correction. We validate the system on superconducting transmon hardware, demonstrating robust discrimination fidelity across multiple qubit and qutrit channels. We further demonstrate conditional qutrit logic driven by FPGA-resident classification, highlighting the potential of low-latency ML-on-FPGA control for NISQ applications and scalable fault-tolerant quantum computing.

Vora, Neel [Lawrence Berkeley National Laboratory ↗

Constraining Cosmology with Simulation-based inference and Optical Galaxy Cluster Abundance

We test the robustness of simulation-based inference (SBI) in the context of cosmological parameter estimation from galaxy cluster counts and masses in simulated optical datasets. We construct ``simulations'' using analytical models for the galaxy cluster halo mass function (HMF) and for the observed richness (number of observed member galaxies) to train and test the SBI method. We compare the SBI parameter posterior samples to those from an MCMC analysis that uses the same analytical models to construct predictions of the observed data vector. The two methods exhibit comparable performance, with reliable constraints derived for the primary cosmological parameters, ($\Omega_m$ and $\sigma_8$), and richness-mass relation parameters. We also perform out-of-domain tests with observables constructed from galaxy cluster-sized halos in the Quijote simulations. Again, the SBI and MCMC results have comparable posteriors, with similar uncertainties and biases. Unsurprisingly, upon evaluating the SBI method on thousands of simulated data vectors that span the parameter space, SBI exhibits worsened posterior calibration metrics in the out-of-domain application. We note that such calibration tests with MCMC is less computationally feasible and highlight the potential use of SBI to stress-test limitations of analytical models, such as in the use for constructing models for inference with MCMC.

79 ASTRONOMY AND ASTROPHYSICS↗

Improved Earthquake Source Parameters with 3D Wavespeed Models in California and Nevada

Seismic tomography harnesses earthquake data to explore the inaccessible structure of the Earth. Adjoint waveform tomography (AWT), a method of seismic tomography, updates the tomographic model by optimizing the fit between observed earthquake data and synthetic waveforms. The synthetic data are calculated by solving the wave equation through a given 3D model. An important requirement to calculating synthetics is the source information (location, centroid time, depth, and moment tensor). Errors in source information affect the quality of the synthetics produced, which in turn can limit how structure can be inferred in the AWT workflow. Here, to test the effect of updating source information, we used MTTime (Chiang, 2020), a time-domain full-waveform moment tensor inversion code, to calculate the moment tensors and depths of 118 earthquakes that occurred in California and Nevada over a 20-yr period. We calculated 3D Green’s functions using a 3D seismic wavespeed model of California and Nevada (Doody et al., 2023b). We show that the inverted solutions provide better waveform fits than the Global Centroid Moment Tensor catalog and increase usable, well-correlated data by up to 7%. Therefore, we argue that recalculating source parameters should be considered in AWT workflows, particularly for smaller magnitude events (⁠M w > 5.0).

58 GEOSCIENCES↗

A Data Processing Pipeline To Extract A Knowledge Graph From Sec Documents For Socio-technical Analysis Of Critical Infrastructure Influence

The code is written in Python and consists of the following pipeline that is implemented in Apache Airflow. This pipeline intends to understand the companies that are directly or indirectly involved with a type of critical infrastructure system at some point in that system's lifecycle. The pipeline takes a configuration file that specifies a list of initial companies to consider, a geographic region of interest (disk) expressed as a latitude/longitude point and distance, and a set of SEC form types from which to extract entities and relations. There are three main components to this pipeline as currently implemented: Social Network Extraction, Critical Infrastructure Network Extraction, and Inference and Fusion. First, Social Network Extraction, implemented as the `organizations_sec` component of the workflow graph queries the SEC EDGAR webservice using the list of initial companies from the configuration file. Given this, it extracts metadata that documents the number of each type of form for the given set of companies and their location. This forms metadata represents a catalog of data sources for the extracted social network knowledge graph. The pipeline then downloads these forms from the website and saves them in a build directory for further processing. These documents are then parsed for entities and relations. Second, the Critical Network Extraction component extracts entities and relations for a critical infrastructure sector. Currently, we focus on Electric Vehicle charging stations and this information is available via the Department of Energy (DOE) database on fueling stations maintained by NREL. Third, the Inference and Fusion component relates the social network graph to the critical infrastructure graph in order to understand the impact of a company within a geographic region. Relations include ownership of the EV Charging Station asset as well as maintenance/ownership of the EV payment networks. The fused network can be represented in many ways and currently we emit a knowledge graph.

Weaver, GabrielA.↗

A map of the rubisco biochemical landscape

Rubisco is the primary CO 2 -fixing enzyme of the biosphere, yet it has slow kinetics. The roles of evolution and chemical mechanism in constraining its biochemical function remain debated. Engineering efforts aimed at adjusting the biochemical parameters of rubisco have largely failed, although recent results indicate that the functional potential of rubisco has a wider scope than previously known. Here we developed a massively parallel assay, using an engineered Escherichia coli in which enzyme activity is coupled to growth, to systematically map the sequence–function landscape of rubisco. Composite assay of more than 99% of single-amino acid mutants versus CO 2 concentration enabled inference of enzyme velocity and apparent CO 2 affinity parameters for thousands of substitutions. This approach identified many highly conserved positions that tolerate mutation and rare mutations that improve CO 2 affinity. These data indicate that non-trivial biochemical changes are readily accessible and that the functional distance between rubiscos from diverse organisms can be traversed, laying the groundwork for further enzyme engineering efforts.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Surface orientation ambiguity for single molecules at dielectric interfaces

Fluorescent molecules emit light in a dipole radiation pattern that can be used to infer their orientation through defocused fluorescence microscopy. Proper measurement of the orientation requires mathematical modeling of the radiation pattern expected for a dipole in the geometry of interest and subsequent comparison against experimental data. We point out an ambiguity in common calculations of these patterns that appears to compromise orientation measurements for molecules that are especially near dielectric surfaces. This results in a rotation of the measured emission dipole toward the surface for near-interface molecules, which can be mistaken for a preferentially horizontal orientation among the emitters. The proper treatment for on-surface emitters requires consideration of finite-sized current elements between two dielectric media, and we show that the theoretical ambiguity can be lifted via finite-element modeling. A prescription is provided for correcting measured orientations at arbitrary interfaces.

Dey, E. [University of Texas, Arlington, TX (Unite↗

Artificial Intelligence for Event Reconstruction and Higgs Physics at CMS and Future Colliders

This dissertation charts a trajectory in which advances in artificial intelligence (AI) play a central role in pushing the high-energy physics frontier, complementing progress driven by higher collision energies and larger colliders. The discovery potential of the LHC and future colliders relies on accurate reconstruction of increasingly complex particle collision events. In the CMS experiment, this task is performed by the particle-flow (PF) algorithm. This dissertation presents the first implementation of a machine-learning-based particle-flow (MLPF) reconstruction in the CMS detector based on transformer architectures. In simulated top quark--antiquark pair (ttbar) events under LHC Run~3 (2023--2024) conditions, MLPF improves jet energy resolution by 10--20\% compared to standard PF for jets with transverse momentum between 30--100\GeV. Runtime performance is evaluated using simulated multijet events, with a median inference time of 20\unit{ms} per event on an NVIDIA L4 GPU, compa red to approximately 110\unit{ms} for standard PF. The MLPF algorithm is also validated on Run~3 collision data, representing the first data-validated ML-based reconstruction pipeline at any LHC experiment. We then extend MLPF toward future electron--positron colliders and introduce the first full-simulation cross-detector transfer learning workflow for PF reconstruction. The model is pre-trained on simulated events from the Compact Linear Collider detector (CLICdet) and fine-tuned on the CLIC-like detector (CLD) proposed for the Future Circular Collider (FCC). This approach achieves up to a 40\% improvement in jet energy resolution over rule-based reconstruction while reducing the required training dataset size by an order of magnitude, demonstrating the potential of AI to accelerate detector development and optimization. This dissertation also demonstrates how modern AI techniques enhance the sensitivity of LHC physics analyses. A CMS search for highly Lorentz-boosted Higgs bosons decaying to \textrm{W} boson pairs is presented, focusing on the single-lepton final state. A dedicated fine-tuning strategy for \ParT yields an approximately 70\% increase in expected sensitivity relative to the baseline model. The analysis uses proton--proton collision data at a center-of-mass energy of \ensuremath{\sqrt{s}=13\TeV} collected by CMS between 2016 and 2018, corresponding to an integrated luminosity of 138\ensuremath{\ \mathrm{fb}^{-1}}. The expected significance of the search is $1.86\sigma$, with an observed signal strength of $-0.19^{+0.48}_{-0.46}$. Finally, explainable AI techniques are applied to the MLPF and \ParticleNet algorithms using layerwise relevance propagation, showing that both models base their predictions on physically meaningful features consistent with our physics intuition. Together, these results demonstrate how advanced AI methods can enhance reconstruction, analysis sensitivity, and interpretability, shaping the next era of experimental parti cle physics.

Mokhtar, Farouk [UC, San Diego]↗

Rapid Inverse Parameter Inference Using Physics-Informed Neural Network

As Li-ion batteries become more essential in today's economy, tools need to be developed to accurately and rapidly diagnose a battery's internal state-of-health. Using a Li-ion battery's (high-rate) voltage response, it is proposed to determine a battery's internal state through Bayesian calibration. However, Bayesian calibration is notoriously slow and requires thousands of model runs. To accelerate parameter inference using Bayesian calibration, a surrogate model is developed to replace the underlying physics-based Li-ion model. Developing a surrogate model for rapid Bayesian calibration analysis is discussed for both the single particle model (SPM) and the pseudo two-dimensional (P2D) model. Surrogate models are constructed using physics-informed neural networks (PINNs) that encode the influence of internal properties on observed voltage responses. In practice, a neural network can be trained by: 1) using simulation results of the physics-based model (i.e., a data-loss approach); 2) using the residuals of the governing equations themselves (i.e., a physics-loss approach); or 3) using a combination of simulation results and governing equation residuals. In the present work, PINNs are developed using a variety of training losses and neural network architectures. In this analysis, it is shown that a PINN surrogate model can be reliably trained with only physics-informed loss. However, using a coupled data-informed and physics-loss approach produced the most accurate PINNs.

Bayesian calibration↗

Codiscovering graphical structure and functional relationships within data: A Gaussian Process framework for connecting the dots

Most problems within and beyond the scientific domain can be framed into one of the following three levels of complexity of function approximation. Type 1: Approximate an unknown function given input/output data. Type 2: Consider a collection of variables and functions, some of which are unknown, indexed by the nodes and hyperedges of a hypergraph (a generalized graph where edges can connect more than two vertices). Given partial observations of the variables of the hypergraph (satisfying the functional dependencies imposed by its structure), approximate all the unobserved variables and unknown functions. Type 3: Expanding on Type 2, if the hypergraph structure itself is unknown, use partial observations of the variables of the hypergraph to discover its structure and approximate its unknown functions. These hypergraphs offer a natural platform for organizing, communicating, and processing computational knowledge. While most scientific problems can be framed as the data-driven discovery of unknown functions in a computational hypergraph whose structure is known (Type 2), many require the data-driven discovery of the structure (connectivity) of the hypergraph itself (Type 3). We introduce an interpretable Gaussian Process (GP) framework for such (Type 3) problems that does not require randomization of the data, access to or control over its sampling, or sparsity of the unknown functions in a known or learned basis. Its polynomial complexity, which contrasts sharply with the super-exponential complexity of causal inference methods, is enabled by the nonlinear ANOVA capabilities of GPs used as a sensing mechanism.

Science & Technology - Other Topics↗

Platform Of Optimal Experiment Management

The platform of optimal experiment management, POEM, powered with automated machine learning to accelerate the discovery of optimal solutions, and automatically guide the design of experiments to be evaluated. POEM currently supports 1) random model explorations for experiment design, 2) sparse grid model explorations with Gaussian Polynomial Chaos surrogate model to accelerate experiment design ,3) time-dependent model sensitivity and uncertainty analysis to identify the importance features for experiment design, 4) model calibrations via Bayesian inference to integrate experiments to improve model performance, and 5) Bayesian optimization for optimal experimental design. In addition, POEM aims to simplify the process of experimental design for users, enabling them to analyze the data with minimal human intervention, and improving the technological output from research activities.

Wang, Congjian [Idaho National Laboratory (INL), I↗

Bridging the Gap Between Astronomical Datasets: From Proof-of-Concept to AI Model Deployment with Domain Adaptation

Artificial Intelligence is transforming astrophysics, from studying stars and galaxies to analyzing cosmic large-scale structures. However, a critical challenge arises when AI models trained on simulations or past observational data are applied to new observation— leading to domain shifts, reduced robustness, and increased uncertainty of model predictions. This talk will explore these issues, highlighting examples such as galaxy morphology classification and cosmological parameter inference, where AI struggles to adapt across different datasets. We will discuss domain adaptation as a strategy to improve model generalization and mitigate biases—essential for making AI-driven discoveries reliable. Notably, these challenges extend beyond astrophysics, affecting AI applications across physics and other scientific domains. Addressing them is essential for maximizing AI’s impact in advancing scientific research.

Ćiprijanović, Aleksandra [Fermilab]↗

Low-latency Jet Tagging for HL-LHC Using Transformer Architectures

Transformers are the state-of-the-art model architectures and widely used in application areas of machine learning. However the performance of such architectures is less well explored in the ultra-low latency domains where deployment on FPGAs or ASICs is required. Such domains include the trigger and data acquisition systems of the LHC experiments. We present a transformer-based algorithm for jet tagging built with the HGQ2 framework, which is able to produce a model with heterogeneous bitwidths for fast inference on FPGAs, as required in the trigger systems at the LHC experiments. The bitwidths are acquired during training by minimizing the total bit operations as an additional parameter. By allowing a bitwidth of zero, the model is pruned in-situ during training. Using this quantization-aware approach, our algorithm achieves state-of-the-art performance while also retaining permutation invariance which is a key property for particle physics applications. Due to the strength of transformers in representation learning, our work also serves as a stepping stone for the development of a larger foundation model for trigger applications.

Laatu, Lauri [Imperial Coll., London]↗

Self-Supervised T-GCN for Detection of Disturbance and Propagation in Power Grid

Urban power systems increasingly rely on dense sensing to monitor grid reliability, yet disturbance labels are scarce and events are rare. We present a self-supervised spatio-temporal method that detects, localizes, and characterizes grid frequency disturbances across urban areas using only unlabeled data. Our approach trains a tiny Temporal Graph Convolutional Network (T-GCN) to forecast per-site frequency residuals (deviation from 60 Hz). The sensor graph is constructed directly from signals using pre-event Pearson correlation with a cross-correlation lag penalty without geocoding. At inference, node-level anomalies are the model's forecast errors; region-level alarms arise from connected components of high-score nodes. We estimate disturbance propagation by computing per-node arrival times (first persistent exceedance), then fit a planar or time-of-arrival model to obtain direction, speed, and an epicenter proxy. With only three real events collected at decisecond resolution across U.S. cities, we evaluate the T-GCN and report time-to-detect, footprint size, and propagation consistency. We further show that short-window embeddings from the T-GCN's hidden states enable few-shot event-vs-background recognition via a simple prototypical classifier. Despite minimal data and no labels, our system yields fast, spatially coherent detection and interpretable propagation maps, offering a practical, lightweight pathway to city-scale grid resilience analytics.

Niu, Haoran [ORNL] (ORCID:0000000155228297)↗

Uncertainty Quantification for Neutron Shield Using Convolutional Neural Networks

Uncertainty quantification from radiation transport calculations was conducted using a Bayesian inference approach. A surrogate model, using a convolutional neural network, was employed to emulate the neutron fluence, which was simulated with a Monte Carlo radiation transport model. This allowed for a computationally cheap approach to evaluate input parameters and to sample their corresponding posterior probability distributions. Experimental data from the literature were employed to perform uncertainty quantification studies for concrete shields. As a result, the method is a nonintrusive approach that enables studies with multiple input parameters and can be applied to any radiation transport model.

Bayesian inference↗

Congo Basin Water Balance and Terrestrial Fluxes Inferred From Satellite Observations of the Isotopic Composition of Water Vapor

Large spatio-temporal gradients in the Congo basin vegetation and rainfall are observed. However, its water-balance (evapotranspiration minus precipitation, or ET - P) is typically measured at basin-scales, limited primarily by river-discharge data, spatial resolution of terrestrial water storage measurements, and poorly constrained ET. We use observations of the isotopic composition of water vapor to quantify the spatio-temporal variability of net surface water fluxes across the Congo Basin between 2003 and 2018. These data are calibrated at basin scale using satellite gravity and total Congo river discharge measurements and then used to estimate time-varying ET - P over four quadrants representing the Congo Basin, providing first estimates of this kind for the region. We find that the multi-year record, seasonality, and interannual variability of ET - P from both the isotopes and the gravity/river discharge based estimates are consistent. Additionally, we use precipitation and gravity-based estimates with our water vapor isotope-based ET - P to calculate time and space averaged ET and net river discharge within the Congo Basin. These quadrant-scale moisture flux estimates indicate (a) substantial recycling of moisture in the Congo Basin (temporally and spatially averaged ET/P > 70%), consistent with models and visible light-based ET estimates, and (b) net river outflow is largest in the Western Congo where there are more rivers and higher flow rates. Our results confirm the importance of ET in modulating the Congo water cycle relative to other water sources.

54 ENVIRONMENTAL SCIENCES↗

Field intercomparison of ice nucleation measurements: the Fifth International Workshop on Ice Nucleation Phase 3 (FIN-03)

Abstract. The third phase of the Fifth International Ice Nucleation Workshop (FIN-03) was conducted at the Storm Peak Laboratory in Steamboat Springs, Colorado, in September 2015 to facilitate the intercomparison of instruments measuring ice-nucleating particles (INPs) in the field. Instruments included two online and four offline measurement systems for INPs, which are a subset of those utilized in the laboratory study that comprised the second phase of FIN (FIN-02). The composition of the total aerosols was characterized using the Particle Analysis by Laser Mass Spectrometry (PALMS) and Wideband Integrated Bioaerosol Sensor (WIBS) instruments, and aerosol size distributions were measured by a laser aerosol spectrometer (LAS). The dominant total particle compositions present during FIN-03 were composed of sulfates, organic compounds, and nitrates, as well as particles derived from biomass burning. Mineral-dust-containing particles were ubiquitous throughout and represented 67 % of supermicron particles. Total WIBS fluorescing particle concentrations for particles with diameters of > 0.5 µm were 0.04 ± 0.02 cm−3 (0.1 cm−3 highest; 0.02 cm−3 lowest), typical of the warm season in this region and representing ≈ 9 % of all particles in this size range as a campaign average. The primary focus of FIN-03 was the measurement of INP concentrations via immersion freezing at temperatures > −33 °C. Additionally, some measurements were made in the deposition nucleation regime at these same temperatures, representing one of the first efforts to include both mechanisms within a field campaign. INP concentrations via immersion freezing agreed within factors ranging from nearly 1 to 5 times on average between matched (time and temperature) measurements, and disagreements only rarely exceeded 1 order of magnitude for sampling times coordinated to within 3 h. Comparisons were restricted to temperatures lower than −15 °C due to the limits of detection related to sample volumes and very low INP concentrations. Outliers of up to 2 orders of magnitude occurred between −25 and −18 °C; a better agreement was seen at higher and lower temperatures. Although the 5–10 factor agreement of INP measurements found in FIN-03 aligned with the results of the FIN-02 laboratory comparison phase, giving confidence in progress of this measurement field, this level of agreement still equates to temperature uncertainties of 3.5 to 5 °C that may not be sufficient for numerical cloud modeling applications that utilize INP information. INP activity in the immersion-freezing mode was generally found to be an order of magnitude or more, making it more efficient than in the deposition regime at 95 %–99 % water relative humidity, although this limited data set should be augmented in future efforts. To contextualize the study results, an assessment was made of the composition of INPs during the late-summer to early-fall period of this study inferred through comparison to existing ice nucleation parameterizations and through measurement of the influence of thermal and organic carbon digestion treatments on immersion-freezing ice nucleation activity. Consistent with other studies in continental regions, biological INPs dominated at temperatures of > −20 °C and sometimes colder, while arable dust-like or other organic-influenced INPs were inferred to dominate below −20 °C.

54 ENVIRONMENTAL SCIENCES↗

Automated Bayesian high-throughput estimation of plasma temperature and density from emission spectroscopy

Here, this paper introduces a novel approach for automated high-throughput estimation of plasma temperature and density using atomic emission spectroscopy, integrating Bayesian inference with sophisticated physical models. We provide an in-depth examination of Bayesian methods applied to the complexities of plasma diagnostics, supported by a robust framework of physical and measurement models. Our methodology is demonstrated using experimental observations in the field of magneto-inertial fusion, focusing on individual and sequential shot analyses of the Plasma Liner Experiment at LANL. The results demonstrate the effectiveness of our approach in enhancing the accuracy and reliability of plasma parameter estimation and in using the analysis to reveal the deep hidden structure in the data. This study not only offers a new perspective of plasma analysis but also paves the way for further research and applications in nuclear instrumentation and related domains.

Bayesian inference↗

Examination of coal combustion management sites for microbiological and chemical signatures of groundwater impacts

Coal combustion accounts for 40% of the world’s electricity and generates more than a billion tons of coal combustion products (CCP) annually, half of which end up in landfills and impoundments. CCP contain mixtures of chemicals that can be mobile in the environment and impact the quality of surface water and potable groundwater. In this investigation, water samples from 14 coal combustion management sites across 4 physiographic regions in the United States, paired with background and down-gradient groundwater samples, were analyzed for water chemistry and microbiology. The objective was to determine if microbiology data alone, or supported by chemistry data, could reliably differentiate source waters and identify sites where CCP is known or expected to be influencing groundwater. Two percent of the total amplicons showed genus level conservation across CCP management sites, regions, and sample types; corresponding to ubiquitous, facultatively aerobic proteobacterial taxa that are generally recognized for the potential to respire using different terminal electron acceptors. Ordination plots did not reveal significant differences ( p > 0.05) in 16S rRNA gene amplicon diversity by CCP management site, water sample types, or physiographic regions. Contrastingly, chemistry distinguished sample types by standard water quality metrics (total dissolved solids, Ca:SO 4 ratio), alkali earth metals (K, Na, Li), selenium, boron, and fluoride. A focused evaluation of 16S rRNA gene amplicons for a subset of CCP management sites revealed microbiological features and chemical drivers (F, Ca, temperature) that positively identified the single CCP management site confirmed to have groundwater impacted by CCP leachate. At this site, 9 genera (>0.5% relative abundance) were exclusive to CCP porewater and downgradient groundwater. Inferred metabolisms for these taxa indicates potential for N and S biogeochemical transformations and 1-C metabolism that are consistent with a reducing environment, as evidenced by low ORP and depleted SO 4 2− . This research contributes to a growing understanding of conditions where these data types, analyses, and interpretation methods could be applied for distinguishing influence from CCP on the surrounding environment, as well as practical limitations.

01 COAL, LIGNITE, AND PEAT↗