Search NASA⌕ Search

SEARCH · Search NASA

Results for “Inference”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Applying Machine Learning and Bayesian Inference to Identify and Locate Moving Anthropogenic Sources Using Distributed Acoustic Sensing Data

Distributed acoustic sensing (DAS) systems, which use existing telecommunication fibers, offer high‐resolution capabilities ideal for recording anthropogenic sources. However, the complexity of urban environments and the large amount of data recorded by DAS require automated methods to efficiently detect and categorize anthropogenic sources. Here, we evaluate how well three machine learning models (k‐nearest neighbor [k‐NN], convolutional neural networks, and recurrent‐convolutional neural networks) can identify various anthropogenic sources recorded by DAS. Our findings reveal that both k‐NN and neural network methods perform well in high signal‐to‐noise ratio (SNR) settings. However, their accuracy decreases at SNRs <4. We also use Kalman filtering, a form of Bayesian inference, on backprojected locations of these sources to recover locations that generally fall within standard smartphone Global Positioning System errors. By combining machine learning and Kalman filter results, we calculate a multidimensional model of moving anthropogenic sources. These results demonstrate the potential of DAS data in urban seismology for accurately identifying and locating such sources. Depending on the research objectives, these sources can be further studied or filtered out to improve the quality of seismic data for earthquake studies. Such methods provide a valuable tool for urban seismology and seismic hazard analysis.

Luckie, Thomas William [Sandia National Laboratori↗

Dynamical Inference and 3D Imaging of Magnetized Dusty Plasmas

This is the final technical report for this DOE award. The motion of dust particles in a laboratory plasma has been studied for nearly 30 years and has led to the discovery of strongly-coupled crystalline structures and nonequilibrium dynamics governed by gravitational, hydrodynamic, and electrostatic forces. Many of the inter-particle forces are complex; they can be non-reciprocal, and non-additive. Deciphering the mélange of particle interactions has been a challenging task, nevertheless, magnetized dusty plasmas remains an open challenge with limited understanding. With applications ranging from magnetically-confined plasmas for fusion to near-surface planetary environments, this research field is ripe for well-controlled, laboratory experiments and new theoretical tools. In collaboration with the MDPX experiment at Auburn University, this proposal aims to tease apart both the known and unknown forces that drive dusty plasmas in magnetized environments using high-resolution, three-dimensional imaging and particle tracking coupled with modern dynamical inference and machine learning techniques.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Conformal Hierarchical Simulation-Based Inference with Local Validity

Trustworthy and interpretable uncertainty quantification is a long-standing challenge in artificial intelligence. Simulation-based inference (SBI) comprises a broad swath of approaches for estimating latent parameters with uncertainties. Although flexible neural density estimators in SBI can be remark- ably expressive capturing highly structured, high-dimensional posteriors their credible regions can be badly mis-calibrated and are often only accompanied by heuristic coverage checks. We present the first SBI framework that delivers finite-sample local valid coverage guarantees that hold in the neighborhood of each observation. Our framework can couple any off-the-shelf hierarchical SBI engine with a confor- mal Bayesian post-processing step that operates on the posterior predictive density. A kernel-weighted conformity score adapts the conformal quantile to the local geometry of the data, yielding prediction sets that are simultaneously (i) marginally calibrated, (ii) locally valid, and (iii) hierarchical, handling global and observation-specific parameters in a single pass. Through experiments on synthetic data and benchmarks from neuroscience and physics, we show that our approach attains 1 − α coverage, where prior SBI methods under- or over-cover. Our approach also maintains a competitive, credible set size with minimal computational overhead. Finally, our approach can be used to make predictions on real data and give valid credible regions modulo weight-initialization-based model mis-specification.

Trivedi, Shubhendu [Fermilab]↗

AlgaeOrtho, a bioinformatics tool for processing ortholog inference results in algae

Introduction: Microalgae constitute a prominent feedstock for producing biofuels and biochemicals by virtue of their prolific reproduction, high bioproduct accumulation, and the ability to grow in brackish and saline water. However, naturally occurring wild type algal strains are rarely optimal for industrial use; therefore, bioengineering of algae is necessary to generate superior performing strains that can address production challenges in industrial settings, particularly the bioenergy and bioproduct sectors. One of the crucial steps in this process is deciding on a bioengineering target: namely, which gene/protein to differentially express. These targets are often orthologs which are defined as genes/proteins originating from a common ancestor in divergent species. Although bioinformatics tools for the identification of protein orthologs already exist, processing the output from such tools is nontrivial, especially for a researcher with little or no bioinformatics experience. Methods: The present study introduces AlgaeOrtho, a user-friendly tool that builds upon the SonicParanoid orthology inference tool (based on an algorithm that identifies potential protein orthologs based on amino acid sequences) and the PhycoCosm database from JGI (Joint Genome Institute) to help researchers identify orthologs of their proteins of interest in multiple diverse algal species. Results: The output of this application includes a table of the putative orthologs of their protein of interest, a heatmap showing sequence similarity (%), and an unrooted tree of the putative protein orthologs. Notably, the tool would be instrumental in identifying novel bioengineering targets in different algal strains, including targets in not-fully annotated algal species, since it does not depend on existing protein annotations. We tested AlgaeOrtho using three case studies, for which orthologs of proteins relevant to bioengineering targets, were identified from diverse algal species, demonstrating its ease of use and utility for bioengineering researchers. Discussion: This tool is unique in the protein ortholog identification space as it can visualize putative orthologs, as desired by the user, across several algal species.

09 BIOMASS FUELS↗

A Strong Gravitational Lens Is Worth a Thousand Dark Matter Halos: Inference on Small-scale Structure Using Sequential Methods

Strong gravitational lenses are a singular probe of the Universe’s small-scale structure—they are sensitive to the gravitational effects of low-mass (<10 10 M ⊙ ) halos even without a luminous counterpart. Recent strong-lensing analyses of dark matter structure rely on simulation-based inference (SBI). Modern SBI methods, which leverage neural networks as density estimators, have shown promise in extracting the halo-population signal. However, it is unclear whether the constraints from these models are limited by the methodology or the data. In this study, we introduce an accelerator-optimized simulation pipeline that can generate lens images with realistic subhalo populations in milliseconds. Leveraging this simulator, we identify the main limitation of our fiducial SBI analysis: training set size. We then adopt a sequential neural posterior estimation (SNPE) approach, allowing us to refine the training distribution to align with the observed data. Using only one-fifth as many mock Hubble Space Telescope images, SNPE matches the constraints on the low-mass halo population produced by our best nonsequential model. Our experiments suggest that an over 3 order-of-magnitude increase in training set size and GPU hours would be required to achieve an equivalent result without sequential methods. While the full potential of the existing lens sample remains to be explored, the notable improvement in constraining power enabled by our sequential approach highlights that current constraints are limited primarily by methodology and not the data itself. Moreover, our results emphasize the need to treat training set generation and model optimization as interconnected stages of any cosmological analysis using SBI.

79 ASTRONOMY AND ASTROPHYSICS↗

Inference of Multichannel r -process Element Enrichment in the Milky Way Using Binary Neutron Star Merger Observations

Observations of GW170817 strongly suggest that binary neutron star (BNS) mergers produce rapid neutron-capture nucleosynthesis ( r -process) elements. However, it remains an open question whether these mergers can account for all the r -process element enrichment in the Milky Way’s history. Here, we constrain the contributions of the BNS channel using astrophysical neutron star observations. The rate and mass distributions are constrained by LIGO/Virgo/Kagra through the latest catalog GWTC-3, the neutron star equation of state by gravitational-wave, radio, and X-ray observations, and the delay time distribution by short gamma-ray burst (GRB) host galaxy associations. We present a Bayesian framework to consistently combine these observations with abundance information to quantify the contribution and uncertainties of single and multiple astrophysical enrichment sources, and obtain a distribution of per-event BNS r -process element yields consistent with geophysical and astrophysical abundance constraints. We then adopt a Galactic chemical evolution model assuming an instantaneous and fixed amount of Fe enrichment from core-collapse supernovae, and show that BNS-only enrichment scenarios remain inconsistent with the observed r-process abundance trend of disk stars in the Galaxy even with the uncertainties in BNS merger observations. Using stellar abundance observations instead of the short GRB constraints, we can infer a shorter BNS delay time distribution with power-law index α ≤ −2.0 and minimum delay time ${t}_{{\rm{\min }}}\leqslant 40$ Myr at 90% confidence, consistent with detailed Galactic chemical evolution models. Such delay times are in tension with those predicted by standard BNS formation models. Alternatively, we confirm that a two-channel scenario, in which the second channel tracks the star formation history without significant delay, can account for both Galactic stellar and short GRB observations. We estimate that 45%–90% of the r -process abundance in the Milky Way today would have been produced by this star formation-tracking channel, rather than BNS mergers with significant delay times.

gravitational wave astronomy↗

A Nonparametric Method for the Inference of Halo Occupation Distributions

The galaxy–halo connection traces processes by which galaxies form and evolve. The halo occupation distribution (HOD) describes the relationship between galaxies and their host dark matter haloes. Measurements of the galaxy two-point correlation function (2PCF) allow us to extract information about the HODs of observed galaxy samples. Several parametric HOD models have been proposed in the literature, but the choice of parameterization restricts the space of possible HODs. To resolve this issue, we introduce a nonparametric HOD fitting method in which we train an emulator to learn the mappings among the galaxy 2PCF, physical properties used to select galaxy samples, and the HOD, all obtained from simulated past light cones constructed with the Santa Cruz semianalytic model. Implementing this emulator within a likelihood analysis framework, we derive constraints on the HOD of a galaxy sample when provided with a measurement of its 2PCF. Using the emulator to accelerate likelihood evaluations, we test the nonparametric HOD approach on a set of 2PCFs for mock galaxy samples drawn from the TNG100-1 simulation and selected above threshold values of stellar mass and star formation rate. Our framework is able to recover TNG100-1 HODs within 0.2 dex. We use the TNG100-1 mocks to tune the reported uncertainties to estimate those expected in the analysis of observations. Comparing to parametric HOD modelling routines applied to the same mock galaxy samples, our approach consistently infers the HOD with comparable or greater precision and accuracy.

Kennedy, Jacob [Rutgers Univ., Piscataway, NJ (Uni↗

Rapid Inference of Logic Gate Neural Networks for Anomaly Detection in High Energy Physics

The increasing data rates and complexity of detectors at the Large Hadron Collider (LHC) necessitate fast and efficient machine learning models, particularly for rapid selection of what data to store, known as triggering. Building on recent work in differentiable logic gates, we present a public implementation of a Convolutional Differentiable Logic Gate Neural Network (CLGN). We apply this to detecting anomalies at the Level-1 Trigger at CMS using public data from the CICADA project. We demonstrate that the CLGN achieves physics performance on par with or superior to conventional quantized neural networks. We also synthesize an LGN for a Field-Programmable Gate Array (FPGA) and show highly promising FPGA characteristics, notably zero Digital Signal Processor (DSP) resource usage. This work highlights the potential of logic gate networks for high-speed, on-detector inference in High Energy Physics and beyond.

FOS: Physical sciences↗

Leveraging Hydropower Multi-Sensor Data for Inference and Age-Informed Modeling

Increased demand of operational flexibility such as faster ramp up/down in generation, and more frequent start/stops are putting hydropower plants and their associated components in unprecedented stress. Consequently, these plants are at the high risk of extended and more frequent outage to accommodate unscheduled, and unexpected maintenance. Therefore, hydropower plants are in critical need of data driven and age-informed analysis for their regular and unscheduled operation. Yet not all hydropower plants are exhaustively equipped with sensors and/or measurement streams for their respective components – demanding solutions on how to detect, identify, and locate the cause of any event from the unobservable. Idaho National Laboratory (INL) analyzed the anonymized measurements and event records from the Hydropower Research Institute (HRI) to address this issue, as part of the Water Power Technologies Office (WPTO) funded one year multi-lab project. First, we investigated how time series of multiple sensor measurements can be leveraged to identify an event “root cause” as well as to develop an inference (i.e., estimate the unobservable) problem. INL also investigated how individual hydropower components’ reaction or response times vary across the pre-event, during event, and post-event conditions – enabling the hydropower dynamic models to be age-informed. Finally, the impact of clustering multi-sensor time series on short-term vibration prediction is analyzed. INL will present key findings from these analyses and recommend next steps for stakeholder adoption.

13 HYDRO ENERGY↗

Automated Membership Inference Attacks: Discovering MIA Signal Computations using LLM Agents

Membership inference attacks (MIAs), which enable adversaries to determine whether specific data points were part of a model's training dataset, have emerged as an important framework to understand, assess, and quantify the potential information leakage associated with machine learning systems. Designing effective MIAs is a challenging task that usually requires extensive manual exploration of model behaviors to identify potential vulnerabilities. In this paper, we introduce AutoMIA -- a novel framework that leverages large language model (LLM) agents to automate the design and implementation of new MIA signal computations. By utilizing LLM agents, we can systematically explore a vast space of potential attack strategies, enabling the discovery of novel strategies. Our experiments demonstrate AutoMIA can successfully discover new MIAs that are specifically tailored to user-configured target model and dataset, resulting in improvements of up to 0.18 in absolute AUC over existing MIAs. This work provides the first demonstration that LLM agents can serve as an effective and scalable paradigm for designing and implementing MIAs with SOTA performance, opening up new avenues for future exploration.

Tran, Toan Viet [Emory University]↗

Sequential Kalman tuning of the t -preconditioned Crank-Nicolson algorithm: efficient, adaptive and gradient-free inference for Bayesian inverse problems

Ensemble Kalman Inversion (EKI) has been proposed as an efficient method for the approximate solution of Bayesian inverse problems with expensive forward models. However, when applied to the Bayesian inverse problem EKI is only exact in the regime of Gaussian target measures and linear forward models. Here, in this work we propose embedding EKI and Flow Annealed Kalman Inversion, its normalizing flow (NF) preconditioned variant, within a Bayesian annealing scheme as part of an adaptive implementation of the t-preconditioned Crank-Nicolson (tpCN) sampler. The tpCN sampler differs from standard pCN in that its proposal is reversible with respect to the multivariate t-distribution. The more flexible tail behaviour allows for better adaptation to sampling from non-Gaussian targets. Within our Sequential Kalman Tuning (SKT) adaptation scheme, EKI is used to initialize and precondition the tpCN sampler for each annealed target. The subsequent tpCN iterations ensure particles are correctly distributed according to each annealed target, avoiding the accumulation of errors that would otherwise impact EKI. We demonstrate the performance of SKT for tpCN on three challenging numerical benchmarks, showing significant improvements in the rate of convergence compared to adaptation within standard SMC with importance weighted resampling at each temperature level, and compared to similar adaptive implementations of standard pCN. The SKT scheme applied to tpCN offers an efficient, practical solution for solving the Bayesian inverse problem when gradients of the forward model are not available. Code implementing the SKT schemes for tpCN is available at https://github.com/RichardGrumitt/KalmanMC.

97 MATHEMATICS AND COMPUTING↗

Inferring precocial Chinook Salmon production through single‐parentage assignments

Abstract Objective Parentage analysis is a routine methodology in fisheries research, but study systems exist where it is impractical to sample both parents. The ability to reliably assign offspring to a single parent is beneficial in these situations. We applied single‐parentage assignments to a naturally spawning population of Chinook Salmon Oncorhynchus tshawytscha to quantify production of anadromous returns by unsampled precocial males. Methods We used an approach that focused on two important aspects of parentage analyses: (1) addressing the presence of family structure within the set of sampled parents and (2) controlling for false‐positive and false‐negative assignments. Result Results indicated that 30% of reproductively successful males were precocial males, which produced 20% of the returning anadromous offspring. Conclusion This study provides a framework for applying single‐parent assignments in a salmonid study system while explicitly addressing sources of assignment errors.

Steele, Craig A.↗

Gaussian processes for inferring parton distributions

The extraction of parton distribution functions (PDFs) from experimental or lattice QCD data is an ill-posed inverse problem, where regularization strongly impacts both systematic uncertainties and the reliability of the results. We study a framework based on Gaussian Process Regression (GPR) to reconstruct PDFs from lattice QCD matrix elements. Within a Bayesian framework, Gaussian processes serve as flexible priors that encode uncertainties, correlations, and constraints without imposing rigid functional forms. We investigate a wide range of kernel choices, mean functions, and hyperparameter treatments. We quantify information gained from the data using the Kullback-Leibler divergence. Synthetic data tests demonstrate the consistency and robustness of the method. Our study establishes GPR as a systematic and non-parametric approach to PDF reconstruction, offering controlled uncertainty estimates and reduced model bias in lattice QCD analyses.

hadronic spectroscopy↗

Inference of phase field fracture models

The phase field approach to modeling fracture uses a diffuse damage field to represent cracks. This representation mollifies singularities that arise in computations with sharp interface models and some of the resultant difficulties in the mathematical and numerical treatment of fracture. Phase field fracture models have proven effective at representing crack propagation, branching, and merging. Specific formulations, beginning with brittle fracture, have also been shown to converge to classical solutions. Extensions to cover the range of material failure, including ductile and cohesive fracture, lead to an array of possible models. There exists a large body of literature focusing on this class of models and on the impact of model form on the predicted crack evolution. However, there have not been systematic studies into how optimal models may be chosen. Here, we take a first step in this direction by developing formal methods for identification of the best parsimonious model of phase field fracture given full-field data on the damage and deformation fields. We consider some of the main models that have been used for the degradation of elastic response due to damage and its propagation. Our approach builds upon Variational System Identification (VSI), a weak form variant of the Sparse Identification of Nonlinear Dynamics (SINDy). Furthermore, in this first communication we focus on synthetically generated data but we also consider central issues associated with the use of experimental full-field data, such as data sparsity and noise.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗