Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Inference”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

PySIDT: Subgraph Isomorphic Decision Trees for Molecular Property Prediction

Accurate molecular property prediction is important across all fields of chemistry. Deep neural networks (DNNs) have become increasingly popular due to their ability to train automatically, avoiding the incredibly tedious process of constructing and extending traditional property estimation schemes. However, DNNs require large amounts of training data, are challenging to interpret, require large amounts of memory to load even during inference, and have severe difficulties incorporating qualitative chemical knowledge, which are often desired for molecular property prediction tasks. Here, in this study, we present PySIDT (https://github.com/zadorlab/PySIDT), a software for training and running inference on Subgraph Isomorphic Decision Trees (SIDTs). SIDTs are graph-based decision trees made of nodes associated with molecular substructures. Inference is done by descending target molecular structures down the decision tree to nodes with matching subgraph isomorphic substructures and making predictions based on the final (most specific) nodes matched. SIDTs scale down well to dataset sizes much smaller than is feasible for DNNs. As trees of molecular substructures, SIDTs are inherently readable and easy to visualize, making them easy to analyze. They are also straightforward to extend and retrain, facilitate uncertainty estimation, and enable easy integration of expert knowledge. We demonstrate the SIDT approach discussing its application to a diverse range of molecular prediction tasks: rate coefficient estimation, diffusion coefficient estimation, thermochemistry estimation, transition state bond stretch prediction, p K a prediction, stability of molecular structures, stability of surface structures, and prediction of surface lateral interaction energetics. Additionally, we demonstrate the power of the SIDT algorithms in two direct learning curve vanilla comparisons with the popular DNN-based software Chemprop and the popular gradient boosted trees-based software XGBoost on enthalpy of formation and rate coefficient prediction tasks. In particular, in the enthalpy of formation case, vanilla PySIDT is able to outperform vanilla Chemprop and XGBoost across the full range of training/validation set sizes out to 11,560 data points.

Johnson, Matthew Sean [Sandia National Laboratorie↗

Protorheology in practice: Avoiding misinterpretation

Protorheology is the paradigm that any observed flow or deformation is a chance to infer quantitative rheological properties. While this creates many opportunities for insight, there is significant risk of misunderstanding the physics involved, e.g. misinterpreting a liquid as a solid or mistaking viscous flow time as viscoelastic relaxation time. We describe these and other potential mistakes, use case studies to show how serious the problems can be, and contrast misinterpretations with correct approaches and interpretations. Some issues are especially important with materials involving colloidal particles and flows involving surface tension. Whether the reader is making inference from a tilted vial, time-lapse gravity-driven flow, a bounce test, die swell, or any other protorheology observation, the examples here serve as a guide for avoiding bad data in protorheology.

42 ENGINEERING↗

Data-driven prediction of scaling and ignition of inertial confinement fusion experiments

Recent advances in inertial confinement fusion (ICF) at the National Ignition Facility (NIF), including ignition and energy gain, are enabled by a close coupling between experiments and high-fidelity simulations. Neither simulations nor experiments can fully constrain the behavior of ICF implosions on their own, meaning pre- and postshot simulation studies must incorporate experimental data to be reliable. Linking past data with simulations to make predictions for upcoming designs and quantifying the uncertainty in those predictions has been an ongoing challenge in ICF research. We have developed a data-driven approach to prediction and uncertainty quantification that combines large ensembles of simulations with Bayesian inference and deep learning. The approach builds a predictive model for the statistical distribution of key performance parameters, which is jointly informed by past experiments and physics simulations. The prediction distribution captures the impact of experimental uncertainty, expert priors, design changes, and shot-to-shot variations. We have used this new capability to predict a 10× increase in ignition probability between Hybrid-E shots driven with 2.05 MJ compared to 1.9 MJ, and validated our predictions against subsequent experiments. We describe our new Bayesian postshot and prediction capabilities, discuss their application to NIF ignition and validate the results, and finally investigate the impact of data sparsity on our prediction results.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Neural Posterior Estimation for Scalable and Accurate Inverse Parameter Inference in Li-Ion Batteries

Diagnosing the internal state of Li-ion batteries is critical for battery research, operation of real-world systems, and prognostic evaluation of remaining lifetime. By using physics-based models to perform probabilistic parameter estimation via Bayesian calibration, diagnostics can account for the uncertainty due to model fitness, data noise, and the observability of any given parameter. However, Bayesian calibration in Li-ion batteries using electrochemical data is computationally intensive even when using a fast surrogate in place of physics-based models, requiring many thousands of model evaluations. A fully amortized alternative is neural posterior estimation (NPE). NPE shifts the computational burden from the parameter estimation step to data generation and model training, reducing the parameter estimation time from minutes to milliseconds, enabling real-time applications. The present work shows that NPE can infer parameters equally or more accurately than Bayesian calibration, even if it leads to higher voltage reconstruction errors. We also demonstrate that the higher computational costs for data generation are tractable even in high-dimensional cases (ranging from 6 to 27 estimated parameters). The NPE method also offers several interpretability advantages over Bayesian calibration, such as local parameter sensitivity to specific regions of the voltage curve. The NPE method is demonstrated using an experimental fast charge dataset, with parameter estimates validated against measurements of loss of lithium inventory and loss of active material. The implementation is made available in a companion repository (https://github.com/NatLabRockies/BatFIT).

25 ENERGY STORAGE↗

Bridging Cloud and Edge Computing at NREL Using CONNECT: Cloud Optimized Networking for Next-Gen Edge Computing Technologies [Slides]

CONNECT is an innovative on-premise hardware and software solution that integrates edge and cloud computing infrastructure at NREL. Built on the AWS Greengrass middleware and leveraging the MQTT protocol, CONNECT enables real-time data streaming from IoT devices and gateways to both cloud and local services, empowering researchers to rapidly capture, analyze, and act upon edge-generated data while leveraging cloud capabilities. The platform addresses research infrastructure challenges by providing a pre-approved platform which is already configured with the correct networking and cybersecurity baselines thus eliminating procurement delays and enabling on-demand availability. CONNECT's hybrid architecture efficiently manages burstable workloads, allowing research teams to dynamically scale computational capacity, handle peak data loads, and reduce operational bottlenecks. Advanced capabilities include built-in GPU support for executing machine learning models which enables low-latency inference at the edge from models trained in the cloud. This architecture supports real-time analytics and filtering, providing a mechanism to allow only transmitting and processing high-value data. Cloud-based configuration management permits engineers to manage on-premise systems remotely, optimizing operational efficiency. By bridging edge and cloud computing, CONNECT provides NREL researchers with a flexible, scalable platform that accelerates scientific discovery while maintaining robust security and performance standards.

97 MATHEMATICS AND COMPUTING↗

Iterative HOMER with uncertainties

We present iHOMER, an iterative version of the HOMER method to extract Lund fragmentation functions from experimental data. Through iterations, we address the information gap between latent and observable phase spaces and systematically remove bias. To quantify uncertainties on the inferred weights, we use a combination of Bayesian neural networks and uncertainty-aware regression. We find that the combination of iterations and uncertainty quantification produces well-calibrated weights that accurately reproduce the data distribution. A parametric closure test shows that the iteratively learned fragmentation function is compatible with the true fragmentation function.

Butter, Anja [Heidelberg Univ. (Germany); Sorbonne↗

Discriminative versus generative approaches to simulation-based inference

Most of the fundamental, emergent, and phenomenological parameters of particle and nuclear physics are determined through parametric template fits. Simulations are used to populate histograms which are then matched to data. This approach is inherently lossy, since histograms are binned and low-dimensional. Deep learning has enabled unbinned and high-dimensional parameter estimation through neural likelihood(-ratio) estimation. We compare two approaches for neural simulation-based inference (NSBI): one based on discriminative learning (classification) and one based on generative modeling. These two approaches are directly evaluated on the same datasets, with a similar level of hyperparameter optimization in both cases. In addition to a Gaussian dataset, we study NSBI using a Higgs boson dataset from the FAIR Universe Challenge. We find that both the direct likelihood and likelihood ratio estimation are able to effectively extract parameters with reasonable uncertainties. For the numerical examples and within the set of hyperparameters studied, we found that the likelihood ratio method is more accurate and/or precise. Both methods have a significant spread from the network training and would require ensembling or other mitigation strategies in practice.

high energy physics↗

A machine learning framework for accurate and robust analysis of radiation detector pulses

The microscopic properties of atomic nuclei are used to study various scientific questions. They are essential for understanding the fundamental forces of nature and the chemical evolution of the universe. Detecting decay radiation from radioactive nuclei makes it possible to probe these fundamental nuclear properties. Detector waveform traces may contain additional information about the radiation. Generally, advanced signal processing techniques are needed to extract this additional information, often involving fitting the waveform with model response functions using non-linear least-squares optimization with second-order gradient methods. While this is a powerful technique, it is also computationally expensive, leading to slow processing time, which scales with the volume of data. To address this problem, we have developed a machine learning (ML) approach that infers the characteristics of traces from a model detector response function. In particular, we are interested in classifying whether a single recorded trace consists of one or two pulse constituents and estimating the pulse parameters. Furthermore, our proposed ML method can precisely extract the pulses’ parameters, such as energy and timing information, and accurately classify the pulse multiplicity of a trace. Unlike non-learning-based approaches, our ML approach uses neural networks that are significantly faster at inference, as they do not require any optimization during this stage.

Curve fitting↗

A physics-constrained deep learning surrogate model of the runaway electron avalanche growth rate

A surrogate model of the runaway electron avalanche growth rate in a magnetic fusion plasma is developed. This is accomplished by employing a physics-informed neural network (PINN) to learn the parametric solution of the adjoint to the relativistic Fokker–Planck equation. The resulting PINN is able to evaluate the runaway probability function across a broad range of parameters in the absence of any synthetic or experimental data. This surrogate of the adjoint relativistic Fokker–Planck equation is then used to infer the avalanche growth rate as a function of the electric field, synchrotron radiation and effective charge. Predictions of the avalanche PINN are compared against first principle calculations of the avalanche growth rate with excellent agreement observed across a broad range of parameters.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Inferring performance metrics for laser direct drive experiments on OMEGA

Quantifying performance improvements on the OMEGA laser facility requires robust inference of established no-alpha performance metrics, which requires, at minimum, a model to infer the shocked fuel mass and pressure of the confined fusion plasma. In this work, we describe the methodology used to infer performance metrics on OMEGA and present the current state-of-the art model used to infer these metrics from OMEGA experiments. In particular, since neutron images of cryogenic implosions are not available on OMEGA at present, we present how x-ray sizes are determined on OMEGA using a Gaussian Process regression model and how the neutron production region's size is inferred from them. As a result, we end by benchmarking the model using synthetic data and 1-D LILAC simulations and test its experimental self-consistency across available x-ray diagnostic channels.

Gopalaswamy, V. [Laboratory for Laser Energetics, ↗

Deep inference of simulated strong lenses in ground-based surveys

The large number of strong lenses discoverable in future astronomical surveys will likely enhance the value of strong gravitational lensing as a cosmic probe of dark energy and dark matter. However, leveraging the increased statistical power of such large samples will require further development of automated lens modeling techniques. We show that deep learning and simulation-based inference (SBI) methods produce informative and reliable estimates of parameter posteriors for strong lensing systems in ground-based surveys. We present the examination and comparison of two approaches to lens parameter estimation for strong galaxy-galaxy lenses — Neural Posterior Estimation (NPE) and Bayesian Neural Networks (BNNs). We perform inference on 1-, 5-, and 12-parameter lens models for ground-based imaging data that mimics the Dark Energy Survey (DES). We find that NPE outperforms BNNs, producing posterior distributions that are more accurate, precise, and well-calibrated for most parameters. For the 12-parameter NPE model, the calibration is consistently within <10% of optimal calibration for all parameters, while the BNN is rarely within 20% of optimal calibration for any of the parameters. Similarly, residuals for most of the parameters are smaller (by up to an order of magnitude) with the NPE model than the BNN model. This work takes important steps in the systematic comparison of methods for different levels of model complexity.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

HostSub_GP: Precise Galaxy Background Subtraction in Transient Long-slit Spectroscopy with Gaussian Processes

We present a novel host galaxy subtraction technique in long-slit spectroscopy for extragalactic transients. Unlike classic methods which generally estimate the background using simple interpolation of local galaxy flux in the 2D spectrum, our approach leverages multi-band archival images of the host galaxies to model the background emission from the galaxy in the 2D spectrum. Such imaging encodes the wavelength-dependent galaxy profile along the slit, and is readily accessible through wide-field imaging surveys. We construct a smooth prior for the 2D galaxy profile with a Gaussian process (GP) based on these reference images, and use another GP to model the correlated deviations from the prior in the observed spectrum. This enables accurate inference of the galaxy flux blended with the transient. On synthetic long-slit data of a spiral galaxy extracted from a Multi Unit Spectroscopic Explorer hyper-spectral cube, the GP method remains robust as long as the host galaxy is spatially resolved and consistently outperforms classic methods. We apply the method to archival Keck spectra of two real transients, SN 2019eix and AT 2019qiz, to further demonstrate how the method uniquely recovers weak spectral features amid strong galaxy contamination, enabling refined constraints on the properties of both transients. We have released the software implementation, HostSub_GP, a scalable toolkit that leverages JAX, with an MIT license.

79 ASTRONOMY AND ASTROPHYSICS↗

Block Lanczos algorithm for lattice QCD spectroscopy and matrix elements

Recent work introduced a new framework for analyzing correlation functions with improved convergence and signal-to-noise properties, as well as rigorous quantification of excited-state effects, based on the Lanczos algorithm and spurious eigenvalue filtering with the Cullum-Willoughby test. Here, we extend this framework to the analysis of correlation-function matrices built from multiple interpolating operators in lattice quantum chromodynamics (QCD) by constructing an oblique generalization of the block Lanczos algorithm, as well as a new physically motivated reformulation of the Cullum-Willoughby test that generalizes to block Lanczos straightforwardly. The resulting block Lanczos method directly extends generalized eigenvalue problem (GEVP) methods, which can be viewed as applying a single iteration of block Lanczos. Block Lanczos provides qualitative and quantitative advantages over GEVP methods analogous to the benefits of Lanczos over the standard effective mass, including faster convergence to ground- and excited-state energies, explicitly computable two-sided error bounds, straightforward extraction of matrix elements of external currents, and asymptotically constant signal-to-noise. No fits or statistical inference are required. Proof-of-principle calculations are performed for noiseless mock-data examples as well as two-by-two proton correlation-function matrices in lattice QCD.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Ratio-preserving approach to cosmological concordance

Cosmological observables are particularly sensitive to key ratios of energy densities and rates, both today and at earlier epochs of the Universe. Well-known examples include the photon-to-baryon and the matter-to-radiation ratios. Equally important, though less publicized, are the ratios of pressure-supported to pressureless matter and the Thomson scattering rate to the Hubble rate around recombination, both of which observations tightly constrain. Preserving these key ratios in theories beyond the Λ Cold-Dark-Matter ( Λ CDM ) model ensures broad concordance with a large swath of datasets when addressing cosmological tensions. We demonstrate that a mirror dark sector, reflecting a partial Z 2 symmetry with the Standard Model, in conjunction with percentage level changes to the visible fine-structure constant and electron mass which represent a phenomenological change to the Thomson scattering rate, maintains essential cosmological ratios. Incorporating this ratio-preserving approach into a cosmological framework significantly improves agreement to observational data ( Δ χ 2 = - 35.72 ) and completely eliminates the Hubble tension with a cosmologically inferred H 0 = 73.80 ± 1.02 km / s / Mpc when including the S H 0 ES calibration in our analysis. While our approach is certainly nonminimal, it emphasizes the importance of keeping key ratios constant when exploring models beyond Λ CDM .

79 ASTRONOMY AND ASTROPHYSICS↗

CCQE-like $\nu_{e}$ Selection in SBND using Convolutional Visual Network

Neutrinos from the Booster Neutrino Beam (BNB) at Fermilab interact with argon in a Liquid Argon Time Projection Chamber (LArTPC) differently based on their flavour. By examining the particles produced in a charged-current interaction, both the interaction type and the neutrino flavour can be inferred. The Short Baseline Near Detector has the largest neutrino-argon cross section data to date, motivating in-depth studies of various cross-section channels and topologies. This project aims to select electron neutrino quasi-elastic-like (QE-like) interactions in SBND using Convolutional Visual Network (CVN) scores. The CVN is a neural network that processes visual information from an event and assigns scores corresponding to its likelihood of being each interaction type. An inclusive study of electron neutrino charged current interactions using CVN has already been conducted. This analysis aims to build on this study, further utilizing CVN scores to isolate electron neutrino QE-like interactions characterized by the presence of an electron and one or more protons ($N > 0$) in the final state. The project s goal is to contribute to the overall cross-section measurement efforts within the SBN program at Fermilab.

Breen, Genevieve [Mt. Holyoke Coll.]↗

Early Inference of Nuclear Technology-Directed Research Activities of Authors from Scientific Publications

Nuclear research articles can provide information about early nuclear proliferation indicators such as influential research entities and technology capability levels of a country, but detection of nuclear activities typically occurs after they have started. We investigate the extent to which nuclear research articles can be used to infer whether a research entity will acquire or develop a nuclear technology before it happens. Early detection of nuclear proliferation or technology development indicators from data is challenging due to partial observability, sparse and unlabeled information, and confounding signals from multiple concurrent activities. This paper presents the early detection problem as a sequential decision-making, goal inference problem, where the objective is to characterize and predict an individual’s, organization’s, or a country’s intent (unobserved goal-directed behavior) towards developing a nuclear capability from partially observed sequences of their research publications, using inverse reinforcement learning and Bayesian goal inference methods. A computational framework is presented, and its application demonstrated using 29,196 Scopus records for a case study related to a civil nuclear capability. The case study results serve as a proof-of-concept demonstration for inference of technology-directed research activity of authors who publish in the nuclear domain. The inference method, combined with advanced computing, may be used to assess and monitor activities pertaining to early developmental stages of a nuclear technology or capability, which in turn can help to identify and prioritize activities with nuclear proliferation potential for further investigation.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Reconstructing the Stripping History of the Sagittarius Stream with Neural Networks

The Sagittarius (Sgr) Stream is produced by the ongoing disruption of the Sgr dwarf spheroidal (dSph) galaxy and is thought to contain multiple wraps that were stripped during different pericentric passages. In this study, we introduce a neural-network–based method trained on N-body simulations to infer the stripping time of Sgr Stream stars directly from their phase-space coordinates. We combine spectroscopic data from SEGUE, APOGEE DR17, and LAMOST DR7 low-resolution spectroscopic (LRS) survey with Gaia EDR3 astrometry and distance estimates from the latest StarHorse catalog to identify high-quality Sgr Stream members. Applying our method to these stars, we measure a clear metallicity gradient with stripping time, well described by a linear relation with slope ∼0.3 dex Gyr −1 . We further predict the stripping times of globular clusters previously suggested to originate from the Sgr dSph. M 54, Terzan 7, Terzan 8, and Arp 2 exhibit stripping times consistent with being currently bound to the Sgr remnant. Pal 12, Whiting 1, and NGC 2419 are inferred to have been stripped 0.9 ± 0.1, 1.1 ± 0.2, and 2.1 ± 0.2 Gyr ago, respectively. For NGC 4147 and NGC 5634, whose membership in the Sgr system remains uncertain, our analysis suggests stripping times of 1.1 ± 0.4 and 1.1 ± 0.1 Gyr, respectively, if they are ultimately confirmed as genuine Sgr members. These results demonstrate that data-driven models of dynamical stripping histories offer a promising approach for reconstructing the formation and chemical evolution of the Sgr Stream.

79 ASTRONOMY AND ASTROPHYSICS↗

Frequentist cosmological constraints from full-shape clustering measurements in DESI DR1

We present a frequentist analysis of clustering measurements from Data Release 1 of the Dark Energy Spectroscopic Instrument (DESI) using the standard profile likelihood method. While Bayesian inferences for effective field theory models of galaxy clustering can be highly sensitive to prior choices for extended cosmological models, frequentist inferences are not susceptible to such effects. We compare frequentist and Bayesian constraints for the parameter set {σ 8 , H 0 , Ω m , w 0 , w a } using the full-shape power spectrum multipoles, post-reconstruction baryon acoustic oscillation (BAO) measurements, and external datasets from the CMB and type Ia supernovae measurements. The frequentist confidence intervals are significantly shifted relative to the Bayesian credible intervals for the w 0 w a CDM model, unless supernovae data are included. When DESI full-shape and BAO data are fit jointly, we obtain the following 1σ frequentist confidence intervals for ΛCDM (w 0 w a CDM): σ 8 = 0.863 +0.048 -0.040 , H 0 = 68.96 +0.81 -0.80 km s -1 Mpc -1 , Ω m = 0.3034 ± 0.0110 (σ 8 = 0.782 +0.060 -0.036 , H 0 = 63.7 +4.2 -2.0 km s -1 Mpc -1 , Ω m = 0.378 +0.024 -0.047 , w 0 = -0.16 +0.10 -0.50 , w a = -3.0 +1.7 ), corresponding to 0.8σ, 0.3σ, 0.7σ (2.1σ, 4.1σ, 6.5σ, 6.3σ, 6.6σ) shifts between the maximum likelihood estimate and the Bayesian posterior mean for ΛCDM (w 0 w a CDM) respectively.

Bayesian reasoning↗