Search NASASearch

SEARCH · Search NASA

Results for “Data Inference”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Precision beam diagnostics at the NuMI facility using muon monitor observations

The Neutrinos at the Main Injector (NuMI) facility at Fermilab delivers an intense neutrino beam for multiple experiments by producing pions that decay into neutrinos, muons, and other particles. Magnetic horns—the primary pion focusing elements in the NuMI beamline—exhibit predominantly linear optics, enabling a predictable relationship between the proton beam and the resulting pion and muon phase spaces. This study has two primary objectives: first, to evaluate and confirm the linearity of the horn focusing mechanism using analytical models and numerical simulations; and second, to demonstrate that key beam parameters—such as proton beam intensity, beam position on target, and horn current—can be extracted from muon monitor observations within this linear optics framework. Using a machine learning model trained on spill-by-spill muon monitor data, we infer the horn current with a precision of ±0.05%, the beam intensity with ±0.1%, and the beam position on target with ±0.018⁢ mm horizontally and ±0.013⁢ mm vertically. This approach provides a reliable cross-check of beam parameters, helping to reduce systematic uncertainties that are critical for future experiments such as the Deep Underground Neutrino Experiment, which will rely on the neutrino beam produced by the Long-Baseline Neutrino Facility.

Beam control

Validity of a finite temperature expansion for dense nuclear matter

In this work we provide a new, well-controlled expansion of the equation of state of dense matter from zero to finite temperatures (𝑇) while covering a wide range of charge fractions (𝑌 𝑄 ), from pure neutron to isospin symmetric nuclear matter. Our expansion can be used to describe neutron star mergers using the equation of state inferred from neutron star observations. We discuss how knowledge from low-energy nuclear experiments and heavy-ion collisions can be directly incorporated into the expansion. We also suggest new thermodynamic quantities of interest that can be calculated from theoretical models or directly inferred by experimental data that can be used to infer the finite temperature equation of state. With our new method, we can quantify the uncertainty in our finite 𝑇 and 𝑌 𝑄 expansions without making assumptions about the underlying degrees of freedom. We can reproduce results from a microscopic equation of state up to 𝑇 = 100 MeV for baryon chemical potential 𝜇 𝐵 ≳ 1100 MeV [≈(1–2)⁢𝑛 sat ] within 5% error, with even better results for larger 𝜇 𝐵 and/or lower 𝑇. We investigate the sources of numerical and theoretical uncertainty and discuss future directions of study.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Middle and Late-Holocene coastal environments of far Southeastern Russia as inferred from palynological and diatom data

Existing discontinuous palynological records from coastal and river valley exposures, with varying quality of radiocarbon control, suggested that the regional Holocene climate in southern areas of the Russian Far East was characterized by as many as 10 fluctuations in temperature and/or precipitation. In this study, palynological data from Zerkalnoye Lake, located on the western coast of the Sea of Japan, indicate a gradual decline in temperature through the Middle and Late-Holocene with a Holocene thermal maximum between 8800 and 5500 cal yr BP. Three wetter than present intervals, marked by an increase in Pinus koraiensis, occurred c. 3600–3500 cal yr BP, 2340–2050 cal yr BP, and 1830–1800 cal yr BP. The Zerkalnoye record shows the dominance of Quercus-broadleaf forests during the Middle and Early Holocene, although Quercus shows a gradual decrease as climate cooled during this interval. Diatom analysis of the Zerkalnoye sediments documents that the site was a shallow bay or coastal lagoon until c. 3540 cal yr BP. After that time, a freshwater lake was established, which had variable marine influences probably caused by sea level changes following deglaciation. The diatom data indicate a cool water interval between 2900 and 2580 cal yr BP, also noted in other sites in the region. Changes in the basin’s depositional environment do not affect the palynological record, indicating such sites can provide reliable paleovegetational and paleoclimatic records. In conclusion, the discrepancies of the various regional paleoclimatic scenarios indicate the need for the further collection of continuous records from coastal to alpine zones in this region of northeastern Asia.

Environmental sciences

A copula-based rank histogram ensemble filter

Serial ensemble filters implement triangular probability transport maps to reduce high-dimensional inference problems to sequences of state-by-state univariate inference problems. The univariate inference problems are solved by sampling posterior probability densities obtained by combining constructed prior densities with observational likelihoods according to Bayes' rule. Many serial filters in the literature focus on representing the marginal posterior densities of each state. However, rigorously capturing the conditional dependencies between the different univariate inferences is crucial to correctly sampling multidimensional posteriors. This work proposes a new serial ensemble filter, called the copula rank histogram filter (CoRHF), that seeks to capture the conditional dependency structure between variables via empirical copula estimates; these estimates are used to rigorously implement the triangular (state-by-state univariate) Bayesian inference. The success of the CoRHF is demonstrated on two-dimensional examples and the Lorenz'63 problem. A practical extension to the high-dimensional setting is developed by localizing the empirical copula estimation, and is demonstrated on the Lorenz'96 problem.

97 MATHEMATICS AND COMPUTING

DELVE Milky Way Satellite Galaxy Census. I. Satellite Population and Survey Selection Function in DES, DELVE, and Pan-STARRS

The properties of Milky Way satellite galaxies have important implications for galaxy formation, reionization, and the fundamental physics of dark matter. However, the population of Milky Way satellites includes the faintest known galaxies, and current observations are incomplete. To understand the impact of observational selection effects on the known satellite population, we perform rigorous, quantitative estimates of the Milky Way satellite galaxy detection efficiency in three wide-field survey datasets: the Dark Energy Survey Year 6, the DECam Local Volume Exploration Data Release 3, and the Pan-STARRS1 Data Release 1. Together, these surveys cover ∼13,600 deg 2 to g ∼ 24.0 and ∼27,700 deg 2 to g ∼ 22.5, spanning ∼91% of the high-Galactic-latitude sky (∣b∣ ≥ 15°). We apply multiple detection algorithms over the combined footprint and recover 49 known satellites above a strict census detection threshold. To characterize the sensitivity of our census, we run our detection algorithms on a large set of simulated galaxies injected into the survey data, which allows us to develop models that predict the detectability of satellites as a function of their properties. We then fit an empirical model to our data and infer the luminosity function, radial distribution, and size–luminosity relation of Milky Way satellite galaxies. Our empirical model predicts a total of $265^{+79}_{-47}$ satellite galaxies with −20 ≤ M V ≤ 0, half-light radii of 15 ≤ r 1/2 , (pc) ≤ 3000, and galactocentric distances of 10 ≤ D GC (kpc) ≤ 300. We also identify a mild anisotropy in the angular distribution of the observed galaxies, at a significance of ∼2σ, which can be attributed to the clustering of satellites associated with the LMC.

Tan, Chin Yi [Univ. of Chicago, IL (United States)

EMPDF : inferring the Milky Way mass with data-driven distribution function in phase space

We introduce the emPDF (empirical distribution function), a novel dynamical modelling method that infers the gravitational potential from kinematic tracers with optimal statistical efficiency under the minimal assumption of steady state. emPDF determines the best-fitting potential by maximizing the similarity between instantaneous kinematics and the time-averaged phase-space distribution function (DF), which is empirically constructed from observation upon the theoretical foundation of oPDF (Han et al. 2016). This approach eliminates the need for presumed functional forms of DFs or orbit libraries required by conventional DF- or orbit-based methods. emPDF stands out for its flexibility, efficiency, and capability in handling observational effects, making it preferable to the popular Jeans equation or other minimal assumption methods, especially for the Milky Way (MW) outer halo where tracers often have limited sample size and poor data quality. We apply emPDF to infer the MW mass profile using Gaia DR3 data of satellite galaxies and globular clusters, obtaining enclosed masses of M (,r) = 26±8, 46±8, 90±13⁠, and 149±40 x 10 10 M ⊙ at r = 30, 50, 100⁠, and 200 kpc, respectively. These are consistent with the updated constraints from simulation-informed DF fitting (Li et al. 2020). While the simulation-informed DF offers superior precision owing to the additional information extracted from simulations, emPDF is independent of such supplementary knowledge and applicable to general tracer populations. emPDF is currently implemented for tracers with complete 6D kinematics within spherical potentials, but it can potentially be extended to address more general problems.

Astrophysics of Galaxies (astro-ph.GA)

Inference of the neutron down-scatter ratio using neutron time-of-flight data for DT-layered implosions on the OMEGA laser

First analysis of the neutron time of flight (nTOF) data is presented to infer the areal density using the neutron down-scatter ratio (DSR) for OMEGA deuterium–tritium (DT)-layered (cryogenic) implosions. The required very-high-dynamic range (nTOF) signal is constructed using an nTOF detector [Forrest et al., Rev. Sci. Ins. 83, 10D919 (2012)] with multiple time-gated photomultiplier tubes. The nTOF data are analyzed using a forward-fit technique to account for the detector responses and infer the DSR and areal density for DT-layered implosions. The areal densities inferred using this analysis are found to be in agreement with the areal densities measured using the magnetic recoil spectrometer detector that is located in a similar line-of-sight. This work establishes the feasibility of the DSR measurement using nTOF data on OMEGA and provides an additional areal-density measurement that will benefit the assessment of the implosion performance and 3D reconstruction of the imploded core.

47 OTHER INSTRUMENTATION

Simultaneous inference of equation of state parameters and unknown data errors with uncertainty quantification via hierarchical Bayesian posterior maximization

Equations of state (EOSs) are a key component in running hydrodynamic simulations as they relate the thermodynamic states for the material. The Davis reactants EOS is commonly used for modeling high explosives (HEs), and the EOS model parameters are calibrated using material specific data. The calibrations are often performed with uncertainty quantification via Bayesian inference to account for uncertainty in the data and generate ensembles of likely parameters. However, there are relatively few HE data sets to use for calibration and many are historical and lack error information. In this work, we simultaneously calibrate the Davis reactants EOS model parameters and unknown data error terms for the high explosive PBX 9501. To quantify the uncertainty in the models and the data, we use a Bayesian framework for the calibration and compute the hierarchical Bayesian posterior distribution with both a posteriori maximization approach and Markov Chain Monte Carlo. In general, we find that, given our assumptions, the two approaches result in similar calibrated parameters, posterior covariance matrices, and insights about the parameters but that the posterior maximization requires far less computational resources.

97 MATHEMATICS AND COMPUTING

Inferring demographic and selective histories from population genomic data using a 2-step approach in species with coding-sparse genomes: an application to human data

Abstract The demographic history of a population, and the distribution of fitness effects (DFE) of newly arising mutations in functional genomic regions, are fundamental factors dictating both genetic variation and evolutionary trajectories. Although both demographic and DFE inference has been performed extensively in humans, these approaches have generally either been limited to simple demographic models involving a single population, or, where a complex population history has been inferred, without accounting for the potentially confounding effects of selection at linked sites. Taking advantage of the coding-sparse nature of the genome, we propose a 2-step approach in which coalescent simulations are first used to infer a complex multi-population demographic model, utilizing large non-functional regions that are likely free from the effects of background selection. We then use forward-in-time simulations to perform DFE inference in functional regions, conditional on the complex demography inferred and utilizing expected background selection effects in the estimation procedure. Throughout, recombination and mutation rate maps were used to account for the underlying empirical rate heterogeneity across the human genome. Importantly, within this framework it is possible to utilize and fit multiple aspects of the data, and this inference scheme represents a generalized approach for such large-scale inference in species with coding-sparse genomes.

Soni, Vivak (ORCID:0000000294969562)

Near-Efficient and Non-Asymptotic Multiway Inference

We establish non-asymptotic efficiency guarantees for tensor decomposition–based inference in count data models. Under a Poisson framework, we consider two related goals: (i) parametric inference , the estimation of the full distributional parameter tensor, and (ii) multiway analysis , the recovery of its canonical polyadic (CP) decomposition factors. Our main result shows that in the rank-one setting, a rank-constrained maximum-likelihood estimator achieves multiway analysis with variance matching the Cramér–Rao Lower Bound (CRLB) up to absolute constants and logarithmic factors. This provides a general framework for studying “near-efficient” multiway estimators in finite-sample settings. For higher ranks, we illustrate that our multiway estimator may not attain the CRLB; nevertheless, CP-based parametric inference remains nearly minimax optimal, with error bounds that improve on prior work by offering more favorable dependence on the CP rank. Numerical experiments corroborate near-efficiency in the rank-one case and highlight the efficiency gap in higher-rank scenarios.

97 MATHEMATICS AND COMPUTING

Evaluating the limitations of Bayesian metabolic control analysis

Bayesian Metabolic Control Analysis (BMCA) is a promising framework for inferring metabolic control coefficients in data-limited scenarios, combining Bayesian inference with linear-logarithmic (lin-log) rate laws. These metabolic control coefficients quantify how changes in enzyme activities affect steady-state fluxes and metabolite concentrations across a metabolic network. However, its predictive accuracy and limitations remain underexplored. This study systematically evaluates BMCA’s ability to infer elasticity values, flux control coefficients (FCC), and concentration control coefficients (CCC) under varying data availability conditions using three synthetic metabolic network models. We demonstrate that BMCA predictions are highly dependent on the inclusion of flux and enzyme concentration data, with the omission of these datasets leading to severe inaccuracies. In our synthetic, enzyme-perturbation datasets, external metabolite concentrations had minimal impact and, in some cases, their exclusion improved predictions; when external-nutrient perturbations were introduced and those concentrations were observed, gains were at most modest. Additionally, we find that posterior estimation with both ADVI and HMC can underestimate large-magnitude elasticities in our synthetic settings, with ADVI showing somewhat higher variance under strong up-regulation; thus, recovering |elasticity| ≳ 1.5 remains challenging regardless of the inference engine. ADVI also fails to accurately infer allosteric interactions, even when regulatory effects are strong. While BMCA maintains reasonable accuracy in partially recovering the rankings of the highest FCC values, its estimates of absolute values remain constrained by prior assumptions and data limitations. Our findings reveal the BMCA algorithm’s strengths and weaknesses, providing guidance on its application in metabolic engineering, and highlighting the need for methodological refinements to enhance its predictive capabilities.

59 BASIC BIOLOGICAL SCIENCES

Evaluating the limitations of Bayesian metabolic control analysis

AbstractBayesian Metabolic Control Analysis (BMCA) has emerged as a promising framework for inferring metabolic control coefficients in data-limited scenarios by integrating Bayesian inference with linlog rate laws. However, its predictive accuracy and limitations remain underexplored. This study systematically evaluates BMCA’s ability to infer elasticity values, flux control coefficients (FCCs), and concentration control coefficients (CCCs) under varying data availability conditions using three synthetic metabolic network models. Our findings highlight the strengths and weaknesses of BMCA, guiding its application in metabolic engineering and emphasizing the need for methodological refinements.Author summaryUnderstanding how enzymes control metabolic pathways is crucial for optimizing biomanufacturing and synthetic biology applications. Bayesian Metabolic Control Analysis (BMCA) is a promising computational method that integrates Bayesian inference with metabolic control analysis to estimate key control parameters, even in cases with limited experimental data. However, the accuracy and limitations of BMCA remain unclear. In this study, we systematically evaluate BMCA using three synthetic metabolic networks to determine how different types of physiological data impact its predictive performance. We find that BMCA requires flux and enzyme concentration data for accurate predictions, while external metabolite concentrations contribute little. Additionally, BMCA fails to predict elasticity values beyond a magnitude of 1.5 and reliably infer allosteric regulation, even when strong regulatory interactions exist. In addition, BMCA does not accurately rank metabolic control points, which may limit its utility in identifying key enzymes in engineered pathways. Our work provides practical insights into when and how BMCA can be applied, guiding future research in metabolic modeling and control analysis.

Shin, Janis (ORCID:0000000216572455)

Real-time Anomaly Detection for Liquid Argon Time Projection Chambers

We present a real-time anomaly detection framework for liquid argon time projection chambers (LArTPCs), targeting applications in particle physics experiments such as the Short Baseline Near Detector (SBND) or the future Deep Underground Neutrino Experiment (DUNE). These experiments employ detectors that generate and stream high-resolution but sparse images of neutrino and other particle interactions. Our approach utilizes anomaly detection with autoencoders, compressed through knowledge distillation (KD), to enable the detection of anomalous signals in the data through efficient inference on resource-constrained hardware. The framework is targeted for deployment on computing platforms equipped with field-programmable gate arrays (FPGAs), GPUs, or CPUs, allowing low-latency selection of relevant activity directly from the raw detector data stream. We demonstrate that our approach is suitable for the detection and localization of anomalously "high-multiplicity" activity, and outline promising applications for LArTPC online data filtering and triggering.

FOS: Physical sciences

Hydra: computer vision for data quality monitoring

Hydra, initially developed for Hall-D in 2019, is a system that utilizes computer vision to perform near real time data quality monitoring. Since then, it has been deployed across all experimental halls at Jefferson Lab, with the CLAS12 collaboration in Hall-B being the first outside of GlueX to fully utilize Hydra. The system comprises back end processes that manage the models, their inferences, and the data flow. Finally, the front-end components, accessible via web pages, allow detector experts and shift crews to view and interact with the system.

47 OTHER INSTRUMENTATION

A Representation Fusion Framework for Decoupling Diagnostic Information in Multimodal Learning

Modern medicine increasingly relies on multimodal data, ranging from clinical notes to imaging and genomics, to guide diagnosis and treatment. However, integrating these heterogeneous data sources in a principled and interpretable manner remains a major challenge. We present MODES (Multi-mOdal Disentangled Embedding Space), a representation fusion framework that explicitly separates shared and modality-specific factors of variation, offering a structured latent space for multimodal information that improves both prediction and interpretability. By leveraging pre-trained unimodal foundation models, MODES mitigates the dependency on extensive paired datasets, crucial in data-scarce clinical settings. We introduce a masking strategy that optimizes representation dimensionality by eliminating low-information dimensions, to achieve compact, information-rich representations. Our framework demonstrates superior performance in predicting diagnoses and phenotypes compared to unimodal and conventional fusion models. MODES also enables robust diagnostic inference in missing data scenarios, offering an opportunity toward interpretable and efficient multimodal diagnostics in personalized healthcare.

60 APPLIED LIFE SCIENCES