Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Inference”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Dimensional Reduction for Sampled Priors and Application to Photometric Redshift Distributions

A typical Bayesian inference on the values of some parameters of interest q from some data D involves running a Markov Chain (MC) to sample from the posterior $p$($q$,$n$|$D$) $\propto$ $\mathcal{L}$($D$|$q$,$n$)$p$(q)$p$($n$), where n are some nuisance parameters with a separable prior. In some cases, the nuisance parameters are high-dimensional, and their prior p(n) is itself defined only by a set of samples that have been drawn from some other MC. The MC for the posterior will typically require evaluation of p(n) at arbitrary values of n, i.e., one needs to provide a density estimator over the full n space from the provided samples. But the high dimensionality of n hinders both the density estimation and the efficiency of the MC for the posterior. We describe a solution to this problem: a linear compression of the n space into a much lower-dimensional space u, which projects away directions in n space that cannot appreciably alter $\mathcal{L}$. The algorithm for doing so is a slight modification to principal components analysis, and is less restrictive on p(n) than other proposed solutions to this issue. We demonstrate this “mode projection” technique using the analysis of 2-point correlation functions of weak lensing fields and galaxy density in the Dark Energy Survey, where n is a binned representation of the redshift distribution n(z) of the galaxies.

79 ASTRONOMY AND ASTROPHYSICS↗

Constraining the Dense Matter Equation of State with New NICER Mass–Radius Measurements and New Chiral Effective Field Theory Inputs

Pulse profile modeling of X-ray data from the Neutron Star Interior Composition Explorer is now enabling precision inference of neutron star mass and radius. Combined with nuclear physics constraints from chiral effective field theory (χEFT), and masses and tidal deformabilities inferred from gravitational-wave detections of binary neutron star mergers, this has led to a steady improvement in our understanding of the dense matter equation of state (EOS). Here, we consider the impact of several new results: the radius measurement for the 1.42 M ⊙ pulsar PSR J0437−4715 presented by Choudhury et al., updates to the masses and radii of PSR J0740+6620 and PSR J0030+0451, and new χEFT results for neutron star matter up to 1.5 times nuclear saturation density. Using two different high-density EOS extensions—a piecewise-polytropic (PP) model and a model based on the speed of sound in a neutron star (CS)—we find the radius of a 1.4 M ⊙ (2.0 M ⊙ ) neutron star to be constrained to the 95% credible ranges ${12.28}_{-0.76}^{+0.50}$ km (${12.33}_{-1.34}^{+0.70}$ km) for the PP model and ${12.01}_{-0.75}^{+0.56}$ km (${11.55}_{-1.09}^{+0.94}$ km) for the CS model. The maximum neutron star mass is predicted to be ${2.15}_{-0.16}^{+0.14}$M ⊙ and ${2.08}_{-0.16}^{+0.28}$M ⊙ for the PP and CS models, respectively. We explore the sensitivity of our results to different orders and different densities up to which χEFT is used, and show how the astrophysical observations provide constraints for the pressure at intermediate densities. Moreover, we investigate the difference R 2.0 − R 1.4 of the radius of 2 M ⊙ and 1.4 M ⊙ neutron stars within our EOS inference.

gravitational wave sources↗

Characterization of DESI fiber assignment incompleteness effect on 2-point clustering and mitigation methods for DR1 analysis

We present an in-depth analysis of the fiber assignment incompleteness in the Dark EnergySpectroscopic Instrument (DESI) Data Release 1 (DR1). This incompleteness is caused by therestricted mobility of the robotic fiber positioner in the DESI focal plane, which limits thenumber of galaxies that can be observed at the same time, especially at small angular separations.As a result, the observed clustering amplitude is suppressed in a scale-dependent manner, which,if not addressed, can severely impact the inference of cosmological parameters. We discuss themethods adopted for simulating fiber assignment on mocks and data. In particular, we introducethe fast fiber assignment (FFA) emulator, which was employed to obtain the power spectrumcovariance adopted for the DR1 full-shape analysis. We present the mitigation techniques,organised in two classes: measurement stage and model stage. We then use high fidelity mocks as areference to quantify both the accuracy of the FFA emulator and the effectiveness of the differentmeasurement-stage mitigation techniques. This complements the studies conducted in a parallelpaper for the model-stage techniques, namely the θ-cut approach. We find that pairwiseinverse probability (PIP) weights with angular upweighting recover the “true” clustering in allthe cases considered, in both Fourier and configuration space. Notably, we present the first everpower spectrum measurement with PIP weights from real data.

79 ASTRONOMY AND ASTROPHYSICS↗

Trust Your Gut: Comparing Human and Machine Inference from Noisy Visualizations

People commonly utilize visualizations not only to examine a given dataset, but also to draw generalizable conclusions about the underlying models or phenomena. Prior research has compared human visual inference to that of an optimal Bayesian agent, with deviations from rational analysis viewed as problematic. However, human reliance on non-normative heuristics may prove advantageous in certain circumstances. We investigate scenarios where human intuition might surpass idealized statistical rationality. In two experiments, we examine individuals’ accuracy in characterizing the parameters of known data-generating models from bivariate visualizations. Our findings indicate that, although participants generally exhibited lower accuracy compared to statistical models, they frequently outperformed Bayesian agents, particularly when faced with extreme samples. Participants appeared to rely on their internal models to filter out noisy visualizations, thus improving their resilience against spurious data. However, participants displayed overconfidence and struggled with uncertainty estimation. They also exhibited higher variance than statistical machines. Our findings suggest that analyst gut reactions to visualizations may provide an advantage, even when departing from rationality. These results carry implications for designing visual analytics tools, offering new perspectives on how to integrate statistical models and analyst intuition for improved inference and decision-making. The data and materials for this paper are available at https://osf.io/qmfv6

human-machine collaboration↗

Selecting Appropriate Model Complexity: An Example of Tracer Inversion for Thermal Prediction in Enhanced Geothermal Systems

Abstract A major challenge in the inversion of subsurface parameters is the ill‐posedness issue caused by the inherent subsurface complexities and the generally spatially sparse data. Appropriate simplifications of inversion models are thus necessary to make the inversion process tractable and meanwhile preserve the predictive ability of the inversion results. In this study, we investigate the effect of model complexity on fracture aperture inversion and thermal performance prediction in a field‐scale EGS model. Principal component analysis was used to map the aperture field to a low‐dimensional latent space. The complexity of the inversion model was quantitatively represented by the percentage of total variance in the original aperture fields preserved by the latent space. Tracer, pressure and flow rate data were used to invert for fracture aperture through an ensemble‐based inversion method, and the inferred aperture field was used to predict thermal performance. With an over‐simplified aperture model, ensemble collapse occurred. The inverted aperture models failed to resolve necessary flow and transport features, leading to a biased thermal performance prediction. A complex aperture model involved excessive features and was prone to overinterpreting the inversion data. Both the tracer/pressure/flow rate data reproduction and thermal prediction showed significant uncertainties, making it difficult to properly estimate long‐term thermal performance. Fortunately, our results indicate that there exists an appropriate model complexity which can simultaneously match inversion data and predict thermal performance with an acceptable uncertainty. The quality of the fit of tracer data appears to be a useful indicator of such an appropriate model complexity.

15 GEOTHERMAL ENERGY↗

Validation of the DESI DR2 Ly$α$ forest full-shape analysis

We present the validation of the Dark Energy Spectroscopic Instrument (DESI) Data Release 2 (DR2) Lyman-$α$ (Ly$α$) forest full-shape analysis. This analysis combines three-dimensional Ly$α$ forest auto-correlations and cross-correlations with quasars to extract information from both the baryon acoustic oscillation (BAO) feature and the broadband clustering signal, with primary emphasis on the Alcock-Paczynski (AP) measurement. Compared to the DESI DR1 analysis, the DR2 validation uses substantially larger and more realistic mock datasets, including CoLoRe 2LPT and AbacusSummit Ly$α$ forest simulations. The modeling framework is also improved through analytic marginalization over small scales ($<10$$h^{-1}$Mpc) and the impact of ultraviolet background fluctuations. The validation program was completed prior to unblinding and defines quantitative requirements for the cosmological parameters of interest, which are evaluated using hundreds of mock realizations. We further test the analysis through independent fits to the auto- and cross-correlations, multiple catalog splits, and a broad suite of analysis and modeling variations applied to both mocks and blinded observational data. We find that the BAO and AP parameters satisfy all validation requirements and remain stable across all tests. In contrast, mock studies reveal a significant bias in the inferred growth-rate parameter $fσ_8$, leading us to exclude this measurement from the final analysis. The consistency across mocks, data splits, and robustness tests demonstrates that the DR2 Ly$α$ full-shape analysis provides a reliable and substantially improved broadband AP measurement over previous Ly$α$ forest studies.

Herbold, M. [Chicago U., KICP; Ohio State U.] (ORC↗

SPARTA: A flux adjustment methodology to interpret complex experiments

For the accurate determination of reactivity from a detector count rate, correction of spatial effects is of prime importance. This spatial correction is often provided using simulation methodologies, but this may introduce a bias if the result of the experiment is also used as input data for the simulation. Here, this work presents a flux adjustment methodology able to infer experimental reactivity and correction of spatial effects without the need for a simulation. It can process the signal from a complex experiment such as a heat balance measurement in the TREAT reactor, where control rods are continuously adjusted to maintain a constant power. In the present work, this methodology successfully computed the reactivity and the local spatial variation of the flux of a generated signal. It also proved to be robust against noise and errors on kinetic parameters and provides a credible interpretation of a heat balance experiment in TREAT. Efficiency of flux adjustment methods for complex experiment enable a better experiment interpretation less reliant on nuclear data evaluation.

73 - NUCLEAR PHYSICS AND RADIATION PHYSICS↗

The DECADE cosmic shear project III: validation of analysis pipeline using spatially inhomogeneous data

We present the pipeline for the cosmic shear analysis of the Dark Energy Camera All Data Everywhere (DECADE) weak lensing dataset: a catalog consisting of 107 million galaxies observed by the Dark Energy Camera (DECam) in the northern Galactic cap. The catalog derives from a large number of disparate observing programs and is therefore more inhomogeneous across the sky compared to existing lensing surveys. First, we use simulated data-vectors to show the sensitivity of our constraints to different analysis choices in our inference pipeline, including sensitivity to residual systematics. Next we use simulations to validate our covariance modeling for inhomogeneous datasets. Finally, we show that our choices in the end-to-end cosmic shear pipeline are robust against inhomogeneities in the survey, by extracting relative shifts in the cosmology constraints across different subsets of the footprint/catalog and showing they are all consistent within 1σ to 2σ. This is done for forty-six subsets of the data and is carried out in a fully consistent manner: for each subset of the data, we re-derive the photometric redshift estimates, shear calibrations, survey transfer functions, the data vector, measurement covariance, and finally, the cosmological constraints. Our results show that existing analysis methods for weak lensing cosmology can be fairly resilient towards inhomogeneous datasets. This also motivates exploring a wider range of image data for pursuing such cosmological constraints.

79 ASTRONOMY AND ASTROPHYSICS↗

Data-driven projection pursuit adaptation of polynomial chaos expansions for dependent high-dimensional parameters

Uncertainty quantification (UQ) and inference involving a large number of parameters are valuable tools for problems associated with heterogeneous and non-stationary behaviors. The difficulty with these problems is exacerbated when these parameters are statistically dependent requiring statistical characterization over joint measures. Probabilistic modeling methodologies stand as effective tools in the realms of UQ and inference. Among these, polynomial chaos expansions (PCE), when adapted to low-dimensional quantities of interest (QoI), provide effective yet accurate approximations for these QoI in terms of an adapted orthogonal basis. These adaptation techniques have been cast as projection pursuits in Gaussian Hilbert space in what has been referred to as a projection pursuit adaptation (PPA) by Xiaoshu Zeng and Roger Ghanem (2023). The PPA method efficiently identifies an optimal low-dimensional space for representing the QoI and simultaneously evaluates an optimal PCE within that space. The quality of this approximation clearly depends on the size of the training dataset, which is typically a function of the adapted reduced dimension. Here, the complexity of the problem is thus mediated by the complexity of the low-dimensional quantity of interest and not the complexity of the high-dimensional parameter space.

Data-driven↗

Forecasting Dark Matter Subhalo Constraints from Stellar Streams using Implicit Likelihood Inference

The evidence for dark matter (DM) remains compelling, although attempts to understand its particle nature remain inconclusive. One promising method to study DM is detecting DM subhalos through their gravitational interactions with stellar streams. In this study, we apply Neural Posterior Estimation (NPE) to constrain subhalo interaction parameters, including mass, scale radius, velocity, and encounter geometry, from stellar stream kinematics. We generate particle spray simulations based on the Lagrange Cloud stripping technique, focusing on the ATLAS-Aliqa Uma stream as a test case. We train multiple NPE models across multiple observational scenarios, quantifying how kinematic completeness affects inference and forecasting constraints from upcoming surveys including LSST, 4MOST, and 10-year Gaia data. Our results demonstrate that NPE can produce accurate and well-calibrated posteriors. In the idealized case with full 6D coordinates, we achieve subhalo mass uncertainties of 15-20% for a $10^7 \, \mathrm{M_\odot}$ subhalo, with 5D coordinates (excluding radial velocities) achieving similar performance. Under realistic observational conditions, mass uncertainties range from 50% (present-day) to 20-40% (future scenarios), with comparable performance between the photometric-only LSST sample and a smaller sample that includes Gaia proper motions and 4MOST radial velocities. Most notably, we find that velocity bimodality emerges when phase space is poorly sampled, whether due to missing kinematic information or limited stellar tracers. Combining large photometric samples with targeted spectroscopic follow-up can effectively resolves this degeneracy. These results demonstrate the power of implicit likelihood inference for optimizing stellar stream observational strategies and forecasting DM subhalo constraints from upcoming surveys.

Nguyen, Tri [Northwestern U. (main); SkAI, Chicago↗

Measurement of the time-integrated 𝐶⁢𝑃 asymmetry in 𝐷 0 → 𝐾$^{0}_{S}$𝐾$^{0}_{S}$ decays using opposite-side flavor tagging at Belle and Belle II

We measure the time-integrated 𝐶⁢𝑃 asymmetry in 𝐷 0 → 𝐾$^{0}_{S}$𝐾$^{0}_{S}$ decays reconstructed in 𝑒 + ⁢𝑒 − → $c\bar{c}$ events collected by the Belle and Belle II experiments. The corresponding data samples have integrated luminosities of 980 and 428 fb −1 , respectively. To infer the flavor of the 𝐷 0 meson, we exploit the correlation between the flavor of the reconstructed decay and the electric charges of particles reconstructed in the rest of the 𝑒 + ⁢𝑒 − → $c\bar{c}$ event. This results in a sample which is independent from any other previously used at Belle or Belle II. The result, 𝐴 𝐶⁢𝑃 ⁡(𝐷 0 → 𝐾$^{0}_{S}$𝐾$^{0}_{S}$)=(1.3±2.0±0.2)%, where the first uncertainty is statistical and the second systematic, is consistent with previous determinations and with 𝐶⁢𝑃 symmetry.

CP violation↗

Characterization of DESI fiber assignment incompleteness effect on 2-point clustering and mitigation methods for DR1 analysis

We present an in-depth analysis of the fiber assignment incompleteness in the Dark Energy Spectroscopic Instrument (DESI) Data Release 1 (DR1). This incompleteness is caused by the restricted mobility of the robotic fiber positioner in the DESI focal plane, which limits the number of galaxies that can be observed at the same time, especially at small angular separations. As a result, the observed clustering amplitude is suppressed in a scale-dependent manner, which, if not addressed, can severely impact the inference of cosmological parameters. We discuss the methods adopted for simulating fiber assignment on mocks and data. In particular, we introduce the fast fiber assignment (FFA) emulator, which was employed to obtain the power spectrum covariance adopted for the DR1 full-shape analysis. We present the mitigation techniques, organised in two classes: measurement stage and model stage. We then use high fidelity mocks as a reference to quantify both the accuracy of the FFA emulator and the effectiveness of the different measurement-stage mitigation techniques. This complements the studies conducted in a parallel paper for the model-stage techniques, namely the θ-cut approach. We find that pairwise inverse probability (PIP) weights with angular upweighting recover the “true” clustering in all the cases considered, in both Fourier and configuration space. Notably, we present the first ever power spectrum measurement with PIP weights from real data.

cosmological simulations↗

Real-Time event reconstruction for Nuclear Physics Experiments using Artificial Intelligence

Charged track reconstruction is a critical task in nuclear physics experiments, enabling the identification and analysis of particles produced in high-energy collisions. Machine learning (ML) has emerged as a powerful tool for this purpose, addressing the challenges posed by complex detector geometries, high event multiplicities, and noisy data. Traditional methods rely on pattern recognition algorithms like the Kalman filter, but ML techniques, such as neural networks, graph neural networks (GNNs), and recurrent neural networks (RNNs), offer improved accuracy and scalability. By learning from simulated and real detector data, ML models can identify and classify tracks, predict trajectories, and handle ambiguities caused by overlapping or missing hits. Moreover, ML-based approaches can process data in near-real-time, enhancing the efficiency of experiments at large-scale facilities like the Large Hadron Collider (LHC) and Jefferson Lab (JLAB). As detector technologies and computational resources evolve, ML-driven charged track reconstruction continues to push the boundaries of precision and discovery in nuclear physics. In these proceedings, we highlight advancements in charged track identification leveraging Artificial Intelligence within the CLAS12 detector, achieving a notable enhancement in experimental statistics compared to traditional methods. Additionally, we showcase real-time event reconstruction capabilities, including the inference of charged particle properties, such as momentum, direction, and species identification, at speeds matching data acquisition rates. These innovations enable the extraction of physics observables directly from the experiment in real-time.

Gavalian, Gagik (ORCID:0000000267385457)↗

Constraining Hamiltonians from chiral effective field theory with neutron-star data

Multi-messenger observations of neutron stars (NSs) and their mergers have placed strong constraints on the dense-matter equation of state (EOS). The EOS, in turn, depends on microscopic nuclear interactions that are described by nuclear Hamiltonians. These Hamiltonians are commonly derived within chiral effective field theory (EFT). Ideally, multi-messenger observations of NSs could be used to directly inform our understanding of EFT interactions, but such a direct inference necessitates millions of model evaluations. This is computationally prohibitive because each evaluation requires us to calculate the EOS from a Hamiltonian by solving the quantum many-body problem with methods such as auxiliary-field diffusion Monte Carlo (AFDMC), which provides very accurate and precise solutions but at a significant computational cost. Additionally, we need to solve the stellar structure equations for each EOS which further slows down each model evaluation by a few seconds. In this work, we combine emulators for AFDMC calculations of neutron matter, built using parametric matrix models, and for the stellar structure equations, built using multilayer perceptron neural networks, with the PyCBC data-analysis framework to enable a direct inference of coupling constants in an EFT Hamiltonian using multi-messenger observations of NSs. We find that astrophysical data can provide informative constraints on two-nucleon couplings despite the high densities probed in NS interiors.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

IntraShuffler: A Privacy Preserving Framework for Heterogeneous DP Federated Learning

Heterogeneous Differential Privacy (HDP) in Federated Learning (FL) allows clients to select individual privacy budgets () according to institutional policies and data sensitivity. In practice, many HDP-FL systems employ -aware server aggregation to improve model utility by re-weighting client updates according to their declared privacy budgets. However, gradient updates in FL retain structural patterns induced by non-independent and identically-distributed (non-IID) data, and these additional signals exposed by -aware aggregation create new opportunities for inference by an honest-but-curious server. In this work, we first show that a server equipped with gradient denoising and surrogate modeling can mount a Privacy Inference Attack that infers distributional attributes of clients and links updates from the same client across training rounds, measured via surrogate inference accuracy and linkage success, under realistic knowledge constraints. The Shuffle-Model has been widely studied as a defense against such inference risks by anonymizing update sources, but it is fundamentally incompatible with HDP-FL -aware aggregation. To address this challenge, we propose IntraShuffler, a middleware defense framework designed for HDP-FL systems. IntraShuffler introduces a privacy-aware shuffling mechanism that groups clients into privacy-compatible buckets and performs parameter-level shuffling within each bucket to disrupt persistent gradient structure while preserving -aware aggregation. Experiments across four different datasets show that IntraShuffler reduces gradient recoverability by over 60% and decreases surrogate inference accuracy from 0.78 to 0.33 while maintaining comparable model utility across multiple FL aggregation rules.

Riya, Farhin Farhad [ORNL]↗

Optimal Transport as a Tool for Scientific Discovery in Radiation Biology

This report summarizes findings from research conducted for the “Exploration of the Poten tial for Artificial Intelligence and Machine Learning to Advance Low-Dose Radiation Biology Re search” (RadBio-AI) program, supported by the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research, under Awards KP1601011/FWP CC121 and KP1601017/FWP CC121. The research reported here was undertaken in an effort to assess the potential of optimal measure transport methods as components within the larger scope of a com putational framework envisioned to support research in the radiation biology domain. Within this effort, our interest centered on enabling a unified generic framework where probabilistic modeling, inference, and statistical learning can be carried out for a wide range of data distributions. As described next in Section 1 (and in more detail in our original publication), optimal measure transport offers the possibility of such unified approach.

97 MATHEMATICS AND COMPUTING↗

Simulation-Based Inference for Neutrino Interaction Model Parameter Tuning

High-energy physics experiments studying neutrinos rely heavily on simulations of their interactions with atomic nuclei. Limitations in the theoretical understanding of these interactions typically necessitate ad hoc tuning of simulation model parameters to data. Traditional tuning methods for neutrino experiments have largely relied on simple algorithms for numerical optimization. While adequate for the modest goals of initial efforts, the complexity of future neutrino tuning campaigns is expected to increase substantially, and new approaches will be needed to make progress. In this paper, we examine the application of simulation-based inference (SBI) to the neutrino interaction model tuning for the first time. Using a previous tuning study performed by the MicroBooNE experiment as a test case, we find that our SBI algorithm can correctly infer the tuned parameter values when confronted with a mock data set generated according to the MicroBooNE procedure. This initial proof-of-principle illustrates a promising new technique for next-generation simulation tuning campaigns for the neutrino experimental community.

Tame-Narvaez, Karla Maria [Fermilab]↗

Simulation-based inference for neutrino interaction model parameter tuning

High-energy physics experiments studying neutrinos rely heavily on simulations of their interactions with atomic nuclei. Limitations in the theoretical understanding of these interactions typically necessitate ad hoc tuning of simulation model parameters to data. Traditional tuning methods for neutrino experiments have largely relied on simple algorithms for numerical optimization. While adequate for the modest goals of initial efforts, the complexity of future neutrino tuning campaigns is expected to increase substantially, and new approaches will be needed to make progress. In this paper, we examine the application of simulation-based inference (SBI) to the neutrino interaction model tuning for the first time. Using a previous tuning study performed by the MicroBooNE experiment as a test case, we find that our SBI algorithm can correctly infer the tuned parameter values when confronted with a mock data set generated according to the MicroBooNE procedure. This initial proof-of-principle illustrates a promising new technique for next-generation simulation tuning campaigns for the neutrino experimental community.

Tame-Narvaez, Karla [Fermilab] (ORCID:000000022249↗