Search NASASearch

SEARCH · Search NASA

Results for “Data Inference”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Domain-Adaptive Neural Posterior Estimation for Strong Gravitational Lens Analysis

Modeling strong gravitational lenses is prohibitively expensive for modern and next-generation cosmic survey data. Neural posterior estimation (NPE), a simulation-based inference (SBI) approach, has been studied as an avenue for efficient analysis of strong lensing data. However, NPE has not been demonstrated to perform well on out-of-domain target data -- e.g., when trained on simulated data and then applied to real, observational data. In this work, we perform the first study of the efficacy of NPE in combination with unsupervised domain adaptation (UDA). The source domain is noiseless, and the target domain has noise mimicking modern cosmology surveys. We find that combining UDA and NPE improves the accuracy of the inference by 1-2 orders of magnitude and significantly improves the posterior coverage over an NPE model without UDA. We anticipate that this combination of approaches will help enable future applications of NPE models to real observational data.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Homomorphic Encryption for Electrical Metering Aggregation: Protecting the Privacy of Building Tenants

Electrical meters are devices that measure consumer electricity usage. The data collected by these meters is necessary for utility billing and electrical grid management but can also be used to assess the environmental impact of buildings. Prior research has found that unprotected metering data could potentially be used to infer some information about the behaviors of building tenants by detecting changes in electricity usage. For example, a period of low electricity usage could suggest that the tenants are not in the building. As smart metering becomes more common, there is a growing need for data privacy protections for metering data that do not negatively impact the quality and availability of data used for energy management and billing applications. To identify potential solutions, we developed a Python-based data aggregation platform to analyze the potential efficacy of privacy-enhancing technologies for energy metering applications. This platform aggregates groups of metering sites into virtual buildings, which could potentially detach changes in electrical activity from individual tenants, making it more difficult to track the activity of a specific tenant. To further protect data during analysis, this project utilizes homomorphic encryption as part of its initial approach. Homomorphic encryption offers a means of protecting energy consumption data while permitting mathematical operations to be performed without the need to know the data contents. This allows for data to be processed into usable statistics without revealing energy consumption information. A series of homomorphic encryption libraries were evaluated to determine their applicability and limitations in the context of metering data. The use of these techniques may help to reassure consumers and encourage further adoption of smart grid infrastructure.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

BatFIT (Battery Feature Inference Toolbox) [SWR-26-034]

This package implements several data-based techniques for parameter fitting in Li-ion battery models. It uses BATMODS-lite to generate the data . The repository contains the code that is used for the paper "Neural posterior estimation is accurate, tractable and scalable for inverse parameter inference in Li-ion batteries", M. Hassanaly, C. R. Randall, P. J. Weddle, P. J. Gasper, C. Kelly, T. R. Tanim, K. Smith.

Hassanaly, Malik [National Laboratory of the Rocki

Overview of oxygen opacity experiments at the National Ignition Facility and investigation of potential systematic errors

Experiments to measure oxygen opacity at stellar interior conditions have been performed at the National Ignition Facility in a Discovery Science campaign. These experiments utilize the Opacity-on-NIF platform with a sample comprised of O, Mg, and Si. The spectral data from the Opacity Spectrometer cover the 1000–2000 eV photon energy range showing bound-free continuum absorption from O and line absorption from Mg and Si. DANTE and the Gated X-ray Detector are employed to measure the sample plasma’s temperature and density, respectively. Initial data show lower transmission than expected by theoretical models, raising questions of whether potential background or data uniformity concerns could produce systematic errors in the inferred transmission. Here, we investigate three concerns thought to be important for the oxygen opacity data, including instrumental scattered background, sample self-emission non-uniformity, and backlight continuum non-uniformity. Additionally, we show the effect of a recently developed method to account for 2nd order crystal reflection. The total effect of these concerns on one experiment is found to be small compared to the observed difference between the inferred transmission and a model calculation at the inferred temperature and density. Thus, we conclude that these potential sources of systematic error cannot account for the observed difference, increasing the likelihood of a real effect due to the high temperature and density conditions. However, because this is only a single experiment, we cannot make a firm conclusion. More experiments measuring the opacity and necessary calibrations are needed to assess the reproducibility and uncertainty of this result.

79 ASTRONOMY AND ASTROPHYSICS

Dark Energy Survey Year 3 results: optimized $w$CDM simulation-based inference with weak lensing map-level hybrid statistics

We present cosmological constraints from the Dark Energy Survey Year 3 (DES Y3) weak lensing data using hierarchical hybrid statistics within a Bayesian simulation-based inference framework that is based on the Gower Street simulations. To maximize the precision of the inference, we have developed a new, information-theory based, data compression of the weak lensing maps to just seven highly informative summary statistics. The hybrid scheme exploits the high information content of the power spectrum, compressing both the power spectrum and neural-based summaries that are designed to extract further information. Our simulation-based approach enables principled forward modelling of all major sources of systematic uncertainty and survey properties into realistic mock observations, including the survey mask, photometric redshift uncertainties, intrinsic galaxy alignments, multiplicative shear calibration bias, source galaxy clustering, non-Gaussian shape noise, and non-linear structure formation. The summary statistics are then used in a Bayesian simulation-based inference pipeline. The inference is validated through coverage tests and checks for robustness against baryonic feedback. Assuming a $w$CDM cosmology, our analysis yields $S_8 = 0.808 \pm 0.017$, $Ω_{\rm m} = 0.325 \pm 0.024$, and $w < -0.766$ (marginalized posterior 68 per cent credible intervals). This rigorous combination of information theory, physics- and neural network-based extreme data compression, and principled Bayesian analysis improves the figure of merit for $(Ω_{\rm m}, S_8, w)$ by 60 per cent over the previous state-of-the-art, and by almost a factor of 3 over two-point analyses of the same data. They are the most precise joint constraints on $(Ω_{\rm m}, S_8, w)$ from weak gravitational lensing data alone of any survey to date. We intend to apply this analysis to the more recent DES Y6 data.

Williamson, J. [University Coll. London]

Revealing Local Structures through Machine-Learning-Fused Multimodal Spectroscopy

Atomistic structures of materials offer valuable insights into their functionality. Determining these structures remains a fundamental challenge in materials science, especially for systems with defects. While both experimental and computational methods exist, each has limitations in resolving nanoscale structures. Core-level spectroscopies, such as X-ray absorption (XAS) or electron energy-loss spectroscopies (EELS), have been used to determine the local bonding environment and structure of materials. Recently, machine learning (ML) methods have been applied to extract structural and bonding information from XAS/EELS data. However, frameworks relying solely on a single data stream, defined as characterization data derived from a single element using one technique, are often insufficient because multiple local environments can yield similar spectral features, making it challenging to differentiate between competing structural hypotheses. Here, in this work, we address this challenge by integrating multimodal ab initio simulations, experimental data acquisition, and ML techniques for structure characterization. Our goal is to determine local structures and properties using EELS and XAS data from multiple elements and edges. To showcase our approach, we use various lithium nickel manganese cobalt (NMC) oxide compounds which are used for lithium ion batteries, including those with oxygen vacancies and antisite defects, as the sample material system. We successfully inferred local element content, ranging from lithium to transition metals, with quantitative agreement with experimental data. Beyond local element inference, we find that ML model based on multimodal spectroscopic data is able to determine whether local defects such as oxygen vacancy and antisites are present, a task which is impossible for single mode spectra or other experimental techniques. Furthermore, our framework is able to provide physical interpretability, bridging spectroscopy with the local atomic and electronic structures.

battery

Accelerating data acquisition with FPGA-based edge machine learning: a case study with LCLS-II

New scientific experiments and instruments generate vast amounts of data that need to be transferred for storage or further processing, often overwhelming traditional systems. Edge machine learning (EdgeML) addresses this challenge by integrating machine learning (ML) algorithms with edge computing, enabling real-time data processing directly at the point of data generation. EdgeML is particularly beneficial for environments where immediate decisions are required, or where bandwidth and storage are limited. In this paper, we demonstrate a high-speed configurable ML model in a fully customizable EdgeML system using a field programmable gate array (FPGA). Our demonstration focuses on an angular array of electron spectrometers, referred to as the ‘CookieBox,’ developed for the Linac Coherent Light Source II project. The EdgeML system captures 51.2 Gbps from a 6.4 GS s −1 analog to digital converter and is designed to integrate data pre-processing and ML inside an FPGA. Our implementation achieves an inference latency of 0.2 µs for the ML model, and a total latency of 0.4 µs for the complete EdgeML system, which includes pre-processing, data transmission, digitization, and ML inference. The modular design of the system allows it to be adapted for other instrumentation applications requiring low-latency data processing.

97 MATHEMATICS AND COMPUTING

A score-based diffusion model approach for adaptive learning of stochastic partial differential equation solutions

In this paper, we propose a novel framework for adaptively learning the time-evolving solutions of stochastic partial differential equations (SPDEs) using score-based diffusion models within a recursive Bayesian inference setting. SPDEs play a central role in modeling complex physical systems under uncertainty, but their numerical solutions often suffer from model errors and reduced accuracy due to incomplete physical knowledge and environmental variability. To address these challenges, we encode the governing physics into the score function of a diffusion model using simulation data and incorporate observational information via a likelihood-based correction in a reverse-time stochastic differential equation. This enables adaptive learning through iterative refinement of the solution as new data becomes available. To improve computational efficiency in high-dimensional settings, we introduce the ensemble score filter, a training-free approximation of the score function designed for real-time inference. Numerical experiments on benchmark SPDEs demonstrate the accuracy and robustness of the proposed method under sparse and noisy observations.

97 MATHEMATICS AND COMPUTING

Unraveling emission line galaxy conformity at z ∼ 1 with DESI early data

Emission line galaxies (ELGs) are now the preeminent tracers of large-scale structure at z > 0.8 due to their high density and strong emission lines, which enable accurate redshift measurements. However, relatively little is known about ELG evolution and the ELG–halo connection, exposing us to potential modelling systematics in cosmology inference using these sources. In this paper, we use a variety of observations and simulated galaxy models to propose a physical picture of ELGs and improve ELG–halo connection modelling in a halo occupation distribution framework. We investigate Dark Energy Spectroscopic Instrument (DESI)-selected ELGs in COSMOS data, and infer that ELGs are rapidly star-forming galaxies with a large fraction exhibiting disturbed morphology, implying that many of them are likely to be merger-driven starbursts. We further postulate that the tidal interactions from mergers lead to correlated star formation in central–satellite ELG pairs, a phenomenon dubbed ‘conformity’. We argue for the need to include conformity in the ELG–halo connection using galaxy models such as IllustrisTNG, and by combining observations such as the DESI ELG autocorrelation, ELG cross-correlation with luminous red galaxies, and ELG–cluster cross-correlation. We also explore the origin of conformity using the UniverseMachine model and elucidate the difference between conformity and the well-known galaxy assembly bias effect.

79 ASTRONOMY AND ASTROPHYSICS

Elemental and isotopic signatures of Asteroid Ryugu support three early Solar System reservoirs

Understanding the number and locations of different reservoirs present in the early Solar System is crucial to understanding the Solar System’s origin and evolution. Previous work has suggested that three unique isotopic reservoirs existed in the early Solar System but subsequent works have challenged that idea. Here we present elemental abundances along with Ca, Ti, Cr, Fe, Ni, and Zn isotopic data from primitive material returned by the Japan Aerospace Exploration Agency’s (JAXA) Hayabusa2 mission to asteroid (162173) Ryugu to make inferences on the Solar System’s early architecture. Data from Ryugu particle A0208 are consistent with a close genetic heritage between Ryugu and CI chondrites. Here, we employ principal component analysis (PCA) on these Ryugu and published meteorite data to demonstrate that Ryugu and CI chondrites are distinct from other known astromaterials, strongly supporting the existence of a third major isotopic reservoir in the early Solar System.

Isotopes

Evaluation of Drilling Performance at The Geysers with Machine Learning Methods Using Geologic Data

A recent well, GDC-36, was drilled in The Geysers Geothermal Field served in a Department of Energy-industry to demonstrate improved drilling performance with polycrystalline diamond compact (PDC) bits. Both PDC and roller cone drill bits were used to drill this well. Key challenges encountered during drilling included lost circulation in the mud-drilled section, and bit damage interfacial severity in the deeper, air-drilled section. The objective of this study is to evaluate the drilling performance in relation to the local geological characteristics using machine learning methods. By applying K-clustering to the sonic log data, we were able to identify areas correlated with measured lost circulation. Also, the boundaries defined by clustering of the mineralogical and lithological data from the mud logs correlate well with interfacial severity during drilling. A random forest model was employed to build correlation between drilling data and rock strength. The confined compressive strength (CCS) of the rock in the training of the machine learning model was inferred from the dipole sonic log. The R-squared of the testing data is 0.78, and the RMSE (Root Mean Squared Error) is 0.06. The trained model was used to forecast rock strength for the section where sonic log data are not available. CCS could also be inferred from mud logs provided the relationship between mineralogy and rock strength is established through core testing data.

15 GEOTHERMAL ENERGY

Combining Observations and Models: A Review of the CARDAMOM Framework for Data‐Constrained Terrestrial Ecosystem Modeling

The rapid increase in the volume and variety of terrestrial biosphere observations (i.e., remote sensing data and in situ measurements) offers a unique opportunity to derive ecological insights, refine process‐based models, and improve forecasting for decision support. However, despite their potential, ecological observations have primarily been used to benchmark process‐based models, as many past and current models lack the capability to directly integrate observations and their associated uncertainties for parameterization. In contrast, data assimilation frameworks such as the CARbon DAta MOdel fraMework (CARDAMOM) and its suite of process‐based models, known as the Data Assimilation Linked Ecosystem Carbon Model (DALEC), are specifically designed for model‐data fusion. This review, motivated by a recent CARDAMOM community workshop, examines the development and applications of CARDAMOM, with an emphasis on its role in advancing ecosystem process understanding. CARDAMOM employs a Bayesian approach, using a Markov Chain Monte Carlo algorithm to enable data‐driven calibration of DALEC parameters and initial states (i.e., carbon pool sizes) through observation operators. CARDAMOM's unique ability to retrieve localized model process parameters from diverse datasets—ranging from in situ measurements to global satellite observations—makes it a highly flexible tool for analyzing spatially variable ecosystem responses to environmental change. However, assimilating these data also presents challenges, including data quality issues that propagate into model skill, as well as trade‐offs between model complexity, parameter equifinality, and predictive performance. We discuss potential solutions to these challenges, such as reducing parameter equifinality by incorporating new observations. This review also offers community recommendations for incorporating emerging datasets, integrating machine learning techniques, strengthening collaboration with remote sensing, field, and modeling communities, and expanding CARDAMOM's relevance for localized ecosystem monitoring and decision‐making. CARDAMOM enables a deep, mechanistic understanding of terrestrial ecosystem dynamics that cannot be achieved through empirical analyses of observational datasets or weakly constrained models alone.

Bayesian inference

Materials Learning Algorithms (MALA): Scalable machine learning for electronic structure calculations in large-scale atomistic simulations

We present the Materials Learning Algorithms (MALA) package, a scalable machine learning framework designed to accelerate density functional theory (DFT) calculations suitable for large-scale atomistic simulations. Using local descriptors of the atomic environment, MALA models efficiently predict key electronic observables, including local density of states, electronic density, density of states, and total energy. The package integrates data sampling, model training and scalable inference into a unified library, while ensuring compatibility with standard DFT and molecular dynamics codes. We demonstrate MALA's capabilities with examples including boron clusters, aluminum across its solid-liquid phase boundary, and predicting the electronic structure of a stacking fault in a large beryllium slab. Scaling analyses reveal MALA's computational efficiency and identify bottlenecks for future optimization. With its ability to model electronic structures at scales far beyond standard DFT, MALA is well suited for modeling complex material systems, making it a versatile tool for advanced materials research.

Density functional theory

MS25: Materials Science-Focused Benchmark Data Set for Machine Learning Interatomic Potentials

Here, we present MS25, a benchmark data set for evaluating machine learning interatomic potentials (MLIPs) across diverse materials-relevant systems including MgO surfaces, liquid water, zeolites, a catalytic Pt surface reaction, high-entropy alloys (HEAs), and disordered Zr-oxides. Five MLIP architectures (MACE, NequIP, Allegro, MTP, and Torch-ANI) are trained and tested, focusing not only on traditional metrics (energies, forces, and stresses) but also explicitly validating derived physical observables such as lattice constants, volumes, and reaction barriers. We find that most models reach comparable accuracy on standard error metrics across the simple systems, although equivariant MLIPs offer 1.5–2× improvements over nonequivariant MLIPs in energy and force error for structurally complex or compositionally disordered environments such as HEAs and Zr–O systems. Our analysis highlights that low errors in energy and force predictions do not guarantee reliable observables, emphasizing the necessity of explicit validation. We demonstrate limitations in cross-framework transferability, as models trained on one zeolite framework (CHA) fail to reliably generalize to predictions of structurally distinct frameworks (e.g., MFI). Size-extensive tests show some dependence on system size for MgO, resulting from forced periodicity. The HEA and Zr–O data sets are identified as challenging tests for future benchmarks and MLIP model architecture developments as they show significant differentiation in error between MLIP architectures and are still relatively difficult at 1000 training images. Moving forward, we recommend that benchmarking efforts shift their focus from marginal accuracy improvements in energy and force errors toward identifying and understanding model failure modes, rigorously assessing transferability, and evaluating how their errors affect observable predictions. For researchers looking to choose an MLIP architecture, we suggest selecting equivariant MLIP architectures if the complexity of the system is a challenge. For simple materials problems, auxiliary features such as integration with molecular dynamics engines, trade-offs between computational data set generation cost vs MLIP inference speed, and framework integration may play a more important decision factor than small differences in error metrics that are unlikely to matter for production-level research.

chemical structure

A Framework for Parametric and Predictive Uncertainty Quantification in the E3SM Land Model: Assessing Site and Observable Generalizability

Quantifying parametric uncertainty using observations from individual sites provides a critical foundation for Earth system modeling, serving as a necessary first step before scaling up to regional or global applications. This study introduces a novel computational framework designed to enhance model predictability by reducing parametric uncertainty and assessing site and observable generalizability using various observational constraints. The framework integrates five components: Model Simulation, Statistical Emulation, Global Sensitivity Analysis (GSA), Model Calibration, and Model Prediction. Using the E3SM land model, we simulated site-level land-atmosphere carbon and energy fluxes from 2003 to 2007 across five evergreen needleleaf FLUXNET sites, perturbing 26 vegetation-related model parameters. Gaussian process emulators were employed to expedite GSA and model calibration. Four critical parameters that strongly influence selected land-atmosphere fluxes were identified by GSA. Bayesian approaches were used to infer parameter probability distributions leveraging synthetic data and FLUXNET observations. The results reveal that posterior parameter distributions vary significantly across different sites and observables within the same plant functional type. Probabilistic predictions indicate that parameters calibrated at one site can enhance predictive accuracy at other sites, although site heterogeneity may sometimes outweigh parametric uncertainty. Additionally, the probabilistic predictions demonstrate that calibration for one variable can also improve predictability for other variables, thereby maximizing predictive capabilities with limited observations. This framework provides a powerful approach for reducing parametric uncertainty in Earth system models and deepening our understanding of carbon dynamics and energy cycles. Its adaptability makes it a valuable tool for broader applications in Earth system modeling.

54 ENVIRONMENTAL SCIENCES

Mapping causal patterns in crystalline solids

The evolution of the atomic structures of the combinatorial library of Sm-substituted thin film BiFeO 3 along the phase transition boundary from the ferroelectric rhombohedral phase to the non-ferroelectric orthorhombic phase is explored using scanning transmission electron microscopy. Localized properties, including polarization, lattice parameter, and chemical composition, are parameterized from atomic-scale imaging, and their causal relationships are reconstructed using a linear non-Gaussian acyclic model. This approach is further extended to explore the spatial variability of the causal coupling using the sliding window transform method, which revealed that new causal relationships emerged at both the expected locations, such as domain walls and interfaces, and at additional regions forming clusters in the vicinity of the walls or spatially distributed features. While the exact physical origins of these relationships are unclear, they likely represent nanophase-separated regions in the morphotropic phase boundaries. Overall, we posit that an in-depth understanding of complex disordered materials away from thermodynamic equilibrium necessitates understanding not only the generative processes that can lead to observed microscopic states but also the causal links between multiple interacting subsystems.

Causal inference

Measurements of the polarization of several instabilities in the DIII-D tokamak

Recently, a method to infer the polarization of modes with frequencies much less than the ion cyclotron frequency was published [X.D. Du et al., Phys. Rev. Lett. 132 (2024) 215101]. The method uses measurements of electron temperature and density fluctuations δTₑ and δnₑ at the same spatial position to infer the local ratio of “acoustic polarization,” |δϕ ∥ |/(|δϕ ∥ |+|δψ|), where δϕ ∥ is the effective parallel potential and δψ is related to the parallel magnetic vector potential A ∥ . This paper summarizes key formulas, with emphasis on their range of validity, and elaborates on the workflow required to infer the acoustic polarization from experimental data. The drift-acoustic polarization of ellipticity-induced, toroidicity-induced, and reversed shear Alfvén eigenmodes is nearly zero, as expected for modes with predominately shear-Alfvénic polarization. The polarization of beta-induced Alfvén eigenmodes contains an acoustic component that increases with poloidal wave number. “Low frequency modes,” (instabilities that appear transiently when the minimum of the safety factor qₘᵢₙ passes through rational values) have large and highly variable acoustic polarization. In both experiment and simulation, fishbones have non-zero acoustic polarization that increases as the mode chirps down in frequency.

Alfven eigenmode

Modeling the Cosmological Lyman-𝛼 Forest at the Field Level

The distribution of absorption lines in the spectra of distant quasars, called the Lyman-𝛼 (Ly-𝛼) forest, is a unique probe of cosmology and the intergalactic medium at high redshifts and small scales. The statistical power of ongoing redshift surveys demands precise theoretical tools to model the Ly-𝛼 forest. We address this challenge by developing an analytic, perturbative forward model to predict the Ly-𝛼 forest at the field level for a given set of cosmological initial conditions. Our model shows a remarkable performance when compared with the Sherwood hydrodynamic simulations: it reproduces the Ly-𝛼 forest flux power spectrum, its cross-correlation with dark matter halos, and the one-point probability distribution function of both fields at the percent level down to scales of a few Mpc. Our work provides crucial tools that bridge analytic modeling on large scales with simulations on small scales, enabling field-level inference from Ly-𝛼 forest data and simulation-based priors for cosmological analyses. Furthermore, this is especially timely for realizing the full scientific potential of the Ly-𝛼 forest measurements by the dark energy spectroscopic instrument.

Cosmological parameters