Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

Integrated machine learning-molecular dynamics framework for electrolyte property prediction

Electrochemical stability windows determine the operating range of battery electrolytes, yet accurate prediction remains challenging because stability emerges from statistical ensembles of local solvation environments rather than single ground-state molecular structures. Traditional density functional theory calculations on energy-minimized clusters cannot capture the thermal variations in local coordination environments and geometries that govern decomposition, while SMILES-based machine learning methods lack explicit representation of three-dimensional solvation structure and ion pairing. Here, we introduce a structure-aware machine learning framework that predicts frontier orbital energies (HOMO and LUMO) directly from molecular dynamics-sampled solvation configurations, achieving sub-0.6 eV accuracy at computational costs 3–4 orders of magnitude lower than first-principles methods. Across twelve representative battery electrolytes, we demonstrate that solvent-separated and contact ion pairs exhibit strong size- and local chemistry dependent electronic stability, with variations in coordination shifts of HOMO or LUMO level by 2–3 eV, and that extended solvation structure and partially desolvated environment further modulate stability by up to 3 eV. By encoding the statistical nature of electrochemical failure through ensemble sampling of explicit solvation geometries, our approach enables high-throughput screening and rational design of next-generation battery electrolytes with mechanistic understanding of structure–property relationships.

Energy - Storage↗

A Computational Review of Privacy-Preserving Mechanisms for the Smart Grid

Smart grid technologies have rapidly become one of the largest and most comprehensive sources of data for the modern utility. For the most part, data streams are seen as an essential tool that enable utilities to carry their day-to-day business operations, but they also create the need for efficient and secure data management strategies. In the context of the smart grid, ensuring data privacy is becoming an increasing concern due to a combination of factors that range from shifts in operational paradigms and rapid technology evolution to changes in legislation. Furthermore, researchers have highlighted the risks associated with improperly protected energy records. For example, energy consumption data from homes could be used to infer the behaviors and habits of home occupants through activity recognition or user profiling (Fan, 2017), which may lead to unfair service pricing, targeted advertising, or other personal security violations. Similarly, Electric Vehicles’ (EVs) charging metadata could be used to reveal private information about the owner such as their payment methods, preferred charging stations, and other locational and timing information that could be used to reconstruct the vehicle owner’s behaviors. The privacy of user data, even when used for statistical analysis or machine learning training processes, also needs to be carefully considered, as an individual’s private traits may still be vulnerable if their inclusion/exclusion greatly impacts the result or could be linked to a public dataset through cross-reference. The breach of user privacy also has severe impacts for organizations that store, transmit, or work on the data in the form of diminishing the public’s trust in them while potentially incurring legal consequences (e.g., fines and suspensions under the European Union General Data Protection Regulation, Health Insurance Portability and Accountability Act, etc.). Because of these risks, several privacy-preserving mechanisms are available to help organizations comply with privacy legislations and prevent the unauthorized and malicious use of user data. In light of these concerns, this report focuses on performing a computational review of privacy-preserving mechanisms that have received a significant amount of interest in literature. It specifically focuses on 1) homomorphic encryption, 2) zero-knowledge proofs, 3) differential privacy, and 4) federated learning. It is worth noting that although many of the methods presented in this document rely on cryptographic primitives, their intent is not to provide perfect secrecy, but rather to enable users to maintain privacy, and thus they shall not be compared or equated to other constructs that are aimed to address cybersecurity constructs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

High-resolution hypernuclear decay pion spectroscopy at MAMI and future

Hypernuclear decay pion spectroscopy was established in 2012 at MAMI as a mass spectroscopy method for light hypernuclei. A monochromatic pion peak from $^4_Λ$H was successfully observed, and the Λ binding energy was determined to be B Λ = 2.157±0.005(stat.)±0.077(syst.) MeV in the 2014 run. In 2022, an upgrade experiment for $^3_Λ$H spectroscopy was conducted using a newly developed Li target. The absolute electron beam energy will be measured by the synchrotron radiation interferometry, which will be applied with the spectrometer calibration to improve the systematic error. The decay pion spectroscopy is planned to be performed at JLab Hall-C, which would significantly increase the statistics thanks to the higher energy beam, the better K + identification, and the faster data acquisition system. This experiment has been submitted as a Letter of Intent in JLab PAC51. The upgraded decay pion spectroscopy method is expected to provide new, accurate hypernuclear data, which will contribute to the advancement of our understanding of hypernuclear physics

47 OTHER INSTRUMENTATION↗

Mitigation of DESI fiber assignment incompleteness effect on two-point clustering with small angular scale truncated estimators

We present a method to mitigate the effects of fiber assignment incompleteness in two-point power spectrum and correlation function measurements from galaxy spectroscopic surveys, by truncating small angular scales from estimators. We derive the corresponding modified correlation function and power spectrum windows to account for the small angular scale truncation in the theory prediction. We validate this approach on simulations reproducing the Dark Energy Spectroscopic Instrument (DESI) Data Release 1 (DR1) with and without fiber assignment. We show that we recover unbiased cosmological constraints using small angular scale truncated estimators from simulations with fiber assignment incompleteness, with respect to standard estimators from complete simulations. Additionally, we present an approach to remove the sensitivity of the fits to high k modes in the theoretical power spectrum, by applying a transformation to the data vector and window matrix. We find that our method efficiently mitigates the effect of fiber assignment incompleteness in two-point correlation function and power spectrum measurements, at low computational cost and with little statistical loss.

79 ASTRONOMY AND ASTROPHYSICS↗

Deep inference of simulated strong lenses in ground-based surveys

The large number of strong lenses discoverable in future astronomical surveys will likely enhance the value of strong gravitational lensing as a cosmic probe of dark energy and dark matter. However, leveraging the increased statistical power of such large samples will require further development of automated lens modeling techniques. We show that deep learning and simulation-based inference (SBI) methods produce informative and reliable estimates of parameter posteriors for strong lensing systems in ground-based surveys. We present the examination and comparison of two approaches to lens parameter estimation for strong galaxy-galaxy lenses — Neural Posterior Estimation (NPE) and Bayesian Neural Networks (BNNs). We perform inference on 1-, 5-, and 12-parameter lens models for ground-based imaging data that mimics the Dark Energy Survey (DES). We find that NPE outperforms BNNs, producing posterior distributions that are more accurate, precise, and well-calibrated for most parameters. For the 12-parameter NPE model, the calibration is consistently within <10% of optimal calibration for all parameters, while the BNN is rarely within 20% of optimal calibration for any of the parameters. Similarly, residuals for most of the parameters are smaller (by up to an order of magnitude) with the NPE model than the BNN model. This work takes important steps in the systematic comparison of methods for different levels of model complexity.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Direct statistical simulation of the Lorenz96 system in model reduction approaches

Direct statistical simulation (DSS) of nonlinear dynamical systems bypasses the traditional route of accumulating statistics by lengthy direct numerical simulations by solving the equations that govern the statistics themselves. DSS suffers, however, from the curse of dimensionality as the statistics (such as correlations) generally have higher dimensions than the underlying dynamical variables. Here we investigate two approaches to reduce the dimensionality of DSS, illustrating each method with numerical experiments with the Lorenz96 dynamical system. The forms of DSS chosen here involve approximate closures at second and third order in the equal-time cumulants. We demonstrate significant reduction in computational effort that can be achieved without sacrificing the accuracy of DSS. The methods developed here can be applied to turbulent fluid and magnetohydrodynamical systems. Published by the American Physical Society 2025

Li, Kuan↗

Monte Carlo Simulations of Crystal Defects in Open Ensembles

Zero- and two-dimensional crystal defects form in open statistical ensembles, such as the grand canonical, that are usually inaccessible with conventional simulation techniques. This longstanding challenge is overcome with a new Hamiltonian Monte Carlo method that samples energy-biased gradual transformations. In conclusion, the method enables free energy calculations for nonideal point defects and the direct prediction of finite-temperature interface structures.

Grain boundaries↗

Characterizing the Uncertainty of Measurement of Traceable Isotope Ratios with Bayesian Statistical Techniques

Analytical techniques such as multicollector—inductively coupled plasma—mass spectrometry (MC-ICP-MS) are routinely employed at SRNL, other National Laboratories, and in academia to determine the precise isotopic composition of diverse natural and anthropogenic samples (e.g., rocks and nuclear materials). Quantifying and reporting uncertainty in such analyses, while regularly performed, have a rigorous statistical foundation. The Guide to the Expression of Uncertainty in Measurement 4 (GUM) outlines conventional techniques used to assess such uncertainty. As the accessibility and speed of statistical computing increase, there is a need to modernize conventional techniques. For example, Supplement 1 to the 3rd to the GUM suggests the use of approximation methods as an updated approach to the GUM.

McLarty, Ellis C.↗

Nonlinear Topological Photonics: Capturing Nonlinear Dynamics and Optical Thermodynamics

Combining multiple optical resonators or engineering dispersion of complex media has provided an effective method for demonstrating topological physics controlling photons in unprecedented ways such as unidirectional light propagation and spatially localized modes between an interface or on a corner. Further, adding nonlinear responses to those topological photonic systems has enabled achieving diverse phases of photons in both space and time, allowing for more functionalities in photonic devices that provide a new playground for studying dynamic features of nonlinear topological systems. However, most methods for describing nonlinear topological photonic systems rely on linear topological theories, making it challenging to accurately characterize the topology of nonlinear systems. Thus, substantial efforts have focused on rigorously describing nonlinear topological phases and developing effective tools to analyze nonlinear topological effects. Meanwhile, coupled multimode optical waveguides with nonlinear dynamic responses provide an excellent platform for the statistical description of photons, opening a new paradigm called “optical thermodynamics”. This review will introduce the basic concepts of nonlinear topological photonics and the recent development of theoretical approaches focusing on data-driven approaches for creating phase diagrams as well as the spectral localizer framework and the pseudospectrum method for understanding optical nonlinearities in topological systems. In addition, the new concept of optical thermodynamics will be introduced with some recent theoretical works.

Topological photonics↗

A chain stretch-based gradient-enhanced model for damage and fracture in elastomers

Similar to quasi-brittle materials, it has been recently shown that elastomers can exhibit a macroscopically diffuse damage zone that accompanies the fracture process. In this study, we introduce a stretch-based gradient-enhanced damage (GED) model that allows the fracture to localize and also captures the development of a physically diffuse damage zone. This capability contrasts with the paradigm of the phase field method for fracture, where a sharp crack is numerically approximated in a diffuse manner. Capturing fracture localization and diffuse damage in our approach is achieved by considering nonlocal effects that encompass network topology, heterogeneity, and imperfections. These considerations motivate the use of a statistical damage function dependent upon the nonlocal deformation state. From this model, fracture toughness is realized as an output. While GED models have been classically utilized for damage modeling of structural engineering materials (e.g., concrete), they face challenges when trying to capture the cascade from damage to fracture, often leading to damage zone broadening (de Borst and Verhoosel, 2016). This deficiency contributed to the popularity of the phase-field method over the GED model for elastomers and other quasi-brittle materials. Other groups have proceeded with damage-based GED formulations that prove identical to the phase-field method (Lorentz et al., 2012), but these inherit the aforementioned limitations. To address this issue in a thermodynamically consistent framework, we implement two modeling features (a nonlocal driving force bound and a simple relaxation function) specifically designed to capture the evolution of a physically meaningful damage field and the simultaneous localization of fracture, thereby overcoming a longstanding obstacle in the development of these nonlocal strain- or stretch-based approaches. Here, we discuss several numerical examples to understand the features of the approach at the limit of incompressibility, and compare them to the phase-field method as a benchmark for the macroscopic response and fracture energy predictions.

Elastomers↗

Dose-efficient automatic differentiation for ptychographic reconstruction

Ptychography, as a powerful lensless imaging method, has become a popular member of the coherent diffractive imaging family over decades of development. The ability to utilize low-dose X-rays and/or fast scans offers a big advantage in a ptychographic measurement (for example, when measuring radiation-sensitive samples), but results in low-photon statistics, making the subsequent phase retrieval challenging. Here, we demonstrate a dose-efficient automatic differentiation framework for ptychographic reconstruction (DAP) at low-photon statistics and low overlap ratio. As no reciprocal space constraint is required in this DAP framework, the framework, based on various forward models, shows superior performance under these conditions. It effectively suppresses potential artifacts in the reconstructed images, especially for the inherent periodic artifact in a raster scan. We validate the effectiveness and robustness of this method using both simulated and measured datasets.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Topological Interpretability for Deep Learning

With the growing adoption of AI-based systems across everyday life, the need to understand their decision-making mechanisms is correspondingly increasing. The level at which we can trust the statistical inferences made from AI-based decision systems is an increasing concern, especially in high-risk systems such as criminal justice or medical diagnosis, where incorrect inferences may have tragic consequences. Despite their successes in providing solutions to problems involving real-world data, deep learning (DL) models cannot quantify the certainty of their predictions. These models are frequently quite confident, even when their solutions are incorrect. This work presents a method to infer prominent features in two DL classification models trained on clinical and non-clinical text by employing techniques from topological and geometric data analysis. We create a graph of a model's feature space and cluster the inputs into the graph's vertices by the similarity of features and prediction statistics. We then extract subgraphs demonstrating high-predictive accuracy for a given label. These subgraphs contain a wealth of information about features that the DL model has recognized as relevant to its decisions. We infer these features for a given label using a distance metric between probability measures, and demonstrate the stability of our method compared to the LIME and SHAP interpretability methods. This work establishes that we may gain insights into the decision mechanism of a DL model. This method allows us to ascertain if the model is making its decisions based on information germane to the problem or identifies extraneous patterns within the data.

Spannaus, Adam↗

ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis

Objective: Electronic health record (EHR) systems contain a wealth of clinical data stored as both codified data and free-text narrative notes (NLP). The complexity of EHR presents challenges in feature representation, information extraction, and uncertainty quantification. Here, to address these challenges, we proposed an efficient Aggregated naRrative Codified Health (ARCH) records analysis to generate a large-scale knowledge graph (KG) for a comprehensive set of EHR codified and narrative features. Methods: Using data from 12.5 million Veterans Affairs patients, ARCH first derives embedding vectors and generates similarities along with associated p-values to measure the strength of relatedness between clinical features with statistical certainty quantification. Next, ARCH performs a sparse embedding regression to remove indirect linkage between features to build a sparse KG. Finally, ARCH was validated on various clinical tasks, including detecting known relationships between entity pairs, predicting drug side effects, disease phenotyping, as well as sub-typing Alzheimer’s disease patients. Results: ARCH produces high-quality clinical embeddings and KG for over 60,000 codified and narrative EHR concepts. The KG and embeddings are visualized in the R-shiny powered web-API.3 ARCH achieved high accuracy in detecting EHR concept relationships, with AUCs of 0.926 (codified) and 0.861 (NLP) for similar EHR concepts, and 0.810 (codified) and 0.843 (NLP) for related pairs. It detected drug side effects with a 0.723 AUC, which improved to 0.826 after fine-tuning. Using both codified and NLP features, the detection power increased significantly. Compared to other methods, ARCH has superior accuracy and enhances weakly supervised phenotyping algorithms’ performance. Notably, it successfully categorized Alzheimer’s patients into two subgroups with varying mortality rates. Conclusion: The proposed ARCH algorithm generates large-scale high-quality semantic representations and knowledge graph for both codified and NLP EHR features, useful for a wide range of predictive modeling tasks.

Electronic health records↗

Signal-preserving CMB component separation with machine learning

Analysis of microwave sky signals, such as the cosmic microwave background, often requires component separation using multifrequency methods, whereby different signals are isolated according to their different frequency behaviors. Many so-called blind methods, such as the internal linear combination (ILC), make minimal assumptions about the spatial distribution of the signal or contaminants, and only assume knowledge of the frequency dependence of the signal. The ILC produces a minimum-variance linear combination of the measured frequency maps. In the case of Gaussian, statistically isotropic fields, this is the optimal linear combination, as the variance is the only statistic of interest. However, in many cases the signal we wish to isolate, or the foregrounds we wish to remove, are non-Gaussian and/or statistically anisotropic (in particular for the case of Galactic foregrounds). In such cases, it is possible that machine learning (ML) techniques can be used to exploit the non-Gaussian features of the foregrounds and thereby improve component separation. However, many ML techniques require the use of complex, difficult-to-interpret operations on the data. We propose a hybrid method whereby we train an ML model using only combinations of the data that , and combine the resulting ML-predicted foreground estimate with the ILC solution to reduce the error from the ILC. We demonstrate our methods on simulations of extragalactic temperature and Galactic polarization foregrounds and show that our ML model can exploit non-Gaussian features, such as point sources and spatially varying spectral indices, to produce lower-variance maps than ILC—e.g., reducing the variance of the B-mode residual by factors of up to 5—while preserving the signal of interest in an unbiased manner. Moreover, we often find improved performance even when applying our ML technique to foreground models on which it was not trained. Published by the American Physical Society 2025

McCarthy, Fiona (ORCID:0000000253893565)↗

Simultaneous inference of equation of state parameters and unknown data errors with uncertainty quantification via hierarchical Bayesian posterior maximization

Equations of state (EOSs) are a key component in running hydrodynamic simulations as they relate the thermodynamic states for the material. The Davis reactants EOS is commonly used for modeling high explosives (HEs), and the EOS model parameters are calibrated using material specific data. The calibrations are often performed with uncertainty quantification via Bayesian inference to account for uncertainty in the data and generate ensembles of likely parameters. However, there are relatively few HE data sets to use for calibration and many are historical and lack error information. In this work, we simultaneously calibrate the Davis reactants EOS model parameters and unknown data error terms for the high explosive PBX 9501. To quantify the uncertainty in the models and the data, we use a Bayesian framework for the calibration and compute the hierarchical Bayesian posterior distribution with both a posteriori maximization approach and Markov Chain Monte Carlo. In general, we find that, given our assumptions, the two approaches result in similar calibrated parameters, posterior covariance matrices, and insights about the parameters but that the posterior maximization requires far less computational resources.

97 MATHEMATICS AND COMPUTING↗

CV4Quantum

CV4Quantum is a statistical technique for reducing the sampling overhead in probabilistic error cancellation, which is an error mitigation technique used in quantum computing. CV4Quantum is based on the control variates method, which is a Monte Carlo variance reduction technique. This repository contains the code and data associated with a demonstration of CV4Quantum using simulation experiments.

Shyamsundar, Prasanth [Fermi National Accelerator ↗

Measuring labor productivity dynamics in U.S. industrial and electric power sectors: a case study (2014–2023)

This study proposes a subsystem methodology for measuring labor productivity in the U.S. industrial and electric power sectors by leveraging public data available between 2014 and 2023. Building on Pasinetti’s framework and subsequent developments, the approach employs Vertically Integrated Sectors (VIS) to account for both direct and indirect productivity effects. The novelty of this work is twofold. First, it enables the estimation of productivity trends over time, providing a robust foundation for empirical analysis. Second, it applies the methodology to a case study of the electric generation sector, highlighting its practical relevance. Using data from the Bureau of Economic Analysis, the Bureau of Labor Statistics, and the Impact Analysis for Planning (IMPLAN) tool, the study reveals significant discrepancies between conventional productivity measures and those derived from the VIS approach. Furthermore, the proposed method aligns with the principles of Integrated Energy Systems by capturing the interrelations among generation, distribution, storage, and consumption. This alignment underscores its utility and applications for energy-related policy and planning. Overall, the findings contribute to more precise labor productivity assessments, supporting informed decision-making and future research. Additionally, the method highlights the importance of considering the whole supply chain, providing interrelated metrics for labor productivity which includes both, direct and indirect effects on the final labor productivity metric. By incorporating intersectoral dependencies, this method offers a more comprehensive and accurate measure of labor productivity compared to traditional metrics.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Advanced Method Optimization for Sampling and Analysis Instrumentation

This work presents a generalized approach for analytical method optimization that branches the gap between techniques historically employed and accurate modern optimization techniques suitable for various applications. The novelty of the described strategy is the utilization of multivariate, multiobjective optimization with Karush-Kuhn-Tucker conditions to bound the optimization space to solutions within the physical limitations of instrumentation. Briefly, the basic steps outlined in this paper are to (1) determine the objective(s) that should be maximized or minimized based on the goals of the analytical application, (2) conduct a screening experiment, (3) perform ANOVA to determine the parameters which have a statistically significant effect on the objective, (4) conduct an experiment (e.g., Box-Behnken design) to collect data for fitting the objective equation, and (5) determine the physical constraints of the parameters and solve the Lagrangian to determine the optimal method parameters. A broad approach to optimization target selection allows for robust method tuning to develop improved data sets amenable for chemometrics and machine learning algorithm development. Gas chromatography-mass spectrometry was selected as a use case due to its broad use across scientific fields and time-consuming method development involving numerous parameters. In conclusion, this strategy can reduce the cost of research, improve data quality, and enable the rapid development of new analytical technique.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗