Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24

Statistical White-Line Analysis in High-Throughput TXM-XANES for Chemical State Quantification

The transmission X-ray microscopy (TXM) based X-ray absorption near-edge structure (XANES) technique provides three-dimensional mapping of element-specific chemical states at nanometer-scale spatial resolution and micrometer-scale fields of view. However, compared to conventional volume-averaged XANES (VA-XANES) measurements, the inherently small voxel size in TXM-XANES leads to a lower signal-to-noise ratio, making full-spectrum analysis computationally demanding and less robust. Here, we present the structural and compositional conditions for a statistical white-line analysis framework under which chemical state information can be directly extracted from the white-line peak position in voxel spectra without the need for voxel-wise background subtraction or normalization, under well-defined structural and compositional conditions. The method is validated on layered oxide cathode materials, where low-order polynomial fitting accurately reproduces white-line features, and the extracted energy distributions correlate strongly with VA-XANES results. This statistical approach enables high-throughput, dose-efficient, and noise-robust chemical state quantification in TXM-XANES, offering broad applicability to functional materials requiring nanoscale oxidation-state mapping.

TXM↗

Stochastic equilibrium Raman spectroscopy (STERS)

In this manuscript, we propose a new method for cavity- and surface-enhanced Raman spectroscopy (SERS) with improved temporal resolution in the measurement of stochastic Raman spectral fluctuations. Our approach combines Fourier spectroscopy and photon correlation to decouple the integration time from the temporal resolution. Using statistical optics Monte Carlo simulations, we establish the relationship between time resolution and Raman signal strength, revealing that typical Raman spectral fluctuations, commensurate with molecular conformational dynamics, can theoretically be resolved on micro- to millisecond timescales. The method can further extract average single-molecule dynamics from small sub-ensembles, thereby potentially mitigating challenges in achieving strictly single-molecule isolation on SERS substrates.

Cobb-Bruno, Colburn [University of California, Ber↗

Bipartite mutual information in classical many-body dynamics

Information theoretic measures have helped to sharpen our understanding of many-body quantum states. As perhaps the most well-known example, the entanglement entropy (or more generally, the bipartite mutual information) has become a powerful tool for characterizing the dynamical growth of quantum correlations. By contrast, although computable, the bipartite mutual information (MI) is almost never explored in classical many particle systems; this owes in part to the fact that computing the MI requires keeping track of the evolution of the full probability distribution, a feat which is rarely done (or thought to be needed) in classical many-body simulations. Here, we utilize the MI to analyze the spreading of information in 1D elementary cellular automata (CA). Broadly speaking, we find that the behavior of the MI in these dynamical systems exhibits a few different types of scaling that roughly correspond to known CA universality classes. Of particular note is that we observe a set of automata for which the MI converges parametrically slowly to its thermodynamic value. We develop a microscopic understanding of this behavior by analyzing a two-species model of annihilating particles moving in opposite directions. Furthermore, our work suggests the possibility that information theoretic tools such as the MI might enable a more fine-grained characterization of classical many-body states and dynamics.

Cellular automata↗

Geospatial analysis of preterm and small-for-gestational age births in Washington D.C.

Background: This study is based on the recognition that adverse pregnancy outcomes significantly affect maternal and infant health, leading to increased morbidity and mortality. These outcomes are shaped by a complex interplay of individual-level factors—like maternal age and education—and community-level influences, including socio-economic status and access to healthcare. Understanding these determinants is crucial for developing effective public health strategies, especially for marginalized populations, by identifying high-risk areas and informing targeted interventions that address both individual and structural barriers. Methods: We utilized geospatial analysis to explore the association between individual- and community-level factors and adverse pregnancy outcomes, specifically preterm birth (PTB) and small-for-gestational-age (SGA) birthweight in Washington, D.C. We used Empirical Bayes smoothing methods to calculate rates of adverse birth outcomes from 2010 to 2018 at the U.S. Census tract–level. Spatial scan statistics were used to investigate if adverse birth outcomes clustered in specific areas. ANOVA tests were conducted for individual- and community-level factors within identified clusters. Results: Spatial analysis identified significant high-risk clusters for PTB and SGA infants primarily in southeastern Washington, D.C., particularly in Wards 7 and 8. Individuals residing within these clusters experienced a 47% increased risk of PTB (RR = 1.467) and a 56% increased risk of SGA (RR = 1.560) compared to those outside clusters. Space–time analysis revealed temporal variation, with PTB clusters persisting from 2011 to 2014 and SGA clusters extending through 2017. Compared to low-risk clusters, high-risk clusters had younger birthing individuals (mean age ~26.5 vs. ~33 years), lower maternal college degree attainment (~20% vs. ~80%), higher rates of late or no prenatal care (~16% vs. 11%), and increased prevalence of smoking and hypertension (all P < 0.001). Community-level indicators showed lower median household incomes ($\$40,000$ vs. ~$\$105,000$), greater poverty (~16% vs. ~7% below $\$10,000$/year), higher public assistance use (~32% vs. ~5%), and reduced healthcare access (greater distances to emergency and specialty care) in high-risk areas (all P < 0.001). Neighborhood deprivation indices were significantly elevated, commutes were longer, and population density was lower in these clusters. These findings highlight that adverse birth outcomes cluster in neighborhoods with pronounced socioeconomic and health disparities. Conclusion: High-risk birth clusters highlight intertwined factors: individual, socio-economic, and geographic. Addressing these requires comprehensive interventions focusing on social and structural determinants of health.

Birth outcomes↗

The Atacama Cosmology Telescope: Mitigating the Impact of Extragalactic Foregrounds for the DR6 Cosmic Microwave Background Lensing Analysis

We investigate the impact and mitigation of extragalactic foregrounds for the cosmic microwave background (CMB) lensing power spectrum analysis of Atacama Cosmology Telescope (ACT) data release 6 (DR6) data. Two independent microwave sky simulations are used to test a range of mitigation strategies. We demonstrate that finding and then subtracting point sources, finding and then subtracting models of clusters, and using a profile bias-hardened lensing estimator together reduce the fractional biases to well below statistical uncertainties, with the inferred lensing amplitude, A lens , biased by less than 0.2σ. We also show that another method where a model for the cosmic infrared background (CIB) contribution is deprojected and high-frequency data from Planck is included has similar performance. Other frequency-cleaned options do not perform as well, either incurring a large noise cost or resulting in biased recovery of the lensing spectrum. In addition to these simulation-based tests, we also present null tests on the ACT DR6 data for sensitivity of our lensing spectrum estimation to differences in foreground levels between the two ACT frequencies used, while nulling the CMB lensing signal. These tests pass whether the nulling is performed at the map or bandpower level. The CIB-deprojected measurement performed on the DR6 data is consistent with our baseline measurement, implying that contamination from the CIB is unlikely to significantly bias the DR6 lensing spectrum. This collection of tests gives confidence that the ACT DR6 lensing measurements and cosmological constraints presented in companion papers to this work are robust to extragalactic foregrounds.

79 ASTRONOMY AND ASTROPHYSICS↗

Rapid measurement of soluble xylo-oligomers using near-infrared spectroscopy (NIRS) and multivariate statistics: calibration model development and practical approaches to model optimization

Rapid monitoring of biomass conversion processes using techniques such as near-infrared (NIR) spectroscopy can be substantially quicker and less labor-, resource-, and energy-intensive than conventional measurement techniques such as gas or liquid chromatography (GC or LC) due to the lack of solvents and preparation methods, as well as removing the need to transfer samples to an external lab for analytical evaluation. The purpose of this study was to determine the feasibility of rapid monitoring of a biomass conversion process using NIR spectroscopy combined with multivariate statistical modeling, and to examine the impact of (1) subsetting the samples in the original dataset by process location and (2) reducing the spectral range used in the calibration model on model performance. We develop multivariate calibration models for the concentrations of soluble xylo-oligosaccharides (XOS), monomeric xylose, and total solids at multiple points in a biomass conversion process which produces and then purifies XOS compounds from sugar cane bagasse. A single model using samples from multiple locations in the process stream showed acceptable performance as measured by standard statistical measures. However, compared to the single model, we show that separate models built by segregating the calibration samples according to process location show improved performance. We also show that combining an understanding of the sample spectra with simple multivariate analysis tools can result in a calibration model with a substantially smaller spectral range that provides essentially equal performance to the full-range model. We demonstrate that real-time monitoring of soluble xylo-oligosaccharides (XOS), monomeric xylose, and total solids concentration at multiple points in a process stream using NIR spectroscopy coupled with multivariate statistics is feasible. Segregation of sample populations by process location improves model performance. Models using a reduced spectral range containing the most relevant spectral signatures show very similar performance to the full-range model, reinforcing the importance of performing robust exploratory data analysis before beginning multivariate modeling.

09 BIOMASS FUELS↗

Disordered Rocksalts as High‐Energy and Earth‐Abundant Li‐Ion Cathodes

To address the growing demand for energy and support the shift toward transportation electrification and intermittent renewable energy, there is an urgent need for low‐cost, energy‐dense electrical storage. Research on Li‐ion electrode materials has predominantly focused on ordered materials with well‐defined lithium diffusion channels, limiting cathode design to resource‐constrained Ni‐ and Co‐based oxides and lower‐energy polyanion compounds. Recently, disordered rocksalts with lithium excess (DRX) have demonstrated high capacity and energy density when lithium excess and/or local ordering allow statistical percolation of lithium sites through the structure. This cation disorder can be induced by high temperature synthesis or mechanochemical synthesis methods for a broad range of compositions. DRX oxides and oxyfluorides containing Earth‐abundant transition metals have been prepared using various synthesis routes, including solid‐state, molten‐salt, and sol‐gel reactions. This review outlines DRX design principles and explains the effect of synthesis conditions on cation disorder and short‐range cation ordering (SRO), which determines the cycling stability and rate capability. In addition, strategies to enhance Li transport and capacity retention with Mn‐rich DRX possessing partial spinel‐like ordering are discussed. Finally, the review considers the optimization of carbon and electrolyte in DRX materials and addresses key challenges and opportunities for commercializing DRX cathodes.

Li-ion batteries↗

Decoding substance use disorder severity from clinical notes using a large language model

Substance use disorder (SUD) poses a major concern due to its detrimental effects on health and society. SUD identification and treatment depend on a variety of factors such as severity, co-determinants (e.g., withdrawal symptoms), and social determinants of health. Existing diagnostic coding systems used by insurance providers, like the International Classification of Diseases (ICD-10), lack granularity for certain diagnoses, but American clinicians will add this granularity (as that found within the Diagnostic and Statistical Manual of Mental Disorders classification or DSM-5) as supplemental unstructured text in clinical notes. Traditional natural language processing (NLP) methods face limitations in accurately parsing such diverse clinical language. Large language models (LLMs) offer promise in overcoming these challenges by adapting to diverse language patterns. This study investigates the application of LLMs for extracting severity-related information for various SUD diagnoses from clinical notes. We propose a workflow employing zero-shot learning of LLMs with carefully crafted prompts and post-processing techniques. Through experimentation with Flan-T5, an open-source LLM, we demonstrate its superior recall compared to the rule-based approach. Focusing on 11 categories of SUD diagnoses, we show the effectiveness of LLMs in extracting severity information, contributing to improved risk assessment and treatment planning for SUD patients.

60 APPLIED LIFE SCIENCES↗

The DESI One-Percent Survey: Modelling the clustering and halo occupation of all four DESI tracers with U CHUU

We present results from a set of mock lightcones for the DESI One-Percent Survey, created from the UCHUU simulation. This 8 h −3 Gpc 3 N-body simulation comprises 2.1 trillion particles and provides high-resolution dark matter (sub)haloes in the framework of the Planck-based ΛCDM cosmology. Employing the subhalo abundance matching (SHAM) technique, we populated the UCHUU (sub)haloes with all four DESI tracers – Bright Galaxy Survey (BGS), luminous red galaxies (LRGs), emission line galaxies (ELGs), and quasars (QSOs) – to z = 2.1. Our method accounts for redshift evolution as well as the clustering dependence on luminosity and stellar mass. The two-point clustering statistics of the DESI One-Percent Survey generally agree with predictions from UCHUU across scales ranging from 0.3 h −1 Mpc to 100 h −1 Mpc for the BGS and across scales ranging from 5 h −1 Mpc to 100 h −1 Mpc for the other tracers. We observed some differences in clustering statistics that can be attributed to incompleteness of the massive end of the stellar mass function of LRGs, our use of a simplified galaxy-halo connection model for ELGs and QSOs, and cosmic variance. We find that at the high precision of UCHUU, the shape of the halo occupation distribution (HOD) of the BGS and LRG samples is smaller bias values, likely due to cosmic variance. The bias dependence on absolute magnitude, stellar mass, and redshift aligns with that of previous surveys. These results provide DESI with tools to generate high-fidelity lightcones for the remainder of the survey and enhance our understanding of the galaxy-halo connection.

cosmology↗

A cross-dimensional analysis of data-driven short-term load forecasting methods with large-scale smart meter data

Electricity load forecasting is essential to utility operation and power grid stability. A wide spectrum of data-driven methods, ranging from linear regression models to more recent deep learning models have been adopted to forecast electric load over the years. However, there still lacks a holistic evaluation of the applicability of conventional statistical and machine learning based algorithms with respect to different temporal and spatial scopes, computational requirements, and sensitivity of model-tuning. Enabled by a large-scale electricity load profile dataset of over 40,000 residential customers in a utility region, we conducted a cross-dimensional analysis of data-driven load forecasting methods. Three regression-based and seven deep learning algorithms with different model configurations were evaluated in terms of their overall and peak load prediction accuracy, and training burdens, across spatial aggregation levels ranging from the transformer, feeder, substation, to neighborhood. We found, first, the load forecasting accuracy is constrained by a predictability boundary, influenced by the forecasting horizon and spatial aggregation level. Specifically, RandomForest, XGBoost, TFT, TSMixer, and TiDE models achieved less than 10 % prediction error for up to 96-h ahead forecasting for district, substation, and feeder levels, while other models struggle at long-horizon predictions; Second, for winter and summer peak load dates, most models were able to predict the peak demand timing within ± 1 h, but the prediction percentage error varied by models, with TFT and TiDE models being the top performers; Third, models with similar prediction accuracy can differ in training burden by an order of magnitude. Therefore, choosing model configurations that balance prediction performance and computational resource is an important practical consideration for large-scale deployment of the machine learning based load forecasting. The outcome of this study can guide researchers and practitioners to choose the proper load forecasting algorithms based on their problem scope, required accuracy, and available resources. The predictability boundary can serve as a benchmark for electricity load forecasting problems with new algorithms and datasets.

Li, Han↗

Adaptive Uncertainty Quantification for Stochastic Hyperbolic Conservation Laws

Here, we propose a predictor-corrector adaptive method for the study of hyperbolic partial differential equations (PDEs) under uncertainty. Constructed around the framework of stochastic finite volume (SFV) methods, our approach circumvents sampling schemes or simulation ensembles while also preserving fundamental properties, in particular hyperbolicity of the resulting systems and conservation of the discrete solutions. Furthermore, we augment the existing SFV theory with a priori convergence results for statistical quantities, in particular push-forward densities, which we demonstrate through numerical experiments. By linking refinement indicators to regions of the physical and stochastic spaces, we drive anisotropic refinements of the discretizations, introducing new degrees of freedom where deemed profitable. To illustrate our proposed method, we consider a series of numerical examples for nonlinear hyperbolic PDEs based on Burgers’ and Euler’s equations.

97 MATHEMATICS AND COMPUTING↗

Dark Energy Survey Year 6 Results: improved mitigation of spatially varying observational systematics with masking

As photometric surveys reach unprecedented statistical precision, systematic uncertainties increasingly dominate large-scale structure probes relying on galaxy number density. Defining the final survey footprint is critical, as it excludes regions affected by artefacts or suboptimal observing conditions. For galaxy clustering, spatially varying observational systematics, such as seeing, are a leading source of bias. Template maps of contaminants are used to derive spatially dependent corrections, but extreme values may fall outside the applicability range of mitigation methods, compromising correction reliability. The complexity and accuracy of systematics modelling depend on footprint conservativeness, with aggressive masking enabling simpler, robust mitigation. We present a unified approach to define the DES Year 6 joint footprint, integrating observational systematics templates and artefact indicators that degrade mitigation performance. This removes extreme values from an initial seed footprint, leading to the final joint footprint. By evaluating the DES Year 6 lens sample MagLim++ plus plus on this footprint, we enhance the Iterative Systematics Decontamination (ISD) method, detecting non-linear systematic contamination and improving correction accuracy. While the mask's impact on clustering is less significant than systematics decontamination, it remains non-negligible, comparable to statistical uncertainties in certain w(theta) scales and redshift bins. Supporting coherent analyses of galaxy clustering and cosmic shear, the final footprint spans 4031.04 deg2, setting the basis for DES Year 6 1x2pt, 2x2pt, and 3x2pt analyses. This work highlights how targeted masking strategies optimise the balance between statistical power and systematic control in Stage-III and -IV surveys.

Rodríguez-Monroy, M. [Madrid, IFT; IJCLab, Orsay]↗

Data Assimilation for Robust UQ Within Agent-Based Simulation on HPC Systems

Agent-based simulation provides a powerful tool for in silico system modeling. However, these simulations do not provide built-in methods for uncertainty quantification (UQ). Within these types of models a typical approach to UQ is to run multiple realizations of the model then compute aggregate statistics. This approach is limited due to the compute time required for a solution. When faced with an emerging biothreat, public health decisions need to be made quickly and solutions for integrating near real-time data with analytic tools are needed. We propose an integrated Bayesian UQ framework for agent-based models based on sequential Monte Carlo sampling. Given streaming or static data about the evolution of an emerging pathogen this Bayesian framework provides a distribution over the parameters governing the spread of a disease through a population. These estimates of the spread of a disease may be provided to public health agencies seeking to abate the spread. By coupling agent-based simulations with Bayesian modeling in a data assimilation, our proposed framework provides a powerful tool for modeling dynamical systems in silico. We propose a method which reduces model error and provides a range of realistic possible outcomes. Moreover, our method addresses two primary limitations of ABMs: the lack of UQ and an inability to assimilate data. Our proposed framework combines the flexibility of an agent-based model with UQ provided by the Bayesian paradigm in a workflow which scales well to HPC systems. We provide algorithmic details and results on a simulated outbreak with both static and streaming data.

Spannaus, Adam [ORNL] (ORCID:0000000225213657)↗

Breaking the curse of dimensionality: Solving configurational integrals for crystalline solids by tensor networks

Accurately evaluating configurational integrals for dense solids remains a central and difficult challenge in the statistical mechanics of condensed systems. Here, we present a tensor network approach that reformulates the high-dimensional configurational integral for identical-particle crystals into a sequence of computationally efficient summations. We represent the integrand as a high-dimensional tensor and apply tensor-train (TT) decomposition together with a custom TT-cross interpolation. This approach circumvents the need to explicitly construct the full tensor. We introduce tailored rank-1 and rank-2 schemes optimized for sharply peaked Boltzmann probability densities, typical for identical-particle crystals. When applied to the calculation of internal energy and pressure-temperature curves for crystalline Cu and Ar at high (GPa) pressures, as well as the alpha-to-beta phase transition diagram of Sn, our method accurately reproduces molecular dynamics simulation results using tight-binding, machine learning, hierarchical interacting particle–neural network, and modified embedded atom method potentials,all within seconds of computation time.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Functional protein mining with conformal guarantees

Molecular structure prediction and homology detection offer promising paths to discovering protein function and evolutionary relationships. However, current approaches lack statistical reliability assurances, limiting their practical utility for selecting proteins for further experimental and in-silico characterization. To address this challenge, we introduce a statistically principled approach to protein search leveraging principles from conformal prediction, offering a framework that ensures statistical guarantees with user-specified risk and provides calibrated probabilities (rather than raw ML scores) for any protein search model. Our method (1) lets users select many biologically-relevant loss metrics (i.e. false discovery rate) and assigns reliable functional probabilities for annotating genes of unknown function; (2) achieves state-of-the-art performance in enzyme classification without training new models; and (3) robustly and rapidly pre-filters proteins for computationally intensive structural alignment algorithms. Our framework enhances the reliability of protein homology detection and enables the discovery of uncharacterized proteins with likely desirable functional properties.

59 BASIC BIOLOGICAL SCIENCES↗

Moments of parton distribution functions of the pion from lattice QCD using gradient flow

We present a nonperturbative determination of the pion valence parton distribution function (PDF) moment ratios ⟨𝑥 𝑛−1 ⟩/⟨𝑥⟩ up to 𝑛 = 6, using the gradient flow in lattice quantum chromodynamics (QCD). As a testing ground, we employ SU(3) isosymmetric gauge configurations generated by the OpenLat initiative with a pseudoscalar mass of 𝑚 𝜋 ≃ 411 MeV. Our analysis uses four lattice spacings and a nonperturbatively improved action, enabling full control over the continuum extrapolation, and the limit of vanishing flow time, 𝑡 →0. The flowed ratios exhibit O(𝑎 2 ) scaling across the ensembles, and the continuum-extrapolated results, matched to the $\overline{MS}$ scheme at 𝜇 = 2 GeV using next-to-next-to-leading order matching coefficients, show only mild residual flow-time dependence. The resulting ratios, computed with a relatively small number of configurations, are consistent with phenomenological expectations for the pion’s valence distribution, with statistical uncertainties that are competitive with modern global fits. These findings demonstrate that the gradient flow provides an efficient and systematically improvable method to access partonic quantities from first principles. Future extensions of this work will target lighter pion masses toward the physical point, and applications to nucleon structure such as the proton PDFs and the gluon and sea-quark distributions.

lattice QCD↗

Prime VI

SAND2025-03757O Prime VI is a distribution-of-disease outbreak model calibration code based on variational inference. It accompanies a publication for submission to Statistics in Medicine journal, and the code will be maintained for open-source use on Sandia's GitLab. The software provides methods for calibrating an epidemiological model to measured case-count data for a multitude of correlated spatial regions. The code solves a Bayesian inverse problem for model calibration where the posterior over-model parameters are approximated through a custom implementation of variational inference. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Safta, Cosmin↗

Differentially Private Adaptive Noise Injection (DP-ANI) v1.0

Location data is collected from users continuously to understand their mobility patterns. Releasing the user trajectories may compromise user privacy. Therefore, the general practice is to release aggregated location datasets. However, private information may still be inferred from an aggregated version of location trajectories. Differential privacy (DP) protects the query output against inference attacks regardless of background knowledge. This software implements a differential privacy-based privacy model that protects the user's origins and destinations from being inferred from aggregated mobility datasets. This is achieved by injecting Planar Laplace noise to the user origin and destination GPS points. The noisy GPS points are then transformed into a link representation using a link-matching algorithm. Finally, the link trajectories form an aggregated mobility network. The injected noise level is selected using the Sparse Vector Mechanism. This DP selection mechanism considers the link density of the location and the functional category of the localized links. Compared to the different baseline models, including a k-anonymity method, our differential privacy-based aggregation model offers query responses that are close to the raw data in terms of aggregate statistics at both the network and trajectory-levels with maximum 9% deviation from the baseline in terms of network length.

Peisert, Sean [Lawrence Berkeley National Laborato↗