Search NASA⌕ Search

SEARCH · Search NASA

Results for “Statistical Methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Localization of infrasonic sources via Bayesian back projection

SUMMARY A Bayesian framework is investigated for event-specific localization of infrasonic sources using back projection ray tracing. Direction-of-arrival information from array-based detection analysis is used to initialize a back projection ray path originating from the detecting array location and quantifying propagation characteristics from hypothetical source locations. The Fisher statistic, computed from the array’s beam coherence, is mapped into uncertainty in the launch angles of the ray path. Auxiliary parameters previously introduced for solving the Transport equation to compute geometric spreading along ray paths are used to map uncertainty in the ray launch angles into spatial and temporal uncertainties in the ray path. An atmospheric ensemble approach is applied to account for atmospheric uncertainty, and the relation between uncertainties in the atmospheric state and confidence in estimated localization are evaluated using several ensembles with specified variances. The method is evaluated using a synthetic event in the western United States constructed via forward propagation simulations as well as a single-station, multi-arrival detection from a surface explosion in the western United States. Localization results using this event-specific approach are more accurate and exhibit improved precision than existing Bayesian localization methods that leverage generalized, pre-computed propagation statistics.

58 GEOSCIENCES↗

Animal movement estimation and network-based epidemic modeling: Illustration for the swine industry in Iowa (US)

Animal movement plays a critical role in disease transmission between farms. However, in the United States, the lack of available animal shipment data, sometimes coupled with a lack of detailed information about farm demographics and characteristics, presents great challenges for epidemic modeling and prediction. In this study, we proposed a new method based on the maximum entropy to generate “synthetic” animal movement networks, considering available statistics about the premises operation type, operation size, and the distance between premises. We illustrated our method for the swine movement networks in Iowa and performed network analyses to gain insights into the swine industry. We then applied the generated networks to a network-based epidemic model to identify potential system vulnerabilities in terms of disease transmission. The model was parameterized for African Swine Fever (ASF) as the US swine industry is quite concerned about this disease. Results show that premises with a central role in the network are more vulnerable to disease outbreaks and play an important role in disease spread. Simulations with outbreaks starting from random farms reveal no significant large outbreaks, indicating the system’s relative robustness against arbitrary disease introductions. However, outbreaks originating from high out-degree farms can lead to large epidemic sizes. This underscores the importance for stakeholders and policymakers to continue improving animal movement records and traceability programs in the US and the value of making that data available to epidemiologists and modelers to better understand risk and inform strategies aimed to cost-effectively prevent and control disease transmission. Our approach could be easily adapted to estimate movement networks in other animal production systems and to inform disease spread models for various infectious diseases.

60 APPLIED LIFE SCIENCES↗

Simulations of classical three-body thermalization in one dimension

One-dimensional systems, such as nanowires or electrons moving along strong magnetic field lines, have peculiar thermalization physics. The binary collision of pointlike particles, typically the dominant process for reaching thermal equilibrium in higher-dimensional systems, cannot thermalize a 1D system. We study how dilute classical 1D gases thermalize through three-body collisions. We consider a system of identical classical point particles with pairwise repulsive inverse power-law potential V ij ∝ 1/|x i –x j | n or the pairwise Lennard-Jones potential. Using Monte Carlo methods, we compute a collision kernel and use it in the Boltzmann equation to evolve a perturbed thermal state with temperature T toward equilibrium. We explain the shape of the kernel and its dependence on the system parameters. Additionally, we implement molecular dynamics simulations of a many-body gas and show agreement with the Boltzmann evolution in the low-density limit. For the inverse power-law potential, the rate of thermalization is proportional to ρ 2 ⁢T$\frac{1}{2}$ – $\frac{1}{n}$, where ρ is the number density. Furthermore, the corresponding proportionality constant decreases with increasing n.

1-dimensional systems↗

A novel conditional generative model for efficient ensemble forecasts of state variables in large-scale geological carbon storage

Integrating monitoring data to efficiently update reservoir pressure and CO 2 plume distribution forecasts presents a significant challenge in geological carbon storage (GCS) applications. Inverse modeling techniques are commonly used to fuse observational data and refine reservoir model parameters, thereby improving state variable forecasts. However, these techniques often rely on linear or Gaussian assumptions, which can limit their effectiveness in accurately predicting state variables. Moreover, simulating large-scale three-dimensional (3D) GCS problems is computationally expensive, making iterative runs in inverse problems prohibitive. To address these challenges, we propose a conditional generative model utilizing the score-based diffusion method for real-time 3D pressure and saturation field distribution predictions. Our approach involves solving the score function with a mini-batch-based Monte Carlo estimator to generate labeled data. This data is subsequently employed to train a fully connected neural network, enabling it to learn the conditional sample generator within a supervised learning framework. This method enables the rapid generation of a large ensemble of predictions, facilitating comprehensive uncertainty quantification of state variables. Here we applied our method to forecast the dynamic 3D distributions of pressure and saturation fields over a 30-year injection period. The statistical assessment with low root mean square error (RMSE) values demonstrates that our method can accurately predict the spatiotemporal distributions of both pressure and saturation fields. Moreover, the developed conditional generative model shows high computational efficiency by generating 100 ensemble forecasts of 3D state variables in less than 10 min. The consistency between ensemble averages and ground truth values further illustrates the model’s capability to capture state variable dynamics during the CO 2 plume injection process. Notably, the ground truth values fall within the ensemble forecasts, indicating that our uncertainty quantification effectively captures variability and potential noise in the observations. Thus, the developed conditional generative model proves to be a more efficient, accurate, and practical tool for GCS applications, facilitating timely risk analysis and informed decision-making.

58 GEOSCIENCES↗

LandScan HD: a high-resolution gridded ambient population methodology for the world

Unwarned population distributions accounting for routine human activities are needed to address many global human security challenges, including disasters, conflict, and infrastructure demand. LandScan High Definition (LSHD) supports this need through gridded ambient population estimates that measure average human presence between daytime and nighttime at a high spatial resolution of 3 arcseconds (approximately 90 m). Although LSHD has traditionally been produced on a country-specific basis, advances in global foundational data and computational resources now enable scaling its methodology to the world. Combining aspects of top-down and bottom-up gridded population methods, LSHD allocates subnational population totals from authoritative statistics to built-up areas based on occupancy estimates for multiple facility types (e.g., residential, commercial) and then reaggregates these estimates to a global population grid. We scale this approach by organizing the LSHD data stack into a 1° resolution tileset of vector analytic features, enabling an efficient and repeatable workflow for all countries worldwide. Examining the Philippines as an output of the global LSHD baseline dataset, we contrast unwarned and residential (WorldPop) population distributions by (1) exploring a practical application of flood risk assessment and (2) evaluating their congruence with outcomes of collective human activities (subnational CO 2 emissions). Finally, we discuss plans to address current LSHD limitations through data/modeling and uncertainty quantification improvements and provide outlook for workflow automation and extending the model to social, demographic and economic population characteristics.

Building morphology↗

Asymptotic inconsistency of the cumulative algorithm for laser-induced damage probability analysis

The “cumulative algorithm” is a data analysis method that has been proposed to provide an objective, nonparametric determination of laser-induced damage probability as a function of fluence from experimental data that contain both damaged sites and undamaged sites (i.e., 1-on-1 or S-on-1 testing protocols). In this work, the limitations of this approach are explored by considering the asymptotic limit of a large number of test sites. It is shown that the cumulative algorithm does not converge to the true probability distribution and significantly underestimates the damage probability near the damage onset. Here, based on the results of this work, the cumulative algorithm is not recommended for accurate estimation of damage probability.

Computational methods↗

Correlation function metrology for warm dense matter: Recent developments and practical guidelines

X-ray Thomson scattering (XRTS) has emerged as a valuable diagnostic for matter under extreme conditions, as it captures the intricate many-body physics of the probed sample. Recent advances, such as the model-free temperature diagnostic of Dornheim et al. [Nat. Commun. 13 , 7911 (2022)], have demonstrated how much information can be extracted directly within the imaginary-time formalism. However, since the imaginary-time formalism is a concept often difficult to grasp, we provide here a systematic overview of its theoretical foundations and explicitly demonstrate its practical applications to temperature inference, including relevant subtleties. Furthermore, we present recent developments that enable the determination of the absolute normalization, Rayleigh weight, and density from XRTS measurements without reliance on uncontrolled model assumptions. Finally, we outline a unified workflow that guides the extraction of these key observables, offering a practical framework for applying the method to interpret experimental measurements.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Intermediate time sub-diffusion and stress relaxation in ring polymer melts

The slow dynamics of non-concatenated ring melts remains a frontier problem in polymer science with implications for many soft material environments including cellular biophysics. Here, in this work, we report large-scale simulations of model ring melts that analyze the monomer and center-of-mass (CM) mean square displacements (MSD) and stress relaxation function on intermediate time and length scales. The degree of dynamical slowing down is characterized by the maximally sub-diffusive fractional time scaling exponents. The data span an exceptionally wide range of ring degrees of polymerization and stiffnesses and are not successfully organized based on the classic measure linear chain entanglement, N/N e . Rather, we find that the crossover degree of polymerization, N D , based on ring macromolecular caging that successfully allows master curves to be constructed for the long-time CM self-diffusion constant also collapses these temporal dynamic scaling exponents. Different properties display different exponents and exhibit one or two regimes of linear variation with the logarithm of N D / N . A distinct crossover of the CM-MSD and stress relaxation exponents emerges at sufficiently large N or stiffness that is not found for the monomer MSD, indicating a novel form of dynamic decoupling. This crossover aligns with the predicted critical degree of polymerization for transitioning from a weak to strong caging regime, indicative of activated transport. The latter may reflect the emergence of an intermolecular collective contribution to stress in analogy with dense soft colloidal matter. Suggestions are made for future theoretical work to address the rich patterns of behavior discovered.

Anomalous diffusion↗

Learning new physics from data: A symmetrized approach

Thousands of person years have been invested in searches for new physics (NP), the majority of them motivated by theoretical considerations. Yet, no evidence of beyond the Standard Model physics has been found. This suggests that model-agnostic searches might be an important key to explore NP, and help discover unexpected phenomena which can inspire future theoretical developments. A possible strategy for such searches is identifying asymmetries between data samples that are expected to be symmetric within the Standard Model. We propose exploiting neural networks (NNs) to quickly fit and statistically test the differences between two samples. Our method is based on an earlier work, originally designed for inferring the deviations of an observed dataset from that of a much larger reference dataset. We present a symmetric formalism, generalizing the original one, avoiding fine-tuning of the NN parameters and any constraints on the relative sizes of the samples. Our formalism could be used to detect small symmetry violations, extending the discovery potential of current and future particle physics experiments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

First observation of antiproton annihilation at rest on argon in the LArIAT experiment

We report the first observation and measurement of antiproton annihilation at rest on argon track and shower multiplicities and particle identification conducted with the LArIAT experiment. Stopping antiprotons from the Fermilab Test Beam Facility’s charged particle test beam are identified using beamline instrumentation and LArIAT’s liquid argon time projection chamber (LArTPC). The charged particle multiplicity from the annihilation vertex is manually evaluated via hand scanning, yielding a mean of 3.2 ± 0.4 tracks and a standard deviation of 1.3 tracks, consistent with a semiautomated reconstruction resulting in 2.8 ± 0.4 tracks and a standard deviation of 1.2 tracks. Both methods are consistent with Monte Carlo simulations within statistical uncertainty. The shower multiplicities and particle identification for outgoing tracks are also consistent with eant4 model predictions. These results, obtained from a low-statistics sample, provide a foundation for higher-statistics studies in larger LArTPCs, which could refine modeling of intranuclear annihilation on argon and inform scenarios such as neutron-antineutron oscillations. Published by the American Physical Society 2025

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Modeling Spatial Asymmetries in Teleconnected Extreme Temperatures

Abstract Combining strengths from deep learning and extreme value theory can help describe complex relationships between variables where extreme events have significant impacts (e.g., environmental or financial applications). Neural networks learn complicated nonlinear relationships from large datasets under limited parametric assumptions. By definition, the number of occurrences of extreme events is small, which limits the ability of the data-hungry, nonparametric neural network to describe rare events. Inspired by recent extreme cold winter weather events in North America caused by atmospheric blocking, we examine several probabilistic generative models for the entire multivariate probability distribution of daily boreal winter surface air temperature. We propose metrics to measure spatial asymmetries, such as long-range anticorrelated patterns that commonly appear in temperature fields during blocking events. Compared to vine copulas, the statistical standard for multivariate copula modeling, deep learning methods show improved ability to reproduce complicated asymmetries in the spatial distribution of ERA5 temperature reanalysis, including the spatial extent of in-sample extreme events.

Krock, Mitchell L.↗

An Integrated Framework for Memory-Centric Analysis: From Trace Collection to Co-Design

The memory wall phenomenon—where advances in processor performance significantly outpace those in memory subsystems—poses a fundamental challenge for contemporary computing systems. In memory-bound applications, memory subsystem behavior dominates performance, yet existing analysis approaches present significant limitations: detailed microarchitectural simulators require days to weeks to simulate modest workloads; hardware performance counters provide only aggregate statistics that obscure temporal and spatial access patterns; and scaled simulation approaches face challenges in capturing certain behaviors that emerge at larger scales. These limitations reflect a processor-centric design philosophy increasingly misaligned with memory-bound workloads where detailed understanding of memory access patterns, cache hierarchy interactions, and contention is critical for effective optimization. This paper presents an integrated framework for memory-centric analysis that enables effective hardware-software co-design. We describe practical trace collection techniques, including hardware-assisted processor tracing with minimal overhead and portable software-based instrumentation with statistical sampling. We present multi-perspective analysis methods that examine memory behavior from temporal, sequential, spatial, and relational viewpoints, revealing distinct optimization opportunities invisible in aggregate metrics. We detail an architectural modeling framework that uses sampled traces with temporal interpolation and confidence-based filtering to evaluate cache and memory configurations. Evaluation on representative benchmarks demonstrates that this framework achieves practical accuracy (L2 cache errors of 2.64\%, confidence-filtered L3 errors of 9.92\%, bandwidth errors of 7.33\%) while providing substantial speedup (26.8×) over cycle-accurate simulation, enabling rapid design space exploration. We demonstrate how this integrated framework enables systematic identification of both hardware optimizations (memory controller tuning, bank partitioning, NUMA configuration) and software optimizations (data layout restructuring, prefetching strategies, memory-aware scheduling). Through this comprehensive treatment of the memory-centric analysis pipeline—from trace collection through architectural modeling to co-design application—we provide researchers and practitioners with practical techniques for addressing memory bottlenecks in contemporary computing systems.

Gajaria, Dhruv Mayur↗

Global Corn Heat Stress: Mean and SD of Degree Days Above 29°C based on NEX-GDDP-CMIP6 Climate Projections

Description This global dataset provides the estimated mean and standard deviation (SD) of corn heat stress (degree days above 29°C) for a set of climate models in NEX-GDDP-CMIP6 at 0.25-degree resolution. The NEX-GDDP-CMIP6 dataset is comprised of global downscaled climate scenarios derived from the General Circulation Model (GCM) runs conducted under the Coupled Model Intercomparison Project Phase 6 (CMIP6). The current dataset includes: Long-Term Average Degree Days Above 29°C- Historical Long-Term Average Degree Days Above 29°C- SSP245 Long-Term Standard Deviation of Degree Days Above 29°C- Historical Long-Term Standard Deviation of Degree Days Above 29°C- SSP245 The mean and SD are calculated over 1985-2014 for the historical period and over 2035-2064 for future projections. A full description of methods, including growing season, daily temperature distribution, and statistical coefficients, can be found in Haqiqi (2024). The source climate data are obtained from https://ds.nccs.nasa.gov/thredds2/catalog/catalog.html and are described in Thrasher et al (2022). The codes used to create this dataset are available at https://github.com/ihaqiqi/dd29c_nex_cmip6. Acknowledgments This work was supported by the US Department of Energy, Office of Science, Biological and Environmental Research Program, Earth and Environmental Systems Modeling, MultiSector Dynamics under Cooperative Agreement DE-SC0022141. The data processing, computation, and storage were completed on Purdue Anvil supercomputer and cyberinfrastructure supported by the National Science Foundation HDR award # 2118329: "NSF Institute for Geospatial Understanding through an Integrative Discovery Environment (I-GUIDE)". References Haqiqi. I. (2024). Trade can buffer climate-induced risks and volatilities in crop supply. Environmental Research: Food Systems. https://doi.org/10.1088/2976-601X/ad7d12 Thrasher, B., Wang, W., Michaelis, A., Melton, F., Lee, T., & Nemani, R. (2022). NASA global daily downscaled projections, CMIP6. Scientific Data, 9(1), 262. https://doi.org/10.1038/s41597-022-01393-4

Climate Change↗

Production of alternate realizations of DESI fiber assignment for unbiased clustering measurement in data and simulations

A critical requirement of spectroscopic large scale structure analyses is correcting for selection of which galaxies to observe from an isotropic target list. This selection is often limited by the hardware used to perform the survey which will impose angular constraints of simultaneously observable targets, requiring multiple passes to observe all of them. In SDSS this manifested solely as the collision of physical fibers and plugs placed in plates. In DESI, there is the additional constraint of the robotic positioner which controls each fiber being limited to a finite patrol radius. A number of approximate methods have previously been proposed to correct the galaxy clustering statistics for these effects, but these generally fail on small scales. To accurately correct the clustering we need to upweight pairs of galaxies based on the inverse probability that those pairs would be observed (Bianchi & Percival 2017). This paper details an implementation of that method to correct the Dark Energy Spectroscopic Instrument (DESI) survey for incompleteness. To calculate the required probabilities, we need a set of alternate realizations of DESI where we vary the relative priority of otherwise identical targets. These realizations take the form of alternate Merged Target Ledgers (AMTL), the files that link DESI observations and targets. We present the method used to generate these alternate realizations and how they are tracked forward in time using the real observational record and hardware status, propagating the survey as though the alternate orderings had been adopted. We detail the first applications of this method to the DESI One-Percent Survey (SV3) and the DESI year 1 data. We include evaluations of the pipeline outputs, estimation of survey completeness from this and other methods, and validation of the method using mock galaxy catalogs.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Two-pion contribution to the hadronic vacuum polarization with staggered quarks

We present results from the first lattice QCD calculation of the two-pion contributions to the light-quark connected vector-current correlation function obtained from staggered-quark operators. We employ the MILC Collaboration’s gauge-field ensemble with 2 + 1 + 1 flavors of highly improved staggered sea quarks at a lattice spacing of a ≈ 0.15 fm with a light sea-quark mass at its physical value. The two-pion contributions allow for a refined determination of the noisy long-distance tail of the vector-current correlation function, which we use to compute the light-quark connected contribution to hadronic vacuum polarization (HVP) with improved statistical precision. We compare our results with traditional noise-reduction techniques used in lattice QCD calculations of the light-quark connected HVP, namely, the so-called fit and bounding methods. We observe a factor of roughly 3 improvement in the statistical precision in the determination of the HVP contribution to the muon’s anomalous magnetic moment over these approaches. We also lay the group theoretical groundwork for extending this calculation to finer lattice spacings with increased numbers of staggered two-pion taste states.

Lahert, Shaun [Utah U.; Illinois U., Urbana] (ORCI↗

Filtered Rayleigh-Ritz is all you need

Recent work has shown that the (block) Lanczos algorithm can be used to extract approximate energy spectra and matrix elements from (matrices of) correlation functions in quantum field theory, and identified exact coincidences between Lanczos analysis methods and others. In this work, we note another coincidence: the Lanczos algorithm is equivalent to the well-known Rayleigh-Ritz method applied to Krylov subspaces. Rayleigh-Ritz provides optimal eigenvalue approximations within subspaces; we find that spurious-state filtering allows these optimality guarantees to be retained in the presence of statistical noise. We explore the relation between Lanczos and Prony's method, their block generalizations, generalized pencil of functions (GPOF), and methods based on the generalized eigenvalue problem (GEVP), and find they all fall into a larger "Prony-Ritz equivalence class", identified as all methods which solve a finite-dimensional spectrum exactly given sufficient correlation function (matrix) data. This equivalence allows simpler and more numerically stable implementations of (block) Lanczos analyses.

97 MATHEMATICS AND COMPUTING↗

A Frequency Domain Methodology for Quantitative Evaluation of Diffuse Wavefield With Applications to Seismic Imaging

Abstract Ambient Noise Imaging (ANI) of subsurface structures relies on seismic interferometry of diffuse seismic wavefields. However, the lack of effective methods to quantify and identify highly diffuse waves hampers applications of ANI, particularly in evaluating seismic attenuation and monitoring structural changes with high temporal resolution. Conventional ANI approaches require data normalization, which effectively suppresses the non‐diffuse component with large amplitude but also results in significant loss of amplitude and phase information in the continuous seismic records. In this study, we propose a frequency domain method to quantitatively evaluate the degree of diffuseness of seismic wavefields by analyzing their statistical characteristics of modal amplitudes for stationarity and randomness. Tests on synthetic waveform and field nodal records show that the proposed method can effectively distinguish between diffuse and non‐diffuse waveforms for either single‐ or three‐component data. As an application, we identify a 60‐s‐long diffuse coda of a local M 2.2 earthquake recorded by a dense nodal array on the San Jacinto Fault Zone, and successfully extract high‐quality dispersion curve andQ‐value without performing data normalization. These results are consistent with those obtained by conventional methods that assess the correlation between coherency and the Green's function, and by modeling ballistic waves generated by road traffic. Our proposed method can advance the imaging of subsurface velocity and attenuation structures as well as monitoring temporal changes for scientific studies and engineering applications.

Geochemistry & Geophysics↗