Search NASA⌕ Search

SEARCH · Search NASA

Results for “data statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Statistical Analysis of Inter-Area Oscillations in the U.S. Eastern Interconnection: A 2017-2023 Perspective

Recent advancements and the accumulation of high-resolution, long-term phasor measurement unit (PMU) data have provided detailed insights into inter-area oscillations in power grids. This study conducts a comprehensive statistical analysis of inter-area oscillations within the United States Eastern Interconnection from 2017 to 2023. Utilizing data captured by the advanced wide-area Frequency Monitoring Network (FNET/GridEye), this investigation examines the occurrence patterns, dominant frequencies, damping ratios, and excitation mechanisms of these oscillations. Our analysis sheds light on the evolving statistical behaviors of inter-area oscillations, offering updated and critical information for grid operators and planners. The insights gained from this study can be instrumental in enhancing the operational resilience of the power network and guiding strategic developments in grid infrastructure to accommodate future challenges. Additionally, the study discusses emerging challenges associated with the modernization of the power grid, including increased renewable penetration, dynamic load variability, and cyber-physical vulnerabilities that complicate oscillation monitoring and control.

Inter-area oscillations↗

Synthesis of ARM User Facility Surface Rainfall Datasets to Construct a Best Estimate Value Added Product (PrecipBE)

Surface precipitation measurements are essential for Earth system model (ESM) evaluation and understanding cloud processes. An ever-growing need for robust, temporally evolving, and easy-to-use statistical datasets provides motivation for a baseline ground-based precipitation properties data product. The U.S. Department of Energy Atmospheric Radiation Measurement (ARM) user facility operates an extensive suite of precipitation instruments with various sensitivities and operating mechanisms, which render the decision of which instrument to use based on one or more fixed thresholds challenging and prone to errors and bias. Using a long-term instrument inter-comparison from a unique per-precipitation event perspective, rather than instantaneous sample comparison, we demonstrate that ARM rainfall-measuring instruments are generally consistent with each other at the statistical level. Inter-instrument deviations at the single event level can be large, especially for specific rainfall event properties such as maximum precipitation rates. A machine-learning (ML) analysis using a random forest regressor indicates that in some cases, depending on instrument, local site climatology, and/or specific deployment configuration, certain atmospheric state variables influence the measured quantities in an unpredictable manner. Thus, a-priori weighting of different instruments does not necessarily lead to more accurate and less biased synthesis of instrument data. These results motivate the design of the ARM precipitation best-estimate (PrecipBE) value-added product, which incorporates all valid precipitation data while considering data quality and other instrument limitations. PrecipBE consists of time series and tabular statistics datasets in an easy-to-use and insightful per-precipitation event format. It provides a large set of precipitation event properties supplemented with ancillary data from ARM datasets that correspond to the detected precipitation events. We describe the PrecipBE algorithm and demonstrate its use via the examination of a single-day output as well as a long-term trend analysis of precipitation events at the ARM Southern Great Plains (SGP) site, covering more than 30 years of data. The trend analysis tentatively suggests a long-term temporal tendency for mainly shorter and less intense precipitation events at the SGP site, but a long-term increase in annual rainfall by more than 36 mm (5 %) per decade. This rainfall trend is catalyzed primarily by more extreme event properties of relatively rare, intense precipitation events, with event total and 1 min maximum precipitation rate at a 1 year timeframe increasing up to 5 mm and 9 mm h −1 (several percent) per decade, respectively. While the currently available PrecipBE datasets (at https://adc.arm.gov/discovery/, last access: 8 December 2025) cover rainfall from multiple ARM deployments up to March 2025, PrecipBE is planned to be expanded to include solid-phase precipitation and will soon become an operational product with a several-day lag from real-time. We invite the ARM user community to leverage this new product and welcome user feedback to further enhance the dataset.

Silber, Israel [Pacific Northwest National Laborat↗

Searching for Neutrino Tridents in the NOvA Near Detector

This dissertation presents a search for neutrino trident production in the NOvA near detector through the coherent ``dimuon" channel: $\nu_\mu +\hspace{1pt}\text{X} \rightarrow \nu_\mu + \mu^- + \mu^+ +\hspace{1pt}\text{X}$. Trident production is a rare, purely electroweak process with sensitivity to physics beyond the Standard Model. The theoretical background, motivation for studying the process, and previous experimental measurements are reviewed. The analysis uses data collected by the NOvA near detector (ND) from Fermilab's Neutrinos at the Main Injector (NuMI) beam between November 2014 and February 2024, corresponding to an exposure of $25.5\times 10^{20}$ protons on target. The ND is a segmented tracking calorimeter located 800~m from the beam target, receiving neutrinos with a mean energy of 2~GeV. A multi-pass background reduction strategy is implemented, including the development of a novel dimuon-specific tracking technique. Trident candidates are identified using a boost ed decision tree classifier trained on simulated signal and background events. Limited background Monte Carlo statistics necessitate the use of functional fits to sideband data, which are extrapolated to estimate backgrounds in the signal region. The unblinded data contain 9 trident-like events, with an estimated background of 5.66 $\pm$ 5.15 events. This yields a best fit estimate of 3.34 tridents compared to the Standard Model prediction of 4.66. A profiled Feldman-Cousins method is used to determine a 90\% confidence interval of [0,9.1] on the number of signal events, corresponding to an upper limit of 1.95$\times$ the Standard Model prediction. This result represents the lowest energy search for trident events to date, and the first experimental contribution to the process in 27 years.

Bowles, Reed Scott [Indiana U.]↗

Dark Energy Survey Year 3 results: Simulation-based cosmological inference with wavelet harmonics, scattering transforms, and moments of weak lensing mass maps. Validation on simulations

Beyond-two-point statistics contain additional information on cosmological as well as astrophysical and observational (systematics) parameters. In this methodology paper we provide an end-to-end simulationbased analysis of a set of Gaussian and non-Gaussian weak lensing statistics using detailed mock catalogs of the Dark Energy Survey (DES). Here, we implement: 1) second and third moments; 2) wavelet phase harmonics (WPH); 3) the scattering transform (ST). Our analysis is fully based on simulations, it spans a space of seven $νw$CDM cosmological parameters, and it forward models the most relevant sources of systematics of the data (masks, noise variations, clustering of the sources, intrinsic alignments, and shear and redshift calibration). We implement a neural network compression of the summary statistics, and we estimate the parameter posteriors using a likelihood-free-inference approach. We validate the pipeline extensively, and we find that WPH exhibits the strongest performance when combined with second moments, followed by ST, and then by third moments. The combination of all the different statistics further enhances constraints with respect to second moments, up to 25 percent, 15 percent, and 90 percent for S 8 , Ω m , and the figure-of-merit FoM S8;Ωm , respectively. We further find that non-Gaussian statistics improve constraints on w and on the amplitude of intrinsic alignment with respect to second moments constraints. The methodological advances presented here are suitable for application to Stage IV surveys from Euclid, Rubin-LSST, and Roman with additional validation on mock catalogs for each survey. In a companion paper we present an application to DES Year 3 data.

79 ASTRONOMY AND ASTROPHYSICS↗

Stochastic 3D reconstruction of cracked polycrystalline NMC particles using 2D SEM data

Li-ion battery performance is strongly influenced by the 3D microstructure of its cathode particles. Cracks within these particles develop during calendaring and cycling, reducing connectivity but increasing reactive surface, making their impact on battery performance complex. Understanding these contradictory effects requires a quantitative link between particle morphology and battery performance. However, informative 3D imaging techniques are time-consuming, costly and rarely available, such that analyses often have to rely on 2D image data. This paper presents a novel stereological approach for generating virtual 3D cathode particles exhibiting crack networks that are statistically equivalent to those observed in 2D sections of experimentally measured particles. Consequently, 2D image data suffices for deriving a full 3D characterization of cracked cathodes particles. Such virtually generated 3D particles could serve as geometry input for spatially resolved electro-chemo-mechanical simulations to enhance our understanding of structure-property relationships of cathodes in Li-ion batteries.

36 MATERIALS SCIENCE↗

Towards next-generation optical potentials for nuclear reactions and structure calculations

Optical-model potentials (OMPs) are critical ingredients for basic and applied nuclear physics. Present-day computational capabilities allow us to generate data-driven nucleon-nucleus OMPs that are non-local and exactly dispersive (as theoretically required to be), include statistically-sound uncertainty quantification, and are trained on both scattering and bound-state data from a wide area of the nuclear chart. Combined together, these features allow for significant improvement in fidelity and extrapolative power of the model. Here, we present preliminary work toward the development and training of such an OMP. The capability of the model to describe data at this first stage is encouraging.

Perrotta, Salvatore Simone [Lawrence Livermore Nat↗

Discovering Dark Matter Clumps and Primordial Particles with Galaxies

Cosmological observations and galaxy dynamics seem to imply that five out of six parts in mass of all matter in the universe is composed of dark matter, which is not accounted for by the Standard Model of particles. The cold dark matter (CDM) paradigm has been extremely successful at describing observations on large-cosmological scale. However, many different dark matter candidates have the same observable effects as CDM on large-length scales. One possible avenue to distinguish between these different models is to look on much smaller-length scales, where usually dark matter models become distinguishable. This research developed the theoretical framework and statistical tools needed to map the detailed distribution of dark matter on subgalactic scales using strong gravitational lensing. A second goal of this research was to build new statistical techniques that efficiently exploit the full power of cosmic-structure data from next-generation surveys. Galaxy clustering on large scales provides significant cosmological information through the power spectrum. Additional information can be gained from higher-order statistics. This research provided new ways to understand the initial conditions of the universe from current and upcoming cosmological surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Site A1 - Event Log / Derived Data

This dataset contains the event log table with 10-minute wind statistics from the scanning lidar at AWAKEN's site A1. This is a good dataset to start from for people unfamiliar with the AWAKEN project.

17 WIND ENERGY↗

Site A2 - Event Log / Derived Data

This dataset contains the event log table with 10-minute wind statistics from the scanning lidar at AWAKEN's site A2. This is a good dataset to start from for people unfamiliar with the AWAKEN project.

17 WIND ENERGY↗

A kinetic-based regularization method for data science applications

We propose a physics-based regularization technique for function learning, inspired by statistical mechanics. By drawing an analogy between optimizing the parameters of an interpolator and minimizing the energy of a system, we introduce corrections that impose constraints on the lower-order moments of the data distribution. This minimizes the discrepancy between the discrete and continuum representations of the data, in turn allowing to access more favorable energy landscapes, thus improving the accuracy of the interpolator. Our approach improves performance in both interpolation and regression tasks, even in high-dimensional spaces. Unlike traditional methods, it does not require empirical parameter tuning, making it particularly effective for handling noisy data. We also show that thanks to its local nature, the method offers computational and memory efficiency advantages over Radial Basis Function interpolators, especially for large datasets.

97 MATHEMATICS AND COMPUTING↗

Recent evolution of risk analyses in atomic bomb survivor studies: new methods and applications

Abstract Several decades ago a dramatic leap forward occurred in the development and application of statistical methods for modeling radiation risk at the Radiation Effects Research Foundation (RERF). Poisson regression analysis for grouped person-year cohort data and the linear excess relative risk model were introduced, and subsequently a devoted software system, Epicure® (https://www.hirosoft.com), was developed by researchers at RERF and at the U.S. National Cancer Institute. Numerous advancements in understanding radiation effects on humans were made possible with these methods, which are still the state-of-the-art for risk assessment at RERF and have remained part of the standard toolbox for radiation—and other environmental—epidemiological studies worldwide. Nevertheless, as our understanding of radiation risk has increased, so have the breadth and depth of questions that require answers based on emerging data that are not amenable to these conventional methods. This overview briefly recounts the conventional methods and then describes our recent diversification into the use or development of new statistical approaches to meet the challenges of burgeoning biological data and emerging mechanistic information. We briefly discuss the development and application of new methods, current and planned, that are part of the RERF Statistics Department’s role in supporting institution-wide research, especially in our collaborations involving the Life Span Study, Adult Health Study, and First-generation Offspring Clinical Study. Some approaches to modeling and assessing radiation risk with newer methods mentioned herein have already been published, while some are still in development or are only beginning at the proposal stage.

Oncology↗

Probabilistic Error Bounds for Low-Rank Tensor Decompositions Used in Large-Scale Data Analysis Applications (LDRD Final Report)

This report documents a research project on analyzing low-rank tensor models for data analysis that took place at Sandia National Laboratories from October 2023–September 2025. The focus of this work was to extend theoretical frameworks from statistics and probability theory for use with models for scalar, vector, and matrix data to models with tensor, or general multi-dimensional array, data. Through this work, we have provided a new set of tools for bounding errors on low-rank tensor models of both complete and sampled data. The remainder of this report is organized as follows. In Section 1, we describe the proposed work at the start of the project. Section 2 describes the research advances made as part of the project. Other research contributions in the form of conference presentations and software development is provided in Section 3. Workforce development at Sandia and Florida Atlantic University (via a subcontract on this project) is provided in Section 4.

97 MATHEMATICS AND COMPUTING↗

Heterogeneous Multi-Domain Dataset Synthesis to Facilitate Privacy and Risk Assessments in Smart City IoT

The emergence of the Smart Cities paradigm and the rapid expansion and integration of Internet of Things (IoT) technologies within this context have created unprecedented opportunities for high-resolution behavioral analytics, urban optimization, and context-aware services. However, this same proliferation intensifies privacy risks, particularly those arising from cross-modal data linkage across heterogeneous sensing platforms. To address these challenges, this paper introduces a comprehensive, statistically grounded framework for generating synthetic, multimodal IoT datasets tailored to Smart City research. The framework produces behaviorally plausible synthetic data suitable for preliminary privacy risk assessment and as a benchmark for future re-identification studies, as well as for evaluating algorithms in mobility modeling, urban informatics, and privacy-enhancing technologies. As part of our approach, we formalize probabilistic methods for synthesizing three heterogeneous and operationally relevant data streams—cellular mobility traces, payment terminal transaction logs, and Smart Retail nutrition records—capturing the behaviors of a large number of synthetically generated urban residents over a 12-week period. The framework integrates spatially explicit merchant selection using K-Dimensional (KD)-tree nearest-neighbor algorithms, temporally correlated anchor-based mobility simulation reflective of daily urban rhythms, and dietary-constraint filtering to preserve ecological validity in consumption patterns. In total, the system generates approximately 116 million mobility pings, 5.4 million transactions, and 1.9 million itemized purchases, yielding a reproducible benchmark for evaluating multimodal analytics, privacy-preserving computation, and secure IoT data-sharing protocols. To show the validity of this dataset, the underlying distributions of these residents were successfully validated against reported distributions in published research. We present preliminary uniqueness and cross-modal linkage indicators; comprehensive re-identification benchmarking against specific attack algorithms is planned as future work. This framework can be easily adapted to various scenarios of interest in Smart Cities and other IoT applications. By aligning methodological rigor with the operational needs of Smart City ecosystems, this work fills critical gaps in synthetic data generation for privacy-sensitive domains, including intelligent transportation systems, urban health informatics, and next-generation digital commerce infrastructures.

IoT↗

Unique & challenging aspects of plutonium metal standards exchange program for actinide measurements

The Los Alamos National Laboratory exchange program is the only program of its kind for the distribution of plutonium (Pu) standards materials with a range of impurity contents to multiple laboratories for destructive measurements of elemental concentration. This paper discusses statistical methods used to address challenges in Pu metal exchange data by way of two case studies. Challenges include how to evaluate a data set when a large fraction of the values are minimum detection limits (MDLs), and how to determine potential outliers with limited in-formation on the true spread of the data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Modelling the impact of host galaxy dust on type Ia supernova distance measurements

Type Ia Supernovae (SNe Ia) are a critical tool in measuring the accelerating expansion of the universe. Recent efforts to improve these standard candles have focused on incorporating the effects of dust on distance measurements with SNe Ia. In this paper, we use the state-of-the-art Dark Energy Survey 5 year sample to evaluate two different families of dust models: empirical extinction models derived from SNe Ia data and physical attenuation models from the spectra of galaxies. In this work, we use realistic simulations of SNe Ia to forward-model different models of dust and compare summary statistics in order to test different assumptions and impacts on SNe Ia data. Among the SNe Ia-derived models, we find that a logistic function of the total-to-selective extinction R V best recreates the correlations between supernova distance measurements and host galaxy properties, though an additional 0.02 mag of grey scatter is needed to fully explain the scatter in SNIa brightness in all cases. These empirically derived extinction distributions are highly incompatible with the physical attenuation models from galactic spectral measurements. From these results, we conclude that SNe Ia must either preferentially select extreme ends of galactic dust distributions, or that the characterization of dust along the SNe Ia line-of-sight is incompatible with that of galactic dust distributions.

79 ASTRONOMY AND ASTROPHYSICS↗

Predicting U.S. federal fleet electric vehicle charging patterns using internal combustion engine vehicle fueling transaction statistics

Utilizing fueling transactions from internal combustion engine vehicles (ICEVs), the authors estimated how frequently midday public charging would be required for U.S. federal fleet battery electric vehicles (BEVs). Fueling transaction summary statistics are more widely available than trip-level telematics data, making this methodology more accessible and transferable to other researchers and fleet managers considering BEV replacements. For example, readers can easily apply a linear model using only the count of back-to-back fueling events at gas stations over 57 straight-line miles apart to predict days exceeding range. This linear regression predicted binned days exceeding 250 miles at 80% accuracy on a hold-out test set from the same fleet as the training data and 66 % accuracy on a new fleet displaying different driving behaviors. The authors additionally provide linear equations for days exceeding 200 and 300 miles as alternative range estimates to account for differences in BEV range and temperature impacts. Beyond the single-feature linear models which readers can apply, the authors tuned and trained other machine learning models on a variety of fueling transaction statistics including consecutive transaction distances, transaction distance from garage, estimated miles traveled from fuel economy and fuel quantity, and transaction periodicity. Utilizing a subset of 1678 light-duty federal fleet vehicles which contained daily vehicle miles traveled (VMT) in addition to fueling statistics, the authors determined which fueling transaction statistics were most relevant in predicting driving days exceeding 250 miles (an approximation of BEV rated driving range). In support of the U.S. federal fleet transition to zero-emission vehicles (ZEVs), the authors used these statistics and machine learning models to predict the frequency of BEV midday charging. After training models on the subset with VMT, the authors predicted days exceeding rated range for 112,902 light-duty vehicles operating in similar circumstances in the federal fleet using a Support Vector Regressor (SVR). In conclusion, they then used the projections as part of the ZEV Planning and Charging (ZPAC) tool to identify optimal candidates for BEVs for the federal fleet. An anonymized version of ZPAC is included in the supplementary materials.

25 ENERGY STORAGE↗

Lowering the barrier to access information-rich transient kinetic data for machine learning methods

Transient kinetic data contain a wealth of information about intrinsic features of a catalyst as well as the reaction mechanism. Currently, high volume transient data is underutilized, and data science methods could both increase the value of information that can be extracted from this data, integrate experimental with theoretical data sources, and accelerate the pace of catalyst technology advancement. Transient kinetic characterizations with simple probe molecules exhibiting reversible adsorption, irreversible adsorption and bulk-surface diffusion are presented as training components for similar experiments with more complex surface reactions. In conclusion, by increasing the availability and accessibility of transient kinetic data through details of its structure and acquisition, we aim to decrease the barrier for data scientists to apply machine learning methods to this valuable data source.

Catalysis↗

The National Climate Database (NCDB): An Unbiased 100-Year Dataset for PV Modeling

In this study, we develop a statistical technique to downscale the future projection of solar irradiance for photovoltaics (PV) energy-related applications. A set of Regional Climate Model (RCM)-based projections obtained from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX) are used as inputs to statistical methods to generate high-resolution global horizontal irradiance (GHI) over the contiguous United States (CONUS). The main steps of the statistical downscaling method include (1) regridding RCM output (0.22 degree and daily resolutions) to handle the modeled-observed data sets on a common grid, (2) correcting bias of RCM GHI using satellite-derived observation, and (3) implementing temporal and spatial downscaling to generate GHI at 8-km and hourly resolution. Basically, complex physical processes and interactions between solar radiation and various atmospheric constituents lead solar irradiance to be highly variable and uncertain. Underrepresentation of clouds from the RCM parameterizations is the main source of error and uncertainty in modeling solar irradiance. Thus, we adapt and use the high-quality satellite-derived data from the National Solar Radiation Database (NSRDB) to analyze the bias and error of RCM GHI as well as estimate the statistical parameters for spatial and temporal downscaling. This presentation will summarize the comprehensive analysis conducted to produce and assess the results under two climate scenarios (RCP4.5 and RCP8.5). We will also present a detailed validation demonstrating the strengths of the downscaling method, a summary of the 100-year dataset from 2001-2100, and future extension of this research.

bias correction↗