Search NASASearch

SEARCH · Search NASA

Results for “data statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Measuring Neutron Polarisation in Deuteron Photo-disintegration with the CLAS Start Counter [Thesis]

Deuteron photo-disintegration (γd → γp) is a reaction that represents the simplest case in which nuclear and hadron physics models can be tested. Despite this, associated polarization analyses are limited in terms of angular coverage and energy ranges, especially in observables related to the recoil neutron. This is largely due to a lack in dedicated polarimetry equipment, and represents a roadblock in global progress to understand high-energy phenomena such as hexaquarks, and quark-gluon degrees of freedom. To address this problem, this PhD thesis pioneers a new methodology for the parasitic measurement of nucleon polarization using kinematic reconstruction of (spin-dependent) nucleon-nucleus scattering of reaction products, prior to their detection in large acceptance particle detector apparatus. Following this novel approach, which requires no dedicated polarimeter, a determination of the double polarization observable, $C^n_{x'}$, from deuteron photo-disintegration is presented, using Jefferson Lab’s CLAS detector. The analysis utilizes the (n,p) charge exchange reaction in CLAS’s "start counter" (plastic scintillator) to determine the final state neutron polarizations. The results present the first ever data for this observable above 0.7 GeV (photon beam energy) and significantly extend the angular range of the world data set. This new data is largely statistically consistent with the previous measurement of $C^n_{x'}$ by Bashkanov et al . in the overlapping energy range of 0.4-0.7 GeV. It is planned for the statistical accuracy of the presented result to be increased by the inclusion of additional data. The analysis herein serves as a key proof of concept for future applications, including a recommended similar analysis to be implemented with data from the more modern CLAS12 detector. This paves the way for a plethora of additional analyses using existing data sets that would provide crucial new constraints for hadron and nuclear physics.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Automated analysis of unlabeled PV data with Solar Data Tools software: Overview and feature updates

Distributed rooftop PV systems: ubiquitous, yet commonly have unlabeled data Difficult or impossible to form a performance index We developed Solar Data Tools (SDT), an open-source Python library for analyzing PV power (and irradiance) time-series data SDT enables analysis of unlabeled PV data—no model, no meteorological data, no performance index required Takes a statistical signal processing approach Data processing steps are largely pre-defined and automatic regardless of system type—from utility tracking systems to multi-pitch rooftop systems

Meyers-Im, Bennet E

Investigating the Effects of Bars on Star Formation and Nuclear Activity of Galaxies Using DESI Survey Data

We present a statistical analysis of the connections between galactic bars, star formation, and active galactic nucleus (AGN) activity using 33,201 disk galaxies (0.01 < z < 0.05) from Dark Energy Spectroscopic Instrument Data Release 1 (DESI DR1) cross-matched with Galaxy Zoo DESI. Based on morphological classifications, we identify 3508 strongly barred and 8335 weakly barred systems. We find that barred galaxies exhibit a clear bimodal distribution in color–mass space: weak bars are preferentially found in bluer, lower-mass disks, whereas strong bars are more common in massive, redder systems. Strongly barred galaxies are on average more massive and metal rich than unbarred systems. In addition, strong bars enhance central star formation rates (SFRs) in low-mass galaxies but reduce specific SFRs in massive systems, reflecting a dual role where bars initially trigger central star formation but eventually promote quenching by accelerating gas consumption. In terms of nuclear activity, barred galaxies display a higher incidence of AGN activity. The presence of a bar is also associated with an increased fraction of powerful AGN, with the highest proportions found in strongly barred systems. However, the correlations between AGN activity and detailed bar structural parameters are weak, suggesting that the link between bars and nuclear activity is indirect and regulated by multiple factors. Overall, our results support a scenario in which bars facilitate angular momentum transport and gas inflow, thereby driving central star formation and fueling supermassive black hole accretion while operating alongside other processes that shape galaxy evolution.

Liu, Jianfei [Chinese Academy of Sciences (CAS), B

Full-stack Quantification of Variability in Predicting Ion Transport Properties using Machine-learned Interatomic Potentials

Machine-learned interatomic potentials (MLIPs) have become the state-of-the-art for performing accurate, scalable molecular dynamics (MD) simulations. It is therefore crucial to understand and quantify the reliability of MLIPs for downstream property predictions. Uncertainty in predicted properties can arise from limitations in first-principles training data, intrinsic MLIP model errors in representing the data, and the statistical noise introduced during subsequent MD simulations. Using ion transport in Li7P3S11 as a case study, we systematically assess the impact of training set size and selection, neural network stochasticity, and MD sampling statistics on predicted diffusivity and activation energy. We find that when using equivariant MLIP architectures with standard MD protocols, uncertainty arising from MD sampling dominates over model-induced errors. In contrast, MLIP errors relative to the underlying first-principles data are consistently minor. Given this, there are two main routes to improving the accuracy of predictions based on MLIP potentials: adopting higher accuracy reference data generation methods, and improving the MD sampling statistics.

36 MATERIALS SCIENCE

Post-irradiation Examination of Eurofer-97 Steel Irradiated to 20 dpa at 200–400°C in HFIR under the EUROfusion (ORNL-KIT) Collaboration Program

The Oak Ridge National Laboratory/Karlsruhe Institute of Technology (ORNL/KIT) collaboration focuses on research involving irradiation experiments and post-irradiation examinations (PIE) of isotopically modified Eurofer-97 steels for fusion reactor applications. This collaboration capitalizes on ORNL’s expertise in radiation effects on materials and its capabilities in irradiation and PIE. The research aims to qualify Fe-9Cr-based Eurofer-97 steel under simulated fusion conditions, which include high doses and high transmutation rates. A unique isotopic modification technique is utilized in the research, wherein high-transmutation isotopes such as Fe-54 and Ni-58 are alloyed into the Eurofer steel to align with the expected helium production rates in fusion reactor conditions. This report presents the results of post-irradiation mechanical testing activities for the EUROFER-97/2 specimens after irradiation to ~20 dpa at various irradiation temperatures. The irradiation doses and temperatures for the ES capsules ranged from 18.4–20.4 dpa and 202–256°C, respectively. The mechanical property datasets, obtained through baseline testing and PIE of the tensile and fracture specimens, include microhardness data from 93 irradiated and non-irradiated tensile specimens and 28 irradiated bend bar specimens, uniaxial tensile property data from 45 irradiated and non-irradiated tensile specimens, and fracture toughness data from the bend bar specimens. Further data analyses provide statistical information on the microhardness values and tensile properties, as well as the reference ductile-brittle transition temperature (T 0 ) data.

36 MATERIALS SCIENCE

Bioenergy Feedstock Library Annual Summary Report 2024

The Bioenergy Feedstock Library (BFL), part of the Biomass Feedstock National User Facility (BFNUF) located at Idaho National Laboratory (INL), is a physical sample repository and a web-accessible electronic database. The BFL stores physical and chemical characteristics of biomass and waste carbon sources for energy use, as well as samples generated from U.S. Department of Energy (DOE) Bioenergy Technologies Office (BETO) and U.S. Department of Agriculture-funded projects. The objective of this Bioenergy Feedstock Library Annual Summary Report for 2024, similar to the 2023 Annual Summary Report , is to focus on the updates to: (1) publicly available analytical data and equipment tracked through the BFNUF, (2) significant increases in the physical samples available for request, (3) sample and data archival progress from recent BETO-funded projects, and (4) publicly available data sets created upon request from BETO, INL projects, or outside entities compared to the previous annual summary reports. This report highlights key statistics and available data and information important for INL, BFL users, academics, and industry.

09 BIOMASS FUELS

A Field Guide to Corralling the Chaos: A Conceptual Framework for Using Models to Guide Opportunistic Field Studies of Natural Disturbances

Watersheds regulate biogeochemical processes and provide ecosystem services to human societies, but disturbances can fundamentally alter these processes across space and time. Determining when and where to sample to capture disturbance impacts in watersheds remains a central challenge. Manipulation studies and long-term monitoring are often constrained by scope, and opportunistic studies often lack pre-disturbance data needed to statistically determine disturbance impacts. We identify a persistent knowledge gap: the absence of a clear, transferable framework to guide opportunistic disturbance research where pre-disturbance data collection is not a feasible option. To address this gap, we present a conceptual framework that intentionally integrates modeling and empirical observation in an iterative, stepwise model–experiment workflow. We demonstrate its application through two contrasting case studies: wildfire impacts on headwater streams using a pre-disturbance preparedness approach, and saltwater flooding impacts on coastal forests using an ‘ex-post-facto’ approach. From these applications, we assess strengths, limitations, and the critical role of team science for transferability across disturbance types and study designs. Broadly, this framework offers a scalable path towards more rigorous, timely, and actionable disturbance science that can inform watershed management, hazard risk reduction, and ecosystem resilience.

Coastal Biogeochemistry

Dark Energy Survey Year 3 results: optimized $w$CDM simulation-based inference with weak lensing map-level hybrid statistics

We present cosmological constraints from the Dark Energy Survey Year 3 (DES Y3) weak lensing data using hierarchical hybrid statistics within a Bayesian simulation-based inference framework that is based on the Gower Street simulations. To maximize the precision of the inference, we have developed a new, information-theory based, data compression of the weak lensing maps to just seven highly informative summary statistics. The hybrid scheme exploits the high information content of the power spectrum, compressing both the power spectrum and neural-based summaries that are designed to extract further information. Our simulation-based approach enables principled forward modelling of all major sources of systematic uncertainty and survey properties into realistic mock observations, including the survey mask, photometric redshift uncertainties, intrinsic galaxy alignments, multiplicative shear calibration bias, source galaxy clustering, non-Gaussian shape noise, and non-linear structure formation. The summary statistics are then used in a Bayesian simulation-based inference pipeline. The inference is validated through coverage tests and checks for robustness against baryonic feedback. Assuming a $w$CDM cosmology, our analysis yields $S_8 = 0.808 \pm 0.017$, $Ω_{\rm m} = 0.325 \pm 0.024$, and $w < -0.766$ (marginalized posterior 68 per cent credible intervals). This rigorous combination of information theory, physics- and neural network-based extreme data compression, and principled Bayesian analysis improves the figure of merit for $(Ω_{\rm m}, S_8, w)$ by 60 per cent over the previous state-of-the-art, and by almost a factor of 3 over two-point analyses of the same data. They are the most precise joint constraints on $(Ω_{\rm m}, S_8, w)$ from weak gravitational lensing data alone of any survey to date. We intend to apply this analysis to the more recent DES Y6 data.

Williamson, J. [University Coll. London]

Applying Transfer Learning for Street-Scale Nuisance Flood Forecasting in Coastal-Urban Cities

An important challenge with Machine Learning (ML) is its transferability; that is, whether a ML model trained on one set of data can be applied to a second set of data without requiring a full re-training of the model. Transfer Learning (TL) addresses this challenge by transferring knowledge learned in the source domain (the data it was trained on) to the target domain (a second set of data that is statistically different but related, which the model was not trained on). This study investigates the use of TL for street-scale nuisance flood forecasting by exploring whether a ML model trained on data collected for one set of streets can effectively forecast flooding for another set of streets in the same city using TL. The envisioned use case is a city deploying a new flood depth monitoring sensor on a street and using TL to apply a ML model, trained on sensor data from an existing flood depth sensor network, to this new street. Eventually, the new flood depth sensor will have a sufficient dataset for training its own ML model, but TL can be used to fill the gap in time while this new dataset is being generated. This method is explored using a Long Short-Term Memory (LSTM) model trained on data for the flood-prone streets of Norfolk City, Virginia. The data used for training includes environmental time series (rainfall, tide), topographic features (Digital Elevation Model (DEM), Topographic Wetness Index (TWI), Depth To Water (DTW)), and street-scale flood depth time series obtained from a high-fidelity physics-based model, acting as a synthetic street-scale stream depth sensor dataset since actual stream depth sensor data is generally unavailable for most cities. A set of 180 flood-prone streets was used to train a base model, while another set of 180 flood-prone streets was used to re-train that model using different TL strategies. The results show that full-weight re-training proved most effective and minimal re-training of only the output layer was insufficient. The advantage of TL was most pronounced when target data was limited, meaning data collected at the new water depth sensor location included generally less than 18 flood events. As target data increased beyond 18 flood events, the benefit of TL diminished relative to training a ML model directly on the local flood events. These findings can assist cities as they implement street-scale flood sensing systems to create accurate forecasts for new sensing locations that do not yet have sufficient data records to train a local ML model.

Roy, Binata [Univ. of Virginia, Charlottesville, V

Personalized Tucker Decomposition: Modeling Commonality and Peculiarity on Tensor Data

In this paper, we propose a personalized Tucker decomposition (perTucker) to address the limitations of traditional tensor decomposition methods in capturing heterogeneity across different datasets. perTucker decomposes tensor data into shared global components and personalized local components. We introduce an order orthogonality assumption and develop a proximal gradient regularized block coordinate descent algorithm guaranteed to converge to a stationary point. The unique and common representations learned by perTucker reveal intrinsic statistical patterns in data and provide valuable information for a wide range of downstream analytics, including anomaly detection, source classification, and clustering. We demonstrate perTucker’s effectiveness through a simulation study and two case studies on solar flare detection and tonnage signal classification.

14 SOLAR ENERGY

Measurements of short-lived fission product yields from photofission of 238 U using 13.0 MeV monoenergetic photons

Photon-induced fission product yield (FPY) measurements were conducted on the isotope 238U. Fission was induced using Eγ = 13.0 MeV monoenergetic photons produced by the Triangle Universities Nuclear Laboratory’s (TUNL’s) High Intensity γ-ray Source (HIγS) facility. Short-lived FPYs were measured by performing cyclic activation of the sample using a rapid target transfer system. Following activation, the 238U target was rapidly (0.4 s) transferred to a counting station consisting of two well-shielded high-purity germanium (HPGe) detectors. The irradiation-counting cycle was repeated until the summed data had sufficient statistical accuracy. Twenty-eight unique fission products with half-lives ranging from 1 s to 450 s were identified, and their cumulative FPYs determined. Furthermore, the results are compared with previous independent FPY measurements using inverse kinematics. Good agreement between the data sets is found despite the different excitation energy distributions of the fissioning nucleus in the experiments.

Physics - Nuclear physics and radiation physics

Divide and conquer: using RhizoVision Explorer to aggregate data from multiple root scans using image concatenation and statistical methods

Roots are important in agricultural and natural systems for determining plant productivity and soil carbon inputs. Sometimes, the amount of roots in a sample is too much to fit into a single scanned image, so the sample is divided among several scans, and there is no standard method to aggregate the data. Here, we describe and validate two methods for standardizing measurements across multiple scans: image concatenation and statistical aggregation. We developed a Python script that identifies which images belong to the same sample and returns a single, larger concatenated image. These concatenated images and the original images were processed with RhizoVision Explorer, a free and open-source software. An R script was developed, which identifies rows of data belonging to the same sample and applies correct statistical methods to return a single data row for each sample. These two methods were compared using example images from switchgrass, poplar, and various tree and ericaceous shrub species from a northern peatland and the Arctic. Most root measurements were nearly identical between the two methods except median diameter, which cannot be accurately computed by statistical aggregation. We believe the availability of these methods will be useful to the root biology community.

59 BASIC BIOLOGICAL SCIENCES

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis - NLR Historical Wind

The U.S. Department of Energy and the National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis from variable sources, hydrogen compression and storage, and hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) initiative. This dataset represents part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure nationwide. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with various energy technologies. Future datasets will demonstrate how existing hydrogen fuel cell technologies can provide controllable, dispatchable, and variable power output for artificial intelligence (AI) data centers and other variable loads. This dataset entry describes hydrogen production by conducting a statistical analysis of historical wind data over a five-year period (2020-2025) from a single 1.5MW turbine manufactured by General Electric (GE) located at NLR’s Flatirons Campus, to generate an experimental test profile that was deployed on a 1.25-MW proton exchange membrane type MC250 electrolyzer system manufactured by Nel Hydrogen . [1] While the electrolyzer balance-of-plant supports up to 2.5 MW of electrolysis, NLR only has a single 1.25-MW electrolysis stack. The historical wind data provided several metrics, however, the analysis particularly focused on the measured power output by the wind turbine. The power output time series of data for each day was categorized by total energy generation and standard deviation, and the day that represented the highest combination of these two metrics was chosen – December 25th, 2022. This process was then repeated for a moving four-hour window within this day to identify the most statistically variable period. Finally, this four-hour period was scaled by 65% to match the 1.25 MW electrolyzer. The electrolysis system controls hydrogen production by varying DC current applied to the stack, from a maximum of 3000 A to a minimum safe operation of 300 A, or 10%. Because the current – voltage characteristic changes as the stack ages and efficiency degrades, the actual minimum safe operating power changes over time. The historical wind profiles were translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1 Hz frequency. For more details on the statistical analysis process, see the presentation labeled “ Public Reference Data for Megawatt-Scale Hydrogen Electrolysis” provided with each data entry. These datasets report relevant hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. Each .zip file represents a single wind turbine electrolysis experiment and is formatted as follows: {technology}_{scaling factor}-{electrolyzer ramp rate in amperes/second} For instance, “wind-GE1.5MW_0.65-400.zip” represents the hour-long experiment using historical data from the wind-GE1.5MW turbine, scaled to 65%, with the electrolyzer power supply set to a maximum ramp rate (gain and slew) of 400 A/s. Each .zip folder contains the following files: A .csv file containing raw data An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production, electrolysis power consumption, and wind power input. A PDF file detailing the historical wind data statistical analysis used to generate the wind profile. An experiment labeled “characterization_200.zip” demonstrates the MC250 electrolyzer steady-state response with 30-minute load steps for a total duration of 5 hours. Finally, a .csv file is provided with all simulated wind experiments combined into one dataset labeled "combined_historical_wind_experiments.csv". NLR also built an AI/machine-learning predictive model based on these datasets. The model ingests the electrolyzer current command in amperes, as well as various pressures and temperatures across the system, and predicts hydrogen output in kilograms per hour. The complete model can be found at https://huggingface.co/NatLabRockies/ptmelt-hydrogen-electrolysis [1] nelhydrogen.com/product/mc-series-electrolyser .

08 HYDROGEN

Periodicity significance testing with null-signal templates: reassessment of PTF’s SMBH binary candidates

Periodograms are widely employed for identifying periodicity in time series data, yet they often struggle to accurately quantify the statistical significance of detected periodic signals when the data complexity precludes reliable simulations. We develop a data-driven approach to address this challenge by introducing a null-signal template (NST). The NST is created by carefully randomizing the period of each cycle in the periodogram template, rendering it non-periodic. It has the same frequentist properties as a periodic signal template, and we show with simulations that the distribution of false positives is the same as with the original periodic template, regardless of the underlying data. Thus, performing a periodicity search with the NST acts as an effective simulation of the null (no-signal) hypothesis, without having to simulate the noise properties of the data. We apply the NST method to the supermassive black hole binaries (SMBHB) search in the Palomar Transient Factory (PTF), where Charisi et al. had previously proposed 33 high signal-to-noise candidates utilizing simulations to quantify their significance. Our approach reveals that these simulations do not capture the complexity of the real data. There are no statistically significant periodic signal detections above the non-periodic background. To improve the search sensitivity, we introduce a Gaussian quadrature based algorithm for the Bayes Factor with correlated noise as a test statistic. We show with simulations that this improves sensitivity to true signals by more than an order of magnitude. However, the Bayes Factor approach also results in no statistically significant detections in the PTF data.

79 ASTRONOMY AND ASTROPHYSICS

Hierarchical semi-Markov models with duration-aware dynamics for activity sequences

Residential electricity demand at granular scales is driven by what people do and for how long. Accurately forecasting this demand for applications like microgrid management and demand response therefore requires generative models for activities that can produce realistic daily activity sequences, capturing both the timing and duration of human behavior. This paper develops a generative model of human activity sequences using nationally representative time-use diaries at a 10-min resolution. We use this model to quantify which demographic factors are most critical for improving predictive performance. We propose a hierarchical semi-Markov framework that addresses two key modeling challenges. First, a time-inhomogeneous Markov router learns the patterns of “which activity comes next.” Second, a semi-Markov hazard component explicitly models activity durations, capturing “how long” activities realistically last. To ensure statistical stability when data are sparse, the model pools information across related demographic groups and time blocks. The entire framework is trained and evaluated using survey design weights to ensure our findings are representative of the U.S. population. On a held-out test set, we demonstrate that explicitly modeling durations with the hazard component provides a substantial and statistically significant improvement over purely Markovian models. Furthermore, our analysis reveals a clear hierarchy of demographic factors: Sex, Day-Type, and Household Size provide the largest predictive gains, while Region and Season, though important for energy calculations, contribute little to predicting the activity sequence itself. The result is an interpretable and robust generator of synthetic activity traces, providing a high-fidelity foundation for downstream energy systems modeling.

24 POWER TRANSMISSION AND DISTRIBUTION

Inference of the linear matter power spectrum at z = 0 using DESI DR1 Full-Shape data

Measurements of galaxy distributions at large cosmic distances capture clustering from the past. In this study, we use a cosmological model to translate these observations into the present-day galaxy distribution. Specifically, we reconstruct the 3D linear matter power spectrum at redshift z = 0 using Dark Energy Spectroscopic Instrument (DESI) Year 1 (DR1) galaxy clustering data and Cosmic Microwave Background (CMB) observations, assuming the ΛCDM model, and compare it to the result assuming the w 0 w a CDM model. Building on previous state-of-the-art methods, we apply Effective Field Theory (EFT) modelling of the galaxy power spectrum to account for small-scale effects in the 2-point statistics of galaxy data. Implementation of the EFT approach improves the modelling of the galaxy power spectrum, providing a more robust consistency test of the assumed cosmological model. By casting both CMB and galaxy clustering observations, spanning distinct redshift regimes, into k-space, we can identify discrepancies between the datasets of different redshifts, which would indicate potential inaccuracies in the assumed expansion history. While previous studies have shown consistency with ΛCDM, this work extends the analysis with higher-quality data to further test the expansion histories of both ΛCDM and w 0 w a CDM. Our findings show that both ΛCDM and w 0 w a CDM provide consistent fits to the linear matter power spectrum recovered from DESI DR1 data.

cosmological parameters from LSS

Exclusive η Electro-Production Beam Spin Asymmetry Measurements using CLAS12 at Jefferson Lab

The exploration of nucleon structure and electromagnetic transitions from ground-state to excited-state is a cornerstone of nuclear physics research. Meson electro-production experiments have opened new avenues for investigating these phenomena, particularly in the 12 GeV era at Jefferson Lab with the CLAS12 spectrometer. The ?N final states, accessible only through isospin resonances I = 1/2, provide a unique tool for studying nucleon excitations. By simplifying the analysis and enabling a cleaner extraction of resonance properties compared to the extensively studied ?N final states, ? electroproduction offers a complementary approach to unraveling the structure of excited nucleons. This work presents the first-ever measurement of the beam spin asymmetry (BSA) in exclusive ? electroproduction, covering a previously unexplored kinematic region with 1.6 ? W ? 2.2 GeV. The BSA is extracted from the CLAS12 data using a comprehensive analysis framework that carefully considers the statistical limitations of the data set. The results are compared to predictions from theoretical models, such as the Jülich-Bonn-Washington (JBW) and MAID, as well as compared to previously published cross-section and spin observable results from CLAS and SLAC. Notably, the extracted BSA exhibits discrepancies with the model predictions, highlighting the potential for refining theoretical descriptions of nucleon resonances and their electromagnetic couplings through the incorporation of these new data. The high precision data obtained in this previously unmeasured kinematic region now serve as valuable input for theorists to refine their models.

Illari, Isabella