Search NASA⌕ Search

SEARCH · Search NASA

Results for “data statistics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

TPCpp-10M: Simulated proton-proton collisions in a time projection chamber for AI foundation models

Scientific foundation models hold great promise for advancing nuclear and particle physics by improving analysis precision and accelerating discovery. Yet, progress in this field is often limited by the lack of openly available large scale datasets, as well as standardized evaluation tasks and metrics. Furthermore, the specialized knowledge and software typically required to process particle physics data pose significant barriers to interdisciplinary collaboration with the broader machine learning community. This work introduces a large, openly accessible dataset of 10 million simulated proton-proton collisions, designed to support self-supervised training of foundation models. To facilitate ease of use, the dataset is provided in a common NumPy format. In addition, it includes 70,000 labeled examples spanning three well defined downstream tasks: track finding, particle identification, and noise tagging, to enable systematic evaluation of the foundation model's adaptability. The simulated data are generated using the Pythia Monte Carlo event generator at a center of mass energy of $\sqrt{s}$ = 200 GeV and processed with Geant4 to include realistic detector conditions and signal emulation in the sPHENIX Time Projection Chamber at the Relativistic Heavy Ion Collider, located at Brookhaven National Laboratory. This dataset resource establishes a common ground for interdisciplinary research, enabling machine learning scientists and physicists alike to explore scaling behaviors, assess transferability, and accelerate progress toward foundation models in nuclear and high energy physics. The complete simulation and reconstruction chain is reproducible with the sPHENIX software stack. All data and code locations are provided under Data Accessibility.

Data Analysis, Statistics and Probability (physics↗

Autonomous platform for solution processing of electronic polymers

The manipulation of electronic polymers’ solid-state properties through processing is crucial in electronics and energy research. Yet, efficiently processing electronic polymer solutions into thin films with specific properties remains a formidable challenge. We introduce Polybot, an artificial intelligence (AI) driven automated material laboratory designed to autonomously explore processing pathways for achieving high-conductivity, low-defect electronic polymers films. Leveraging importance-guided Bayesian optimization, Polybot efficiently navigates a complex 7-dimensional processing space. In particular, the automated workflow and algorithms effectively explore the search space, mitigate biases, employ statistical methods to ensure data repeatability, and concurrently optimize multiple objectives with precision. The experimental campaign yields scale-up fabrication recipes, producing transparent conductive thin films with averaged conductivity exceeding 4500 S/cm. Feature importance analysis and morphological characterizations reveal key design factors. This work signifies a significant step towards transforming the manufacturing of electronic polymers, highlighting the potential of AI-driven automation in material science.

Wang, Chengshi [Argonne National Laboratory (ANL),↗

Observation of Kolmogorov turbulence due to multiscale vortices in dusty plasma experiments

We report the experimental observation of fully developed Kolmogorov turbulence originating from self-excited vortex flows in a three-dimensional (3D) dust cloud. The characteristic -5/3 scaling of 3D Kolmogorov turbulence is consistent in both the spatial and temporal energy spectra within a statistical variation of experimental data. Additionally, the 2/3 scaling in the second-order structure function further supports the presence of Kolmogorov turbulence. We also identified a slight deviation in the tails of the probability distribution functions for velocity gradients, a reflection of intermittency. The experiment showed the formation of a dust cloud in the diffused plasma region away from the electrodes. The dust rotation was observed in multiple experimental campaigns under different discharge conditions at different spatial locations and background plasma environments.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A hybrid neural architecture: Online attosecond x-ray characterization

The emergence of high-repetition-rate x-ray free-electron lasers (XFELs), such as SLAC’s LCLS-II, serves as our canonical example for autonomous controls that necessitate high-throughput diagnostics paired with streaming computational pipelines capable of single-shot analysis with extremely low latency. We present the deterministic characterization with an integrated parallelizable hybrid resolver architecture, a hybrid machine learning framework designed for fast, accurate analysis of XFEL diagnostics using angular streaking-based sinogram images. This architecture integrates convolutional neural networks and bidirectional long short-term memory models to denoise input, identify x-ray sub-spike features, and extract sub-spike relative delays with sub-30 attosecond temporal resolution. Deployed on low-latency hardware, it achieves over 10 kHz throughput with 168.3 μs inference latency, indicating scalability to 14 kHz with field-programmable gate array integration. By transforming regression tasks into classification problems and leveraging optimized error encoding, we achieve high precision with low-latency performance that is critical for real-time streaming event selection and experimental control feedback signals. This represents a key development in real-time control pipelines for next-generation autonomous science, generally, and high repetition-rate x-ray experiments in particular.

Accelerator Physics (physics.acc-ph)↗

Optimal binning of correlated measurements

Experimental measurements are commonly represented on a discrete grid, requiring a balance between granularity and statistical noise. Two strategies have traditionally been used to improve such representations: selecting an appropriate bin width to control discretization error and applying kernel-based smoothing to suppress fluctuations. Despite their shared goal, these approaches have largely developed independently, without a unified statistical description of how discretization and correlation jointly determine measurement precision. Here, we extend the discussion of optimal interval averaging to a correlation-aware setting by Gaussian process regression, which explicitly accounts for correlations among neighboring bins. Starting from first principles, we derive the mean-squared error of discretized measurements and obtain closed-form asymptotic expressions for the optimal bin width and correlation length. When recast in reduced variables, the theory reveals distinct universal scaling laws governing the error in the correlation-free and correlation-controlled regimes. Characterized by intrinsically smooth intensity profiles and counting-based statistics, neutron scattering measurements are well suited for demonstrating the enhanced error contraction enabled by inter-bin correlations. We show that such improvement is achievable over the experimentally accessible Q-range and across multiple instruments and material systems. These results show that explicitly accounting for correlations systematically reshapes the limits of precision in discretized, noise-limited measurements. More broadly, the framework provides a transferable statistical foundation for optimizing data representation, inference, and experimental design across the physical and data sciences.

Tung, Chi-Huan [ORNL] (ORCID:0000000221972074)↗

Optimizing spin dressing sensitivity for the nEDMSF experiment

nEDMSF aims to measure the neutron electric dipole moment (d n ) with unprecedented precision. In this paper we explore the experiment's sensitivity when operating with an implementation of the critical dressing method in which the angle between the neutron and Helium-3 spins (ϕ 3n ) is subjected to a square modulation by an amount ϕ d (the “dressing angle”). Several parameters can be tuned to optimize sensitivity. We find roughly 10% improvement over a previous estimate, resulting primarily from the addition of a waiting period between the π/2 pulse that initiates d n -driven ϕ 3n growth and the start of ϕ3n modulation. We find negligible further improvement by allowing ϕ d to vary continuously over the course of a run, and no degradation resulting from the addition of an in situ background measurement into each ϕ3n modulation sequence. A complete simulation confirms a 300 live-day sensitivity ofσ = 1.45×10 -28 e ·cm. At this level of sensitivity, σ ϕ3n0 = 1 mrad precision on the initial n/ 3 He angle difference is not negligible.

47 OTHER INSTRUMENTATION↗

Tools for unbinned unfolding

Machine learning has enabled differential cross section measurements that are not discretized. Going beyond the traditional histogram-based paradigm, these unbinned unfolding methods are rapidly being integrated into experimental workflows. Here, in order to enable widespread adaptation and standardization, we develop methods, benchmarks, and software for unbinned unfolding. For methodology, we demonstrate the utility of boosted decision trees for unfolding with a relatively small number of high-level features. This complements state-of-the-art deep learning models capable of unfolding the full phase space. To benchmark unbinned unfolding methods, we develop an extension of existing dataset to include acceptance effects, a necessary challenge for real measurements. Additionally, we directly compare binned and unbinned methods using discretized inputs for the latter in order to control for the binning itself. Lastly, we have assembled two software packages for the OmniFold unbinned unfolding method that should serve as the starting point for any future analyses using this technique. One package is based on the widely-used RooUnfold framework and the other is a standalone package available through the Python Package Index (PyPI).

47 OTHER INSTRUMENTATION↗

Neural posterior unfolding

Differential cross section measurements are the currency of scientific exchange in particle and nuclear physics. A key challenge for these analyses is the correction for detector distortions, known as deconvolution or unfolding. Binned unfolding of cross section measurements traditionally rely on the regularized inversion of the response matrix that represents the detector response, mapping pre-detector (`particle level') observables to post-detector (`detector level') observables. In this paper we introduce Neural Posterior Unfolding, a modern, Bayesian approach that leverages normalizing flows for unfolding. By using normalizing flows for neural posterior estimation, NPU offers several key advantages including implicit regularization through the neural network architecture, fast amortized inference that eliminates the need for repeated retraining, and direct access to the full uncertainty in the unfolded result. In addition to introducing NPU, we implement a classical Bayesian unfolding method called Fully Bayesian Unfolding (FBU) in modern Python so it can also be studied. These tools are validated on simple Gaussian examples and then tested on simulated jet substructure examples from the Large Hadron Collider (LHC). We find that the Bayesian methods are effective and worth additional development to be analysis ready for cross section measurements at the LHC and beyond.

Analysis and statistical methods↗

Experimental search for the chiral magnetic effect in relativistic heavy-ion collisions: A perspective

The chiral magnetic effect (CME) refers to generation of the electric current along a magnetic field in a chirally imbalanced system of quarks. The latter is predicted by quantum chromodynamics to arise from quark interaction with nontrivial topological fluctuations of the vacuum gluonic field. The CME has been actively searched for in relativistic heavy-ion collisions, where such gluonic field fluctuations and a strong magnetic field are believed to be present. The CME-sensitive observables are unfortunately subject to a possibly large non-CME background, and firm conclusions on a CME observation have not yet been reached. In this perspective, we review the experimental status and progress in the CME search, from the initial measurements more than a decade ago to the dedicated program of isobar collisions in 2018 and the release of the isobar blind analysis result in 2022 to intriguing hints of a possible CME signal in Au + Au collisions, and discuss future prospects of a potential CME discovery in the anticipated high-statistic Au + Au collision data at the Relativistic Heavy-Ion Collider by 2025. We hope such a perspective will help sharpening our focus on the fundamental physics of the CME and steer its experimental search.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Imaging Bragg Edge Analysis TooLs for Engineering Structures (iBeatles)

The Spallation Neutron Source (SNS) at Oak Ridge National Laboratory (ORNL) provides pulsed neutrons with energies varying from epithermal to cold. In preparation for VENUS, the neutron imaging beamline to be located at beam port 10, we have performed a series of experiments focused on wavelength-dependent radiography and computed tomography for a broad range of applications, from materials science to biological tissues.One of the time-of-flight (TOF) techniques that is of interest to the scientific community is the 2-dimensional mapping of phases and average crystalline plane orientation in samples both ex-situ and during applied stresses such as tensile loading and heating. This technique is known as Bragg edgeimaging and relies on the identification of changes of transmission values, fitting of the edge to measure its displacement, and thus identify the shift in lattice parameter due to stresses. One of the challenges of TOF imaging measurements is the amount of data and the inability to observe Bragg edge shifts in real time during an experiment. Thus, we have been focusing on creating a Python-based interface that allows fast data processing and instantaneous mapping and fitting of the Bragg edges, and their evolution through time. Python libraries and Jupyter notebooks have been implemented to facilitate decision making during an experiment. The advantage of the notebooks is the possibility to guide an experiment as they can quickly process and display Bragg edge data. These notebooks can be used independently, or can be combined in a Python Graphical User Interface (GUI) tool called iBeatles. This interface permits visualization and fitting of the Bragg edges, and ultimately back-projects the fitting results onto the radiographs to display a strain map. Assuming data collection has sufficient statistics, the strain mapping analysis can be performed on a pixel-by-pixel basis. This development is a step forward toward a better user experience at the future VENUS beamline in terms of live feedback and productivity. Analysis that used to take days of switching between different applications can now be done in minutes within the

Bilheux, JeanChristophe [Oak Ridge National Labora↗

Full event particle-level unfolding with variable-length latent variational diffusion

The measurements performed by particle physics experiments must account for the imperfect response of the detectors used to observe the interactions. One approach, unfolding, statistically adjusts the experimental data for detector effects. Recently, generative machine learning models have shown promise for performing unbinned unfolding in a high number of dimensions. However, all current generative approaches are limited to unfolding a fixed set of observables, making them unable to perform full-event unfolding in the variable dimensional environment of collider data. A novel modification to the variational latent diffusion model (VLD) approach to generative unfolding is presented, which allows for unfolding of high- and variable-dimensional feature spaces. The performance of this method is evaluated in the context of semi-leptonic t\bar{t} t t ‾ production at the Large Hadron Collider.

Shmakov, Alexander↗

Improving Additive Manufactured Component Performance through Multi-Scale Microstructure Simulation and Process Optimization

The purpose of this project was to utilize computational tools to understand the relationships between processing, microstructure, and properties for additively manufactured (AM) aluminum alloys for automotive applications, and to provide an engineering solution for helping to optimize process conditions. The project leverages ORNL developments in computational modeling, including AM process modeling, phase-field based microstructure evolution predictions, and data analytics techniques for mapping process conditions to material outcomes. The project utilized an Al-Cu-Mn-Zr alloy as a model material for studying formation of defects and microstructural features in response to variations in process conditions. Based on both pre-existing experimental data and simulation results, statistical process maps were constructed to identify regions of process space with minimal defect formation and advantageous microstructures and properties. The software tools used for this purpose were successful disseminated to GM, who were able to successful compile the relevant HPC codes within their own computing ecosystem and perform initial calculations to reproduce ORNL results.

36 MATERIALS SCIENCE↗

Status and prospects of Muon g-2 experiment

This article reviews the muon g-2 experiment, a cornerstone in precision tests of the Standard Model of particle physics. The experiment measures the anomalous magnetic moment of the muon with unprecedented accuracy, seeking potential discrepancies between theoretical predictions and experimental results that might indicate physics beyond the Standard Model. We trace the evolution of this measurement from its beginnings at CERN in the 1960s to the current state-of-the-art experiment at Fermilab, highlighting the remarkable engineering achievements required to achieve parts-per-billion precision. Recent results from Runs 1-3 have achieved a systematic uncertainty of 70 ppb, exceeding design goals, while ongoing theoretical calculations continue to refine predictions. Despite these advances, the analysis remains statistics-limited, with continued data collection and novel experimental approaches at MUonE and J-PARC promising further insights into this fundamental physical quantity.

Karuza, Marin [INFN, Trieste; Rijeka U.] (ORCID:00↗

The Information Length Concept Applied to Plasma Turbulence

A methodology to study statistical properties of anomalous transport in fusion plasma is investigated. Three time traces generated by the full-f gyrokinetic code GKNET are analyzed for this purpose. The time traces consist of heat flux as a function of the radial position, which is studied in a novel manner using statistical methods. The simulation data exhibit transport processes with both medium and long correlation length along the radius. A typical example of a phenomenon with long correlation length is avalanches. In order to investigate the evolution of the turbulent state, two basic configurations are studied, one flux-driven and one gradient-driven with decaying turbulence. The information length concept in tandem with Boltzmann–Gibbs and Tsallis entropy is used in the investigation. It is found that the dynamical states in both flux-driven and gradient-driven cases are surprisingly similar, but the Tsallis entropy reveals differences between them. This indicates that the types of probability distribution function are nevertheless quite different since the higher moments are significantly different.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Algorithm to extract direction in 2D discrete distributions and a continuous Frobenius norm

In this study, we present a novel algorithm for determining directionality in 2D distributions of discrete data. We compare a reference dataset with a known direction to a measured dataset with an unknown direction by the Frobenius norm of the difference (FND) to find the unknown direction. To generalize this concept, we develop a continuous Frobenius norm of the difference (CFND) as a continuous analog of the FND and derive its analytical expression. By relating fitted and normalized 2D Gaussian distributions, we show that the CFND approximates the FND, and we validate this relationship with computer simulations. We find that a first-order approximation of the CFND between two similar Gaussian distributions takes the form of an absolute sine function, offering a simple analytical form with potential for specialized applications in segmented inverse beta decay (IBD) neutrino detectors, astronomy, machine learning, and more. Although this method may easily extend to 3D scalar fields, our focus here is on 2D real-valued fields as it directly applies to directionality. Our methodology consists of modeling a 2D Gaussian distribution, binning the data into a histogram, and encoding it as a square matrix. Rotating this matrix around its geometric center and comparing it to a measured dataset using the FND gives us rotational data that we fit with an absolute sine function. The location of the minimum of this fit is the angle closest to the true angle of the direction in the measured dataset. We present the derivation and discuss initial applications of the CFND in our novel algorithm, demonstrating its success in approximating directionality in 2D distributions.

Data Analysis, Statistics and Probability (physics↗

Global Carbon Budget 2024

Abstract. Accurate assessment of anthropogenic carbon dioxide (CO2) emissions and their redistribution among the atmosphere, ocean, and terrestrial biosphere in a changing climate is critical to better understand the global carbon cycle, support the development of climate policies, and project future climate change. Here we describe and synthesize datasets and methodologies to quantify the five major components of the global carbon budget and their uncertainties. Fossil CO2 emissions (EFOS) are based on energy statistics and cement production data, while emissions from land-use change (ELUC) are based on land-use and land-use change data and bookkeeping models. Atmospheric CO2 concentration is measured directly, and its growth rate (GATM) is computed from the annual changes in concentration. The global net uptake of CO2 by the ocean (SOCEAN, called the ocean sink) is estimated with global ocean biogeochemistry models and observation-based fCO2 products (fCO2 is the fugacity of CO2). The global net uptake of CO2 by the land (SLAND, called the land sink) is estimated with dynamic global vegetation models. Additional lines of evidence on land and ocean sinks are provided by atmospheric inversions, atmospheric oxygen measurements, and Earth system models. The sum of all sources and sinks results in the carbon budget imbalance (BIM), a measure of imperfect data and incomplete understanding of the contemporary carbon cycle. All uncertainties are reported as ±1σ. For the year 2023, EFOS increased by 1.3 % relative to 2022, with fossil emissions at 10.1 ± 0.5 GtC yr−1 (10.3 ± 0.5 GtC yr−1 when the cement carbonation sink is not included), and ELUC was 1.0 ± 0.7 GtC yr−1, for a total anthropogenic CO2 emission (including the cement carbonation sink) of 11.1 ± 0.9 GtC yr−1 (40.6 ± 3.2 GtCO2 yr−1). Also, for 2023, GATM was 5.9 ± 0.2 GtC yr−1 (2.79 ± 0.1 ppm yr−1; ppm denotes parts per million), SOCEAN was 2.9 ± 0.4 GtC yr−1, and SLAND was 2.3 ± 1.0 GtC yr−1, with a near-zero BIM (−0.02 GtC yr−1). The global atmospheric CO2 concentration averaged over 2023 reached 419.31 ± 0.1 ppm. Preliminary data for 2024 suggest an increase in EFOS relative to 2023 of +0.8 % (−0.2 % to 1.7 %) globally and an atmospheric CO2 concentration increase by 2.87 ppm, reaching 422.45 ppm, 52 % above the pre-industrial level (around 278 ppm in 1750). Overall, the mean of and trend in the components of the global carbon budget are consistently estimated over the period 1959–2023, with a near-zero overall budget imbalance, although discrepancies of up to around 1 GtC yr−1 persist for the representation of annual to semi-decadal variability in CO2 fluxes. Comparison of estimates from multiple approaches and observations shows the following: (1) a persistent large uncertainty in the estimate of land-use change emissions, (2) low agreement between the different methods on the magnitude of the land CO2 flux in the northern extra-tropics, and (3) a discrepancy between the different methods on the mean ocean sink. This living-data update documents changes in methods and datasets applied to this most recent global carbon budget as well as evolving community understanding of the global carbon cycle. The data presented in this work are available at https://doi.org/10.18160/GCP-2024 (Friedlingstein et al., 2024).

Friedlingstein, Pierre (ORCID:0000000333094739)↗

Neutron detector response modeling in NOvA

Neutrons can present a significant challenge for neutrino experiments in which energy reconstruction is critical. With the ability to escape detection completely and with a weak correlation between their kinetic energy and any eventual energy deposition, it is difficult to fully account for neutrons produced in neutrino interactions. This in turn leads to significant model dependence when evaluating neutron-related systematic uncertainties. The NOvA experiment is a long-baseline neutrino oscillation experiment with a high-statistics sample of antineutrino data collected by its near detector. We report an excess relative to data of simulated neutron candidates with low energy depositions when using standard Geant4 physics lists. The simulation excess is traced to an overabundance of secondary photons produced from interactions of neutrons with kinetic energy greater than \SI{20}{\mega\eV}. Improved agreement with data is obtained by applying the data-driven neutron-on-carbon \menate model for neutrons between \SI{20}{\mega\eV} and ${\sim}$\SI{100}{\mega\eV}. With \menate, the residual oversimulation is more uniform across the calorimetric neutron energy spectrum, suggesting possible overproduction of primary neutrons by the GENIE neutrino interaction generator. These results motivate the adoption of \menate-supplemented Geant4 simulation as the nominal simulation in the production of future \nova simulation.

Abubakar, S.↗

PNNL-Predictive-Phenomics/ProteoMeter

ProteoMeter is a Python package that assists in the statistical analysis of global proteomics, protein post-translation modification (PTM), and limited proteolysis (LiP) data. It contains batch correction, normalization, and statistical testing methods, as well as functions that "roll up" peptide-level data to the single-site level. It has a robust user configuration system, allowing it to flexibly integrate different types of experiment designs. For basic usage, a simple configuration file provides the essential functionality. Advanced users have access to the entire statistical pipeline for fine-tuning analyses. Processed data is easily exported to many common spreadsheet and data-frame formats.

Rozum, Jordan [Pacific Northwest National Lab]↗