Search NASA⌕ Search

SEARCH · Search NASA

Results for “data distributions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

CHELAX-BNF: Turbulent Parameters by airborne measurements

The original data were collected on board the ARM Aerial Facility ArcticShark uncrewed aerial system (UAS; https://www.arm.gov/capabilities/observatories/aaf/uas ) during the “Characterizing HEterogeneous Land-Atmosphere eXchanges at BNF” field campaign (CHEAX-BNF; https://arm.gov/research/campaigns/aaf2025CHELAX-BNF ). The ARM Aerial Facility ArcticShark UAS was based at the public-use airport of Posey Field, AL (FAA LID: 1M4, 34.28027778° N, 87.60055556° W, 283m MSL) from May 28 through June 23, 2025. The ArcticShark UAS performed 5 flights, including 4 research flights over the BNF Main Site (ARM Mobile Facility 3, https://arm.gov/capabilities/observatories/amf ) and Supplemental Facilities to measure atmospheric state, turbulence, surface IR temperature and imagery, aerosol number concentration, and aerosol size distribution. The current data set presents a collection of turbulent parameters in the atmospheric boundary layer or lower free troposphere based on airborne measurement throughout the field campaign. The primary instruments used to create the current data set were the Aircraft Integrated Meteorological Measurement System (AIMMS-30) and the fine-wire thermocouple probe.

Aircraft Integrated Meteorological Measurement Sys↗

Continuous-variable quantum Boltzmann machine

Here, we propose a continuous-variable quantum Boltzmann machine (CVQBM) using a powerful energy-based neural network. It can be realized experimentally on a continuous-variable (CV) photonic quantum computer. We used a CV quantum imaginary time evolution (QITE) algorithm to prepare the essential thermal state and then designed the CVQBM to proficiently generate continuous probability distributions. We applied our method to both classical and quantum data. Using real-world classical data, such as synthetic-aperture radar (SAR) images, we generated probability distributions. For quantum data, we used the output of CV quantum circuits. We obtained high fidelity and low Kullback–Leibler (KL) divergence showing that our CVQBM learns distributions from given data well and generates data sampling from that distribution efficiently. We also discussed the experimental feasibility of our proposed CVQBM. Our method can be applied to a wide range of real-world problems by choosing an appropriate target distribution (corresponding to, e.g., SAR images, medical images, and risk management in finance). Moreover, our CVQBM is versatile and could be programmed to perform tasks beyond generation, such as anomaly detection.

SAR images↗

NPFTURBULENCE: Turbulent Parameters by airborne measurements

The original data were collected during the field campaign of “Turbulent layers promoting New Particle Formation” experiment (NPFTURBULENCE; https://www.arm.gov/research/campaigns/aaf2024npfturbulence) over the Atmospheric Radiation Measurement (ARM) user facility's Southern Great Plains (SGP) atmospheric observatory (https://www.arm.gov/capabilities/observatories/sgp ) in north-central Oklahoma. The ARM Aerial Facility ArcticShark uncrewed aerial system (UAS, https://www.arm.gov/capabilities/observatories/aaf/uas) was based at Blackwell–Tonkawa Municipal Airport (IATA: BWL, ICAO: KBKN, FAA LID: BKN, 36.74475° N, 97.34918° W, 313.9 m MSL), for the field campaign from May 5 through May 29, 2024. The ArcticShark UAS performed 11 flights, including 10 research flights over the Central Facility of the ARM SGP to measure atmospheric state, turbulence, surface IR temperature and imagery, and aerosol number concentration and size distribution. The current data set presents a comprehensive collection of turbulent parameters in the atmospheric boundary layer or lower free troposphere based on airborne measurement throughout the field campaign. The primary instruments used to create the current data set were the Aircraft Integrated Meteorological Measurement System (AIMMS-30) and the fine-wire thermocouple probe. For user convenience, the current data set includes several parameters commonly used in turbulent research for normalization and/or scaling: atmospheric boundary-layer height, surface conditions, convective scales for temperature, and velocity, etc.

54 ENVIRONMENTAL SCIENCES↗

Data-driven high-dimensional statistical inference with generative models

Crucial to many measurements at the LHC is the use of correlated multi-dimensional information to distinguish rare processes from large backgrounds, which is complicated by the poor modeling of many of the crucial backgrounds in Monte Carlo simulations. In this work, we introduce HI-SIGMA, a method to perform unbinned high-dimensional statistical inference with data-driven background distributions. In contradistinction to many applications of Simulation Based Inference in High Energy Physics, HI-SIGMA relies on generative ML models, rather than classifiers, to learn the signal and background distributions in the high-dimensional space. These ML models allow for interpretable inference while also incorporating model errors and other sources of systematic uncertainties. We showcase this methodology on a simplified version of a di-Higgs measurement in the bbγγ final state, where the di-photon resonance allows for background interpolation from sidebands into the signal region. We demonstrate that HI-SIGMA provides improved sensitivity as compared to standard classifier-based methods, and that systematic uncertainties can be straightforwardly incorporated by extending methods which have been used for histogram based analyses.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Data efficiency assessment of generative adversarial networks in energy applications

This study investigates the data requirements of generative artificial intelligence (AI), particularly generative adversarial networks (GANs), for reliable data augmentation in energy applications. Generative AI, though seen as a solution to data limitations, requires substantial data to learn meaningful distributions—a challenge often overlooked. This study addresses the challenge through synthetic data generation for critical heat flux (CHF) and power grid demand, focusing on renewable and nuclear energy. Two variants of GAN employed are conditional GAN (cGAN) and Wasserstein GAN (wGAN). Our findings include the strong dependency of GAN on data size, with performance declining on smaller datasets and varying performance when generalizing to unseen experiments. Mass flux and heated length significantly influence CHF predictions. wGAN is more robust to feature exclusion, making it suitable for constrained synthetic data generation. In energy demand forecasting, wGAN performed well for solar, wind, and load predictions. Longer lookback hours and larger datasets improved predictions, especially for load power. Seasonal variations posed challenges, with wGAN achieving a relatively high error of Root Mean Squared Error (RMSE) of 0.32 for load power prediction, compared to RMSE of 0.07 under same-season conditions. Feature exclusions impacted cGAN the most, while wGAN showed greater robustness. This study concludes that, while generative AI is effective for data augmentation, it requires substantial data and careful training to generate realistic synthetic data and generalize to new experiments in engineering applications.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Collins function for pion-in-jet production in polarized pp collisions: a test of universality and factorization

We present an updated study of the Collins azimuthal asymmetries for pion-in-jet production in polarized pp collisions. To this aim, we employ a recent extraction of the transversity and Collins fragmentation functions from semi-inclusive deep inelastic scattering and e + e - annihilation into hadron pairs processes, obtained within a simplified transverse momentum dependent (TMD) approach at leading order in the strong coupling constant α s . In the present case we adopt a collinear configuration for the initial state, keeping transverse momentum effects only in the fragmentation mechanism. Our theoretical estimates, when compared against 200 GeV and 510 GeV data from the STAR Collaboration, show a generally good agreement for the distributions in the transverse momentum of the jet, the pion longitudinal momentum fraction and its transverse momentum with respect to the jet direction. While not being a proof, due the assumptions and limitations behind the entire approach, these findings corroborate the hypothesis of TMD factorization for such processes as well as of the universality of the Collins function and, once again, of a reduced impact of the proper TMD evolution on azimuthal asymmetries. We will also present predictions based on an extraction of the Collins and transversity distributions where information from data on single spin asymmetry for inclusive pion production in p ↑ p collisions is included through a Bayesian reweighting procedure.

Azimuthal asymmetries↗

Elucidating the geometric and electronic structure of a fully sulfided analog of an Anderson polyoxomolybdate cluster

The catalytic activity of transition metal sulfide (TMS) clusters in small molecule activation, redox transformations, and charge transfer has inspired the design of novel TMS-based materials for energy-related catalysis and chemical applications. Polyoxometalates (POMs), known for their structural diversity, can in principle be transformed into TMS clusters; however, fully sulfided analogs are rarely isolated, likely due to the strong tendency of uncapped TMS clusters to agglomerate. Here, we report the geometric and electronic structure of a capping ligand-free fully sulfided analog of heptamolybdate Anderson POM [Mo VI 7 O 24 ] 6− , synthesized through the sulfidation of a nanoconfined POM secured within a porous Zr-metal organic framework (NU-1000). A combined computational and experimental analysis indicates that the sulfided counterpart of the Anderson POM is geometrically and electronically more sophisticated than the parent POM. Comparison of experimental pair distribution function (PDF) data with computational simulations confirms that, unlike the oxygen-only [Mo VI 7 O 24 ] 6− cluster, the [Mo IV 7 (μ 3 -S) 6 (μ 2 -SH) 6 (S 2 ) 6 ] 2− polythiometalate (PTM) exhibits diverse sulfur anions (S 2− , HS − , S 2 2− ). DFT calculations indicate that H 2 S acts as a reducing agent, and together with terminal disulfide (S 2 2− ) ligands in the PTM structure, facilitates the complete reduction of all seven Mo VI centers in the parent POM to Mo IV . These findings are supported by X-ray photoelectron spectroscopy (XPS), which confirms exclusive Mo IV , and elemental analysis, which shows quantitative sulfur incorporation. Difference envelope density (DED) mapping further reveals that the PTM clusters are spatially confined within the MOF pores, preventing agglomeration and preserving molecular integrity.

Rabbani, S. M. Gulam [The Ohio State University, C↗

Impact of survey spatial variability on galaxy redshift distributions and the cosmological 3 × 2-point statistics for the Rubin Legacy Survey of Space and Time (LSST)

We investigate the impact of spatial survey non-uniformity on the galaxy redshift distributions for forthcoming data releases of the Rubin Observatory Legacy Survey of Space and Time (LSST). Specifically, we construct a mock photometry data set degraded by the Rubin OpSim observing conditions, and estimate photometric redshifts of the sample using a template-fitting photo-z estimator, BPZ, and a machine learning method, FlexZBoost. We select the Gold sample, defined as $i\lt 25.3$ for 10 yr LSST data, with an adjusted magnitude cut for each year and divide it into five tomographic redshift bins for the weak lensing lens and source samples. We quantify the change in the number of objects, mean redshift, and width of each tomographic bin as a function of the coadd i-band depth for 1-yr (Y1), 3-yr (Y3), and 5-yr (Y5) data. In particular, Y3 and Y5 have large non-uniformity due to the rolling cadence of LSST, hence provide a worst-case scenario of the impact from non-uniformity. We find that these quantities typically increase with depth, and the variation can be $10\!-\!40~{{\rm per\ cent}}$ at extreme depth values. Using Y3 as an example, we propagate the variable depth effect to the weak lensing $3\times 2$ pt analysis, and assess the impact on cosmological parameters via a Fisher forecast. We find that galaxy clustering is most susceptible to variable depth, and non-uniformity needs to be mitigated below 3 per cent to recover unbiased cosmological constraints. There is little impact on galaxy–shear and shear–shear power spectra, given the expected LSST Y3 noise.

cosmology↗

Operando pair distribution function analysis of nanocrystalline functional materials: the case of TiO 2 -bronze nanocrystals in Li-ion battery electrodes

Structural modelling of operando pair distribution function (PDF) data of complex functional materials can be highly challenging. To aid the understanding of complex operando PDF data, this article demonstrates a toolbox for PDF analysis. The tools include denoising using principal component analysis together with the structureMining , similarityMapping and nmfMapping apps available through the online service `PDF in the cloud' ( PDFitc , https://pdfitc.org/). The toolbox is used for both ex situ and operando PDF data for 3 nm TiO 2 -bronze nanocrystals, which function as the active electrode material in a Li-ion battery. The tools enable structural modelling of the ex situ and operando PDF data, revealing two pristine TiO 2 phases (bronze and anatase) and two lithiated Li x TiO 2 phases (lithiated versions of bronze and anatase), and the phase evolution during galvanostatic cycling is characterized.

Chemistry↗

Gradient Coding With Iterative Block Leverage Score Sampling

Gradient coding is a method for mitigating straggling servers in a centralized computing network that uses erasure-coding techniques to distributively carry out first-order optimization methods. Randomized numerical linear algebra uses randomization to develop improved algorithms for large-scale linear algebra computations. In this study, we propose a method for distributed optimization that combines gradient coding and randomized numerical linear algebra. The proposed method uses a randomized ℓ 2 -subspace embedding and a gradient coding technique to distribute blocks of data to the computational nodes of a centralized network, and at each iteration the central server only requires a small number of computations to obtain the steepest descent update. The novelty of our approach is that the data is replicated according to importance scores, called block leverage scores, in contrast to most gradient coding approaches that uniformly replicate the data blocks. Furthermore, we do not require a decoding step at each iteration, avoiding a bottleneck in previous gradient coding schemes. We show that our approach results in a valid ℓ 2 -subspace embedding, and that our resulting approximation converges to the optimal solution.

97 MATHEMATICS AND COMPUTING↗

I/O in Machine Learning Applications on HPC Systems: A 360-degree Survey

Growing interest in Artificial Intelligence (AI) has resulted in a surge in demand for faster methods of Machine Learning (ML) model training and inference. This demand for speed has prompted the use of high performance computing (HPC) systems that excel in managing distributed workloads. Because data is the main fuel for AI applications, the performance of the storage and I/O subsystem of HPC systems is critical. In the past, HPC applications accessed large portions of data written by simulations or experiments or ingested data for visualizations or analysis tasks. ML workloads perform small reads spread across a large number of random files. This shift of I/O access patterns poses several challenges to modern parallel storage systems. In this paper, we survey I/O in ML applications on HPC systems, and target literature within a 6-year time window from 2019 to 2024. We define the scope of the survey, provide an overview of the common phases of ML, review available profilers and benchmarks, examine the I/O patterns encountered during offline data preparation, training, and inference, and explore I/O optimizations utilized in modern ML frameworks and proposed in recent literature. Lastly, we seek to expose research gaps that could spawn further R&D.

97 MATHEMATICS AND COMPUTING↗

A Probabilistic Reasoner Based on Bayes Risk for Damage Detection in Structural Systems

Structural health monitoring (SHM) systems are used to inform operation of structural systems subject to loads and environments that may affect their integrity. SHM systems rely on continuous monitoring of the structure to determine its health state. These systems are often coupled with a model of the deployed structure to determine the consequences of changes in the system by forecasting the response to future states. These models, which may be thought of as digital twins, need to be updated to reflect the latest state of the structural system. This work makes use of an uncertainty-aware machine learning model that enforces distance preservation of the original input space to determine deviations from the training data input space distributions. This workflow enables domain shift detection to determine whether damage is present in the structure. The uncertainty metrics generated by this network are then used in a Bayes risk framework to design an optimal damage detector given cost and risk considerations. The approach is demonstrated on a computational example with simulated damage.

Najera-Flores, David [ATA Engineering, Inc.]↗

Full Rod Length Gamma Spectroscopy Measurements of LWR Used Fuel Rods

Twenty-five high-burnup pressurized water reactor used fuel rods were scanned in the Irradiated Fuels Examination Laboratory at Oak Ridge National Laboratory with a high-resolution gamma spectrometer. A planar high-purity germanium detector was used across an energy range from 60 keV to 3 MeV. Gamma spectra of individual fuel rods were collected in 4.5 mm steps along the entire length of each rod. Additional spectra were collected on select rods in rotational increments of 30 degrees to provide data on the distribution of fission products along the rod perimeter.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

1988/1989 Maricopa Household Travel Study

This study was commissioned by the Maricopa Association of Governments (MAG) Transportation and Planning Office. The primary objectives of this study were to update the trip generation rates used in the MAG travel demand forecasting process and to provide data to validate the MAG trip distribution model. Demographic, socioeconomic, and travel data was collected for 2,992 households residing within the MAG Urban Planning Area. Respondents were asked to record their travel and activities for a 24-hour period. A total of 26,733 trips were recorded by 6,463 people.

1Hz data↗

Representing Complex Systems as Graphs for Debugging and Predictive Maintenance-Preliminary Thoughts

Representing complex systems as graphs enables use of mathematical tools to identify faults or predict failures. Graph nodes correspond to individual modules or subsystems, and edges link coupled system parts. ‘Probes’ measure the node outputs, monitoring the system health for unexpected behavior. Assuming one cannot probe every point, within a system, the fault correlates to a region—not necessarily the specific location. Bayesian networks trained to understand fault patterns can accurately identify the source. The diagnostic tool described aides debugging by pinpointing system failure causes. For predictive maintenance, probe data develop probability distribution functions describing subsystem mean time to failure. Unit lifetime can be estimated through these probability distributions. Two approaches include using Bayesian classifiers to infer the system failure source and developing maintenance schedules by treating systems as collections of random variables. When failure behavior does not follow a closed form function, use of similarity models is proposed.

97 MATHEMATICS AND COMPUTING↗

"Isotopics by nuclides" tool performance in InterSpec

This report provides a performance summary of the “Isotopics by Nuclides” tool in InterSpec. The primary objective of the NA-241 FY25 project was the detection and characterization of non-homogeneous uranium samples. However, the tool also demonstrates effectiveness in determining the enrichment of homogeneous uranium or plutonium from single gamma spectra. This paper presents a subset of evaluated data to avoid distribution limitations while offering potential users’ insight into the tool's performance.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Global QCD analysis of spin PDFs in the proton with high-𝑥 and lattice constraints

We perform a comprehensive global QCD analysis of spin-dependent parton distribution functions (PDFs), combining all available data on inclusive and semi-inclusive deep-inelastic scattering (DIS), as well as inclusive weak boson and jet production in polarized 𝑝⁢𝑝 collisions, simultaneously extracting spin-averaged PDFs and fragmentation functions. Including recent Jefferson Lab DIS data at high 𝑥, together with subleading power corrections to the leading-twist framework, allows us to verify the stability of the PDFs for 𝑊 2 ≥ 4 GeV 2 and quantify the uncertainties on the spin structure functions more reliably. We explore the use of new lattice QCD data on gluonic pseudo-Ioffe time distributions, which, together with jet production and high-𝑥 DIS data, improve the constraints on the polarized gluon PDF. The expanded kinematic reach afforded by the data into the high-𝑥 region allows us to refine the bounds on higher-twist contributions to the spin structure functions, and test the validity of the Bjorken sum rule.

Perturbative QCD↗

Wasserstein normalized autoencoder for anomaly detection

A novel anomaly detection algorithm is presented. The Wasserstein normalized autoencoder (WNAE) is a normalized probabilistic model that minimizes the Wasserstein distance between the learned probability distribution—a Boltzmann distribution where the energy is the reconstruction error of the autoencoder (AE)—and the distribution of the training data. This algorithm has been developed and applied to the identification of semivisible jets—conical sprays of visible standard model (SM) particles and invisible dark matter states—with the CMS experiment at the CERN LHC. Trained on jets of particles from simulated SM processes, the WNAE is shown to learn the probability distribution of the input data in a fully unsupervised fashion, such that it effectively identifies new physics jets as anomalies. The model exhibits stable, convergent training and recovers strong classification performance for a wide range of signals against the selected background process, for which a standard AE fails because of outlier reconstruction. In addition, the model improves upon standard normalized autoencoders while remaining fully agnostic to the signal. The WNAE directly tackles the problem of outlier reconstruction, a common failure mode of autoencoders in anomaly detection tasks.

Hayrapetyan, Aram [Yerevan Phys. Inst.]↗