Search NASA⌕ Search

SEARCH · Search NASA

Results for “data distributions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

1988/1989 Maricopa Household Travel Study

This study was commissioned by the Maricopa Association of Governments (MAG) Transportation and Planning Office. The primary objectives of this study were to update the trip generation rates used in the MAG travel demand forecasting process and to provide data to validate the MAG trip distribution model. Demographic, socioeconomic, and travel data was collected for 2,992 households residing within the MAG Urban Planning Area. Respondents were asked to record their travel and activities for a 24-hour period. A total of 26,733 trips were recorded by 6,463 people.

1Hz data↗

Representing Complex Systems as Graphs for Debugging and Predictive Maintenance-Preliminary Thoughts

Representing complex systems as graphs enables use of mathematical tools to identify faults or predict failures. Graph nodes correspond to individual modules or subsystems, and edges link coupled system parts. ‘Probes’ measure the node outputs, monitoring the system health for unexpected behavior. Assuming one cannot probe every point, within a system, the fault correlates to a region—not necessarily the specific location. Bayesian networks trained to understand fault patterns can accurately identify the source. The diagnostic tool described aides debugging by pinpointing system failure causes. For predictive maintenance, probe data develop probability distribution functions describing subsystem mean time to failure. Unit lifetime can be estimated through these probability distributions. Two approaches include using Bayesian classifiers to infer the system failure source and developing maintenance schedules by treating systems as collections of random variables. When failure behavior does not follow a closed form function, use of similarity models is proposed.

97 MATHEMATICS AND COMPUTING↗

"Isotopics by nuclides" tool performance in InterSpec

This report provides a performance summary of the “Isotopics by Nuclides” tool in InterSpec. The primary objective of the NA-241 FY25 project was the detection and characterization of non-homogeneous uranium samples. However, the tool also demonstrates effectiveness in determining the enrichment of homogeneous uranium or plutonium from single gamma spectra. This paper presents a subset of evaluated data to avoid distribution limitations while offering potential users’ insight into the tool's performance.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Global QCD analysis of spin PDFs in the proton with high-𝑥 and lattice constraints

We perform a comprehensive global QCD analysis of spin-dependent parton distribution functions (PDFs), combining all available data on inclusive and semi-inclusive deep-inelastic scattering (DIS), as well as inclusive weak boson and jet production in polarized 𝑝⁢𝑝 collisions, simultaneously extracting spin-averaged PDFs and fragmentation functions. Including recent Jefferson Lab DIS data at high 𝑥, together with subleading power corrections to the leading-twist framework, allows us to verify the stability of the PDFs for 𝑊 2 ≥ 4 GeV 2 and quantify the uncertainties on the spin structure functions more reliably. We explore the use of new lattice QCD data on gluonic pseudo-Ioffe time distributions, which, together with jet production and high-𝑥 DIS data, improve the constraints on the polarized gluon PDF. The expanded kinematic reach afforded by the data into the high-𝑥 region allows us to refine the bounds on higher-twist contributions to the spin structure functions, and test the validity of the Bjorken sum rule.

Perturbative QCD↗

Wasserstein normalized autoencoder for anomaly detection

A novel anomaly detection algorithm is presented. The Wasserstein normalized autoencoder (WNAE) is a normalized probabilistic model that minimizes the Wasserstein distance between the learned probability distribution—a Boltzmann distribution where the energy is the reconstruction error of the autoencoder (AE)—and the distribution of the training data. This algorithm has been developed and applied to the identification of semivisible jets—conical sprays of visible standard model (SM) particles and invisible dark matter states—with the CMS experiment at the CERN LHC. Trained on jets of particles from simulated SM processes, the WNAE is shown to learn the probability distribution of the input data in a fully unsupervised fashion, such that it effectively identifies new physics jets as anomalies. The model exhibits stable, convergent training and recovers strong classification performance for a wide range of signals against the selected background process, for which a standard AE fails because of outlier reconstruction. In addition, the model improves upon standard normalized autoencoders while remaining fully agnostic to the signal. The WNAE directly tackles the problem of outlier reconstruction, a common failure mode of autoencoders in anomaly detection tasks.

Hayrapetyan, Aram [Yerevan Phys. Inst.]↗

LandScan Global 30 Arcsecond Annual Global Gridded Population Datasets from 2000 to 2022

Abstract Oak Ridge National Laboratory (ORNL) annually develops the LandScan Global (LSG) dataset, a 30 arcsecond global gridded population dataset representing global ambient human population distribution. This multivariable dasymetric model disaggregates census counts within administrative boundaries using ancillary data. Each country’s distribution reflects cultural and socioeconomic patterns; manual validations yield a unique global dataset for assessing populations at risk. For over two decades, LSG has been a standard for estimating populations at risk, aiding U.S. federal government, academia and humanitarian organizations. During disasters such as the 2004 Indian Ocean tsunami and the 2010 Haiti earthquake and geopolitical crises such as the Syrian civil war and the 2022 Russian invasion of Ukraine, LSG supported scientific and operational communities in emergency response and recovery. In 2022, LSG datasets from 2000 onward were made publicly available through ORNL’s LandScan Portal. This data descriptor details our methodology and the application of geospatial science and machine learning to geographic and demographic data, highlighting uses in urban resiliency, emergency management, disaster response, and human health and security.

Science & Technology - Other Topics↗

Neural Network‐Based Methods for Ocean Surface Wave Measurement Using Submarine Distributed Acoustic Sensing (DAS)

Two new data-driven models for estimating ocean surface waves from distributed acoustic sensing (DAS) submarine cable strain rate are developed using supervised machine learning on a 10-day data set collected offshore of Oliktok Point, Alaska. The new models were trained on target data from seafloor pressure moorings at three sites spaced evenly along 27.1 km of cable and were benchmarked against an empirical transfer function method previously used to estimate waves from DAS. A model which uses convolutional neural networks to transform 2-km frequency-wavenumber strain spectra to seafloor pressure spectra outperforms the benchmark in wave height prediction (RMSE of 0.15 vs. 0.41 m) and period prediction (0.29 vs. 0.37 s) when evaluated on a held-out test data set. When applied to a DAS data set collected on the same cable 2 years prior, the CNN-based model maintained similar significant wave height performance (RMSE = 0.23 m) relative to available satellite altimetry data. A two-hidden-layer, fully connected neural network which transforms 1-D strain spectra to seafloor pressure spectra also outperforms the benchmark in wave height prediction (RMSE of 0.19 vs. 0.41 m), but does not generalize as well to the prior data. Regression-based machine learning is useful for estimating waves from DAS data when the pressure-strain relationship varies temporally and spatially across different wave conditions. Models can be applied to DAS data to measure waves with higher spatial resolution and longer temporal coverage than traditional methods, which often measure waves only at a single point.

Davis, Jacob R. [Univ. of Washington, Seattle, WA ↗

Estimating IDP Origins Using ACLED Political Violence Events: Validation with Lebanon IDP Flow Data

A key input in modeling population distribution and flows following conflict events is the inclusion of Internally Displaced Person (IDP) flows between administrative units within a country. These flows are critical for capturing population movement and redistribution driven by current events, particularly conflict. In some cases, IDP destination data are available while origin data is incomplete or unavailable. This creates a gap in understanding where displacement is occurring, limiting the ability to model population redistribution accurately. Without origin data, it is not possible to reallocate population flows or accurately represent where displacement is occurring within the country. This report evaluates whether Armed Conflict Location Event Data (ACLED) political violence event data can be used to estimate IDP origin distributions when direct origin data are unavailable. The approach is validated using historical IDP flow data from Lebanon, where both origins and destinations are observed. Results show that ACLED event distributions strongly correspond to observed IDP-origin patterns, particularly when using cumulative 60-day event windows. The method is most reliable for identifying major origin districts and approximating proportional origin shares. However, it is not intended to reconstruct exact individual displacement flows, but rather to provide a probabilistic spatial allocation of displacement origins.

99 GENERAL AND MISCELLANEOUS↗

Measurement of three-dimensional inclusive muon-neutrino charged-current cross sections on argon with the MicroBooNE detector

We report the measurement of the triple-differential cross section d 3 σ/dE vis d cos(θ μ )dP μ for inclusive muon-neutrino charged-current scattering on argon. This measurement utilizes data from 6.4 x 10 20 protons on target of exposure collected using the MicroBooNE liquid argon time projection chamber located along the Fermilab Booster Neutrino Beam with a mean neutrino energy of approximately 0.8 GeV. The mapping from reconstructed kinematics to truth quantities is validated within uncertainties by comparing the distribution of reconstructed hadronic energy in data to that of the model prediction in different muon scattering angle bins after applying a conditional constraint from the muon momentum distribution in data. The success of this validation provides confidence that the energy transfer in the MicroBooNE detector is well-modeled within simulation uncertainties, enabling a reliable unfolding to a triple-differential cross section defined at the nominal neutrino flux over muon momentum, muon scattering angle, and visible neutrino energy. This validation not only supports accurate cross-section extraction, but also establishes a critical foundation for tuning interaction models used in future neutrino oscillation measurements. The unfolded measurement covers an extensive phase space, providing a wealth of information useful for future liquid argon time projection chamber experiments measuring neutrino oscillations. Comparisons against a number of commonly used model predictions are included and their performance in different parts of the available phase-space is discussed.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

From Modular ADMS to Plug-and-Play Ops: Distribution Grid Operations with Platform-Level Orchestration to Enable Ambitious App Hosting

The core function of the distribution grid is to provide electricity to consumers affordably, reliably, and securely. In pursuing these core objectives, distribution utilities are accountable to customers, regulators, and in some cases, shareholders. Other third parties such as aggregators and microgrids can also have a stake in the smooth operation of the grid. Each of these stakeholders has economic, business, and/or governance objectives that inform their expectations of the distribution grid. This multi-objective, multi-stakeholder environment creates tension that must be reconciled to successfully design and operate the distribution grid. Innovative companies are competing to bring high-tech solutions to electric utilities and their customers that address each of these objectives. Many developers of advanced distribution management systems (ADMS) and distributed energy resource management systems (DERMS) have adopted a modular architecture that allows grid operators to select functions and features according to their individual system needs. A modular platform also allows the solution provider to develop and integrate specific new product modules; however, the need to pursue multiple objectives with a fixed set of controllable devices makes integration expensive whether it is done at the product development stage or the deployment stage. This cost creates a significant barrier to adoption and can lengthen the product to market time of new solutions. To fundamentally address the complexity of system integration for distribution grid operations, the U.S. Department of Energy Office of Electricity has funded the GridAPPS-D project at PNNL, which streamlines integration by contributing to standards development, defining system architecture, applying advanced mathematics, and developing open-source software to demonstrate the concept of an open data-integration platform for distribution operations. The open data-integration platform concept enables system operators and solution providers to deploy ambitious, best-of-breed applications (or apps) without continually reengineering for integration. Ambitious apps developed by different solution providers will inevitably attempt to achieve different control objectives with the same set of controllable devices. If the open platform itself can resolve these conflicts in a way that achieves the best available outcomes for all apps, doesn’t restrict the ambitious design of apps, and ensures safe and secure operations, apps will be able to plug-and-play with the platform at the same time as other ambitious apps. In this paper, we describe a framework called App Deconfliction that empowers a platform to assign setpoints to controllable devices based on the values preferred by different apps (and even external stakeholder entities like customers or aggregators). The App Deconfliction framework is compatible with several methods for determining setpoint values. We present two methods based on game theory that provide a subtle built-in incentive structure for developers to adapt their apps to the fact that they will be operating in a moderated multi-app environment and to favor device setpoints that have the most effect on their objectives over those that have the least effect. Our simulation-based demonstrations have shown that game-theory-based deconfliction can lead to a 7% improvement in control space utilization compared to design-based methods.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Continuous Baseline Microphysical Retrieval (MICROBASE) Value-Added Product Report

This technical report describes the Continuous Baseline Microphysical Retrieval (MICROBASE) Value-Added Product (VAP) produced operationally by the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) User Facility. MICROBASE provides a continuous estimate of cloud microphysical properties at ARM fixed observatories and ARM Mobile Facility (AMF) sites. It is designed to run operationally and provide data to the ARM Data Center for scientific distribution. This technical report presents an overview of the VAP as a resource for data users and ongoing records for major updates to these products.

54 ENVIRONMENTAL SCIENCES↗

Ocelot: An Interactive, Efficient Distributed Compression-As-a-Service Platform With Optimized Data Compression Techniques

Large volumes of data generated by scientific simulations, genome sequencing, and other applications need to be moved among clusters for data collection/analysis. Data compression techniques have effectively reduced data storage and transfer costs. However, users' requirements on interactively controlling both data quality and compression ratios are non-trivial to fulfill. Here, we propose a novel Compression-as-a-Service (CaaS) platform called Ocelot with four important contributions: (1) It offers real-time visualization, interactive compression, and transfer of scientific datasets. (2) It incorporates new strategies for compressing diverse types of datasets more effectively than traditional methods. (3) It provides an effective method for estimating the compression ratio and execution time of compression tasks. (4) Experiments on multiple real-world datasets on geographically distributed computers show that Ocelot can significantly improve data transfer efficiency with a performance gain of more than 10x in computing clusters with relatively slow networks.

compression as a service (CaaS)↗

Analysis and optimization of seismic monitoring networks with Bayesian optimal experimental design

SUMMARY Monitoring networks increasingly aim to assimilate data from a large number of diverse sensors covering many sensing modalities. Bayesian optimal experimental design (OED) seeks to identify data, sensor configurations or experiments which can optimally reduce uncertainty and hence increase the performance of a monitoring network. Information theory guides OED by formulating the choice of experiment or sensor placement as an optimization problem that maximizes the expected information gain (EIG) about quantities of interest given prior knowledge and models of expected observation data. Therefore, within the context of seismo-acoustic monitoring, we can use Bayesian OED to configure sensor networks by choosing sensor locations, types and fidelity in order to improve our ability to identify and locate seismic sources. In this work, we develop the framework necessary to use Bayesian OED to optimize a sensor network’s ability to locate seismic events from arrival time data of detected seismic phases at the regional-scale. This framework requires five elements: (i) A likelihood function that describes the distribution of detection and traveltime data from the sensor network, (ii) A prior distribution that describes a priori belief about seismic events, (iii) A Bayesian solver that uses a prior and likelihood to identify the posterior distribution of seismic events given the data, (iv) An algorithm to compute EIG about seismic events over a data set of hypothetical prior events, (v) An optimizer that finds a sensor network which maximizes EIG. Once we have developed this framework, we explore many relevant questions to monitoring such as: how to trade off sensor fidelity and earth model uncertainty; how sensor types, number and locations influence uncertainty; and how prior models and constraints influence sensor placement.

58 GEOSCIENCES↗

Bayesian inference of nuclear-matter density from proton scattering

Background: Proton elastic scattering at intermediate energy is widely employed as a tool for determining the matter radius of atomic nuclei. Here, the sensitivity of the approach relies on high-resolution measurements at small scattering angles and low-momentum transfer. Under these conditions, the Glauber multiple scattering theory accurately describes the proton-nucleus elastic cross section. Purpose: Investigate the sensitivity of the Glauber multiple scattering theory to uncertainties associated with input parameters such as the nuclear-matter density distribution and nucleon-nucleon data. Method: A joint Bayesian inference was performed using 12 angular distributions of elastic scattering at different energies on 58 Ni, 90 Zr, and 208 Pb targets. A Metropolis-Hastings algorithm was implemented to make an uncertainty quantification analysis for the input parameters used in the Glauber multiple scattering theory. Results: The experimental cross sections were fitted simultaneously using a joint Bayesian inference approach. Posterior probability density distributions of 42 input parameters were obtained from the analysis. A moderate correlation between the nuclear density parameters and the nucleon-nucleon cross sections was found. This correlation impacts the extraction of the nuclear-matter radius. Conclusions: The present analysis provided a consistent method for extracting the nuclear-matter density distribution of 58 Ni, 90 Zr, and 208 Pb from data across different incident energies. Due to the correlation of the nucleon-nucleon cross sections with the other input parameters, a constrained Bayesian inference using free nucleon-nucleon cross section data was performed. The nuclear-matter radii obtained from the analysis are in good agreement with multiple results reported in the literature.

190 ≤ A ≤ 219↗

Automated analysis of unlabeled PV data with Solar Data Tools software: Overview and feature updates

Distributed rooftop PV systems: ubiquitous, yet commonly have unlabeled data Difficult or impossible to form a performance index We developed Solar Data Tools (SDT), an open-source Python library for analyzing PV power (and irradiance) time-series data SDT enables analysis of unlabeled PV data—no model, no meteorological data, no performance index required Takes a statistical signal processing approach Data processing steps are largely pre-defined and automatic regardless of system type—from utility tracking systems to multi-pitch rooftop systems

Meyers-Im, Bennet E↗

Bootstrap-determined p values in lattice QCD

We present a general method to determine the probability that stochastic Monte Carlo data, in particular those generated in a lattice QCD calculation, would have been obtained were that data drawn from the distribution predicted by a given theoretical hypothesis. Such a probability, or p -value, is often used as an important heuristic measure of the validity of that hypothesis. The proposed method offers the benefit that it remains usable in cases where the standard Hotelling T 2 methods based on the conventional χ 2 statistic do not apply, such as for uncorrelated fits. Specifically, we analyze q 2 , defined as the correlated χ 2 statistic obtained using an arbitrary covariance matrix estimator, and show how to use the bootstrap as a data-driven method to determine the expected distribution of q 2 for a given hypothesis with minimal assumptions. This distribution can then be used to determine the p -value for a fit to the data. We also describe a bootstrap approach for quantifying the impact upon this p -value of estimating population parameters from a single ensemble of N samples. The overall method is accurate up to a 1 / N bias which we do not attempt to quantify. Published by the American Physical Society 2025

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Diaspora: Resilience-enabling services for science from HPC to edge

Scientific applications of interest to DOE must increasingly engage distributed resources (e.g., instruments, remote computers, data stores, edge devices) and deliver more stringent levels of service (e.g., uninterrupted processing of experiment data streams). In such systems, state is distributed and components can fail in many ways, often silently, making application resilience a major concern. Addressing the resilience needs of such applications requires methods for gaining knowledge of resources and applications and for translating that knowledge into action. We are working on addressing these needs in the context of multi-messenger astronomy, where detecting and responding to unusual transient events in multiple cosmic messengers (gravitational wave, electromagnetic, high- energy particles) from different instruments leads to a federated learning problem.

47 OTHER INSTRUMENTATION↗