Search NASASearch

SEARCH · Search NASA

Results for “data distributions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Deep Neural Networks are Adaptive to Function Regularity and Data Distribution in Approximation and Estimation

Deep learning has exhibited remarkable results across diverse areas. To understand its success, substantial research has been directed towards its theoretical foundations. Nev- ertheless, the majority of these studies examine how well deep neural networks can model functions with uniform regularities. In this paper, we explore a different angle: how deep neural networks can adapt to varying degrees of smoothness in functions and nonuni- form data distributions across different locations and scales. More precisely, we focus on a broad class of functions defined by nonlinear tree-based approximation methods. This class encompasses a range of function types, such as functions with uniform regularities and discontinuous functions. We develop nonparametric approximation and estimation theories for this class using deep ReLU networks. Our results show that deep neural networks are adaptive to the nonuniform smoothness of functions and nonuniform data distributions at different locations and scales. We apply our results to several function classes, and derive the corresponding approximation and generalization errors. The validity of our results is demonstrated through numerical experiments.

97 MATHEMATICS AND COMPUTING

Recommendations for developing, documenting, and distributing data products derived from NEON data

The National Ecological Observatory Network (NEON) provides over 180 distinct data products from 81 sites (47 terrestrial and 34 freshwater aquatic sites) within the United States and Puerto Rico. These data products include both field and remote sensing data collected using standardized protocols and sampling schema, with centralized quality assurance and quality control (QA/QC) provided by NEON staff. Such breadth of data creates opportunities for the research community to extend basic and applied research while also extending the impact and reach of NEON data through the creation of derived data products—higher level data products derived by the user community from NEON data. Derived data products are curated, documented, reproducibly-generated datasets created by applying various processing steps to one or more lower level data products—including interpolation, extrapolation, integration, statistical analysis, modeling, or transformations. Derived data products directly benefit the research community and increase the impact of NEON data by broadening the size and diversity of the user base, decreasing the time and effort needed for working with NEON data, providing primary research foci through the development via the derivation process, and helping users address multidisciplinary questions. Creating derived data products also promotes personal career advancement to those involved through publications, citations, and future grant proposals. However, the creation of derived data products is a nontrivial task. Here we provide an overview of the process of creating derived data products while outlining the advantages, challenges, and major considerations.

54 ENVIRONMENTAL SCIENCES

Applying Machine Learning and Bayesian Inference to Identify and Locate Moving Anthropogenic Sources Using Distributed Acoustic Sensing Data

Distributed acoustic sensing (DAS) systems, which use existing telecommunication fibers, offer high‐resolution capabilities ideal for recording anthropogenic sources. However, the complexity of urban environments and the large amount of data recorded by DAS require automated methods to efficiently detect and categorize anthropogenic sources. Here, we evaluate how well three machine learning models (k‐nearest neighbor [k‐NN], convolutional neural networks, and recurrent‐convolutional neural networks) can identify various anthropogenic sources recorded by DAS. Our findings reveal that both k‐NN and neural network methods perform well in high signal‐to‐noise ratio (SNR) settings. However, their accuracy decreases at SNRs <4. We also use Kalman filtering, a form of Bayesian inference, on backprojected locations of these sources to recover locations that generally fall within standard smartphone Global Positioning System errors. By combining machine learning and Kalman filter results, we calculate a multidimensional model of moving anthropogenic sources. These results demonstrate the potential of DAS data in urban seismology for accurately identifying and locating such sources. Depending on the research objectives, these sources can be further studied or filtered out to improve the quality of seismic data for earthquake studies. Such methods provide a valuable tool for urban seismology and seismic hazard analysis.

Luckie, Thomas William [Sandia National Laboratori

Materials data science using CRADLE: A distributed, data-centric approach

Abstract There is a paradigm shift towards data-centric AI, where model efficacy relies on quality, unified data. The common research analytics and data lifecycle environment (CRADLE™) is an infrastructure and framework that supports a data-centric paradigm and materials data science at scale through heterogeneous data management, elastic scaling, and accessible interfaces. We demonstrate CRADLE’s capabilities through five materials science studies: phase identification in X-ray diffraction, defect segmentation in X-ray computed tomography, polymer crystallization analysis in atomic force microscopy, feature extraction from additive manufacturing, and geospatial data fusion. CRADLE catalyzes scalable, reproducible insights to transform how data is captured, stored, and analyzed. Graphical abstract

97 MATHEMATICS AND COMPUTING

End-to-end microgrid protection using distributed data-driven methods

This paper introduces an end-to-end microgrid protection framework that offers real-time system monitoring, fault-related decision making, and circuit breaker control. This is achieved through the design of distributed data-driven techniques based on the support vector machine method, where each relay is responsible for distributed data collection, fault detection, fault localization, and fault isolation. Local communication is established among neighboring relays, fostering cooperative fault localization and isolation. This decentralized design not only reduces the computational and communication requirements but also enables the adaptability of each relay under varying operational dynamics. The proposed end-to-end protection framework was validated using MATLAB/Simulink simulations on a 100% renewable microgrid, achieving an accuracy of 93.1% with response time of 0.0523 s, in protecting against a range of fault scenarios that are characterized by various types, locations, impedances, load conditions, photovoltaic power levels, and microgrid operating modes.

24 POWER TRANSMISSION AND DISTRIBUTION

Performance Analysis of Data Processing in Distributed File Systems with Near Data Processing

In the era of big data, the escalating volume and velocity of data generation pose significant challenges in data processing. Traditional systems like Spark and Hadoop manage the increasing amount and velocity of data by improving data placement and processing speeds. However, they face inherent limitations due to the essential data movement required for processing. In this paper, we explore the Skyhook framework, a novel extension of the Ceph distributed system, which significantly reduces the need for data movement. We present an extensive case study using the Skyhook framework, applying it with the TPC-H and K-means clustering algorithms. More specifically, we leverage the TPC-H benchmark to distinguish between CPU-intensive and I/O-intensive tasks. We explore the integration of K-means clustering into SQL, coupled with a near-data processing system to offload the computational burden of the K-means clustering algorithm to storage nodes. We conduct a comprehensive performance evaluation of distributed data processing applications across three processing approaches: traditional layout (baseline), optimized layout, and near-data processing. Additionally, we introduce the use of the FIO tool to simulate real-world system workloads, enabling the measurement of performance metrics such as average latency and CPU utilization. Our research is a significant advance in understanding how to optimize data processing systems to meet the demands of the modern data landscape.

Hou, Shiyue

Bayesian Inference for the Seismic Moment Tensor Using Regional Waveforms and Teleseismic- P Polarities with a Data-Derived Distribution of Velocity Models and Source Locations

The largest source of uncertainty in any source inversion is the velocity model used in the transfer function that relates observed ground motion to the seismic moment tensor. However, standard inverse procedure often does not quantify uncertainty in the seismic moment tensor due to error in the Green’s functions from uncertain event location and Earth structure. Here, we incorporate this uncertainty into an estimation of the seismic moment tensor using a data-derived distribution of velocity models based on complementary geophysical data sets, including thickness constraints, velocity profiles, gravity data, surface-wave group velocities, and regional body-wave travel times. The data-derived distribution of velocity models is then used as a prior distribution of Green’s functions for use in Bayesian inference of an unknown seismic moment tensor using regional and teleseismic-P waveforms. The use of multiple data sets is important for gaining resolution to different components of the moment tensor. The combined likelihood is estimated using data-specific error models and the posterior of the seismic moment tensor is estimated and interpreted in terms of the most probable source type.

58 GEOSCIENCES

Solar Forecasting, Net Load Forecasting, and Data-Driven Distributed Solar Visibility Prizes (Final Technical Report)

The American-Made Solar Forecasting Prize, Net Load Forecasting Prize, and Data-Driven Distribution (3D) Solar Visibility Prize is a multimillion-dollar prize competition designed to energize U.S. solar innovation through a series of contests that accelerate the entrepreneurial process from years to months. The activities incentivized by these three prizes will support the governmentwide approach to increase American energy dominance by promoting innovation and early deployment of energy technologies, resulting in wider adoption, which is critical for secure, affordable, and reliable solar energy.

14 SOLAR ENERGY

Effectiveness of denoising diffusion probabilistic models for fast and high-fidelity whole-event simulation in high-energy heavy-ion experiments

Artificial intelligence (AI) generative models, such as generative adversarial networks (GANs), variational autoencoders, and normalizing flows, have been widely used and studied as efficient alternatives for traditional scientific simulations. However, they have several drawbacks, including training instability and inability to cover the entire data distribution, especially for regions where data are rare. This is particularly challenging for whole-event, full-detector simulations in high-energy heavy-ion experiments, such as sPHENIX at the Relativistic Heavy Ion Collider and Large Hadron Collider experiments, where thousands of particles are produced per event and interact with the detector. This work investigates the effectiveness of denoising diffusion probabilistic models (DDPMs) as an AI-based generative surrogate model for the sPHENIX experiment that includes the heavy-ion event generation and response of the entire calorimeter stack. DDPM performance in sPHENIX simulation data is compared with a popular rival, GANs. Results show that both DDPMs and GANs can reproduce the data distribution where the examples are abundant (low-to-medium calorimeter energies). Nonetheless, DDPMs significantly outperform GANs, especially in high-energy regions where data are rare. Additionally, DDPMs exhibit superior stability compared to GANs. The results are consistent between both central and peripheral centrality heavy-ion collision events. Moreover, DDPMs offer a substantial speedup of approximately a factor of 100 compared to the traditional Geant4 simulation method.

42 ENGINEERING

Considerations for Distributed Edge Data Centers and Use of Building Loads to Support Large Interconnections

The rapid expansion of artificial intelligence (AI) and machine learning is driving unprecedented electricity demand from data centers. It is predicted that by 2030, 90% of AI workloads will be inference-based, requiring interconnection of multiple low-latency edge data centers (<20 MW) sited closer to end users - often on already constrained distribution feeders. Although individually small, these loads can aggregate to large loads per feeder, straining infrastructure, creating multi-year interconnection delays, and driving up customer costs. This paper proposes a data center-focused grid-integration framework that combines feeder hosting capacity analysis with building energy efficiency, building load flexibility, and waste heat reuse to expand effective feeder and substation headroom. Such approaches can reduce interconnection delays, lower costs for ratepayers, and accelerate AI-ready infrastructure deployment.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Merged aerosol size distribution from SMPS and OPC for SAIL

This dataset contains merged aerosol number size distribution data for the Surface Atmosphere Integrated Field Laboratory (SAIL) campaign. The merged size distribution data were constructed by combining measurements from a scanning-mobility particle sizer (SMPS) and an optical particle counter (OPC), covering a size range of 0.01–35 µm. The merging methodology follows the approach described by Hand and Kreidenweis (2002) and Marinescu et al. (2019). All aerosol data from the ARM archive were corrected to standard temperature (273.15 K) and pressure (101.3 kPa).

merged size distribution

Outlook towards deployable continual learning for particle accelerators

Particle accelerators are high power complex machines. To ensure uninterrupted operation of these machines, thousands of pieces of equipment need to be synchronized, which requires addressing many challenges including design, optimization and control, anomaly detection and machine protection. With recent advancements, machine learning (ML) holds promise to assist in more advance prognostics, optimization, and control. While ML based solutions have been developed for several applications in particle accelerators, only few have reached deployment and even fewer to long term usage, due to particle accelerator data distribution drifts caused by changes in both measurable and non-measurable parameters. In this paper, we identify some of the key areas within particle accelerators where continual learning can allow maintenance of ML model performance with distribution drifts. Particularly, we first discuss existing applications of ML in particle accelerators, and their limitations due to distribution drift. Next, we review existing continual learning techniques and investigate their potential applications to address data distribution drifts in accelerators. By identifying the opportunities and challenges in applying continual learning, this paper seeks to open up the new field and inspire more research efforts towards deployable continual learning for particle accelerators.

43 PARTICLE ACCELERATORS

Influence of Alkyne Precursor Structure on Carbon Nanotube Chiral Distribution: Data-Dense Analysis Across Multiple Catalyst Types

Carbon nanotubes (CNTs) are a desirable material in the field of optoelectronics and semiconductors due to electronic properties (e.g., bandgap) that are dependent upon their chirality, defined by their diameter and lattice angle. Unfortunately, industrial-scale syntheses have yet to realize growth of a single desired chirality and instead rely on postsynthetic separation techniques to refine a chiral mixture, which increases process complexity and cost. Here, we studied the influence of precursor structure on chiral distribution, using a series of terminal alkyne precursors (acetylene, methylacetylene, vinylacetylene, 1-butyne, two enantiomers of 3-butyn-2-ol and a racemic mixture thereof) to grow CNTs across five transition-metal catalysts (Fe, FeMo, and three proportions of CoMo). Multiwavelength Raman spectroscopy on 5,145 spots (5 catalysts, 7 precursors, 3 lasers, and 49 distinct substrate locations on each) determined that acetylene grew the smallest diameter CNTs, while vinylacetylene produced fewer subnanometer CNTs. Though precursor structure did not dictate a uniform chiral shift, it was shown to broaden or narrow chiral distribution, while catalyst structure played a dominant role. In conclusion, this is consistent with metal-precursor binding occurring through unsaturated bonds in the hydrocarbons via the alkyne polymerization mechanism.

Carbon nanotubes

Visualization of Noisy and Less Noisy Computational Basis States in Quantum Computing

Quantum computing technology holds substantial promise as a reliable computational paradigm. However, current noisy intermediate scale quantum (NISQ) systems, are significantly impacted by noise originating from hardware inconsistencies. This noise causes errors and lowers output fidelity. So we must find which basis states cause errors. However, there are two main challenges in analyzing noise corresponding to basis states. First, the noise distribution data is high dimensional in nature, thereby making its analysis challenging. Second, although functional box plots have been used in the state of the art research to understand such a high dimensional data, they suffer from clutter and occlusion issues because of overplotting. In this study, we introduce an innovative visualization pipeline to address the aforementioned challenges to provide a clear depiction of noisy and less-noisy basis states. Specifically, our proposed visualization pipeline comprises three stages namely, low dimensional embedding, clustering, and violin plot visualization, to reduce visual clutter and effectively analyze high-dimensional noise distribution data. Our analysis uses quantum machine learning (QML) circuits as case study for drawing a distinction between noisy and less noisy basis states.

Senapati, Priyabrata [Kent State University]

EMPDF : inferring the Milky Way mass with data-driven distribution function in phase space

We introduce the emPDF (empirical distribution function), a novel dynamical modelling method that infers the gravitational potential from kinematic tracers with optimal statistical efficiency under the minimal assumption of steady state. emPDF determines the best-fitting potential by maximizing the similarity between instantaneous kinematics and the time-averaged phase-space distribution function (DF), which is empirically constructed from observation upon the theoretical foundation of oPDF (Han et al. 2016). This approach eliminates the need for presumed functional forms of DFs or orbit libraries required by conventional DF- or orbit-based methods. emPDF stands out for its flexibility, efficiency, and capability in handling observational effects, making it preferable to the popular Jeans equation or other minimal assumption methods, especially for the Milky Way (MW) outer halo where tracers often have limited sample size and poor data quality. We apply emPDF to infer the MW mass profile using Gaia DR3 data of satellite galaxies and globular clusters, obtaining enclosed masses of M (,r) = 26±8, 46±8, 90±13⁠, and 149±40 x 10 10 M ⊙ at r = 30, 50, 100⁠, and 200 kpc, respectively. These are consistent with the updated constraints from simulation-informed DF fitting (Li et al. 2020). While the simulation-informed DF offers superior precision owing to the additional information extracted from simulations, emPDF is independent of such supplementary knowledge and applicable to general tracer populations. emPDF is currently implemented for tracers with complete 6D kinematics within spherical potentials, but it can potentially be extended to address more general problems.

Astrophysics of Galaxies (astro-ph.GA)

Bridging the Gap on Data and Analysis for Distribution System Planning: Information That Utilities Can Provide Regulators, State Energy Offices and Other Stakeholders

Electric utilities conduct planning annually to ensure their distribution system meets technical standards, policies, and regulations; addresses forecasted grid conditions; satisfies customer needs; and advances utility priorities. The plan identifies grid deficiencies, analyzes potential solutions, and prioritizes capital investments and other expenditures. About 20 U.S. states and jurisdictions require regulated utilities to file some type of distribution system plan with the public utility commission for review. Requirements for sharing distribution system data and analyses vary widely, from few specific requirements to a detailed list of information that must be provided. While utilities conduct extensive analysis to develop distribution system plans, in most jurisdictions regulators and stakeholders do not know what data are available and how the utility uses the data in planning and investing. This report aims to bridge the gap by increasing understanding of the types of data and analyses utilities employ to develop distribution system plans and how the information affects their decision-making. The report describes information that states and stakeholders can ask for related to 11 data categories: -Forecasting loads and distributed energy resources (DERs) -Scenario analysis -Worst-performing circuits -Asset management strategy -Hosting capacity analysis -Value of DERs -Grid needs assessment -Cost-effectiveness framework for investments -Distribution system investment strategy and implementation -Geotargeted programs -Non-wires alternatives procurements.

24 POWER TRANSMISSION AND DISTRIBUTION

SIDDA: SInkhorn Dynamic Domain Adaptation for image classification with equivariant neural networks

Modern neural networks (NNs) often do not generalize well in the presence of a ‘covariate shift’; that is, in situations where the training and test data distributions differ, but the conditional distribution of classification labels given the data remains unchanged. In such cases, NN generalization can be reduced to a problem of learning more robust, domain-invariant features. Domain adaptation (DA) methods include a broad range of techniques aimed at achieving this; however, these methods have struggled with the need for extensive hyperparameter tuning, which then incurs significant computational costs. In this work, we introduce SInkhorn Dynamic Domain Adaptation (SIDDA), an out-of-the-box DA training algorithm built upon the Sinkhorn divergence, that can achieve effective domain alignment with minimal hyperparameter tuning and computational overhead. We demonstrate the efficacy of our method on multiple simulated and real datasets of varying complexity, including simple shapes, handwritten digits, real astronomical observations, and remote sensing data. These datasets exhibit covariate shifts due to noise, blurring, differences between telescopes, and variations in imaging wavelengths. SIDDA is compatible with a variety of NN architectures, and it works particularly well in improving classification accuracy and model calibration when paired with symmetry-aware equivariant NNs (ENNs). We find that SIDDA consistently enhances the generalization capabilities of NNs, achieving up to a ${\approx}40\%$ improvement in classification accuracy on unlabeled target data, while also providing a more modest performance gain of $\lesssim 1\%$ on labeled source data. We also study the efficacy of DA on ENNs with respect to the varying group orders of the dihedral group DN, and find that the model performance improves as the degree of equivariance increases. Finally, if SIDDA achieves proper domain alignment, it also enhances model calibration on both source and target data, with the most significant gains in the unlabeled target domain—achieving over an order of magnitude improvement in the expected calibration error and Brier score. SIDDA’s versatility across various NN models and datasets, combined with its automated approach to domain alignment, has the potential to significantly advance multi-dataset studies by enabling the development of highly generalizable models.

79 ASTRONOMY AND ASTROPHYSICS