Search NASA⌕ Search

SEARCH · Search NASA

Results for “data set”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Uncertainty-Informed Volume Visualization using Implicit Neural Representation

The increasing adoption of Deep Neural Networks (DNNs) has led to their application in many challenging scientific visualization tasks. While advanced DNNs offer impressive generalization capabilities, understanding factors such as model prediction quality, robustness, and uncertainty is crucial. These insights can enable domain scientists to make informed decisions about their data. However, DNNs inherently lack ability to estimate prediction uncertainty, necessitating new research to construct robust uncertainty-aware visualization techniques tailored for various visualization tasks. In this work, we propose uncertainty-aware implicit neural representations to model scalar field data sets effectively and comprehensively study the efficacy and benefits of estimated uncertainty information for volume visualization tasks. We evaluate the effectiveness of two principled deep uncertainty estimation techniques: (1) Deep Ensemble and (2) Monte Carlo Dropout (MC-Dropout). These techniques enable uncertainty-informed volume visualization in scalar field data sets. Our extensive exploration across multiple data sets demonstrates that uncertainty-aware models produce informative volume visualization results. Moreover, integrating prediction uncertainty enhances the trustworthiness of our DNN model, making it suitable for robustly analyzing and visualizing real-world scientific volumetric data sets.

Saklani, Shanu↗

Streaming Compression of Scientific Data via Weak-SINDy

Here, in this paper, a streaming weak-SINDy algorithm is developed specifically for compressing streaming scientific data. The production of scientific data, either via simulation or experiments, is undergoing a stage of exponential growth, which makes data compression important and often necessary for storing and utilizing large scientific data sets. As opposed to classical “offline” compression algorithms that perform compression on a readily available data set, streaming compression algorithms compress data “online” while the data generated from simulation or experiments is still flowing through the system. This feature makes streaming compression algorithms well suited for scientific data compression, where storing the full data set offline is often infeasible. This work proposes a new streaming compression algorithm, streaming weak-SINDy, which takes advantage of the underlying data characteristics during compression. The streaming weak-SINDy algorithm constructs feature matrices and target vectors in the online stage via a streaming integration method in a memory efficient manner. The feature matrices and target vectors are then used in the offline stage to build a model through a regression process that aims to recover equations that govern the evolution of the data. For compressing high-dimensional streaming data, we adopt a streaming proper orthogonal decomposition (POD) process to reduce the data dimension and then use the streaming weak-SINDy algorithm to compress the temporal data of the POD expansion. We propose modifications to the streaming weak-SINDy algorithm to accommodate the dynamically updated POD basis. By combining the built model from the streaming weak-SINDy algorithm and a small amount of data samples, the full data flow could be reconstructed accurately at a low memory cost, as shown in the numerical tests.

97 MATHEMATICS AND COMPUTING↗

All-sky Search for Transient Astrophysical Neutrino Emission with 10 Years of IceCube Cascade Events

Abstract Neutrino flares in the sky are searched for in data collected by IceCube between 2011 and 2021 May. This data set contains cascade-like events originating from charged-current electron neutrino and tau neutrino interactions and all-flavor neutral-current interactions. IceCube’s previous all-sky searches for neutrino flares used data sets consisting of track-like events originating from charged-current muon neutrino interactions. The cascade data set is statistically independent of the track data sets, and while inferior in angular resolution, the low-background nature makes it competitive and complementary to previous searches. No statistically significant flare of neutrino emission was observed in an all-sky scan. Upper limits are calculated on neutrino flares of varying duration from 1 hr to 100 days. Furthermore, constraints on the contribution of these flares to the diffuse astrophysical neutrino flux are presented, showing that multiple unresolved transient sources may contribute to the diffuse astrophysical neutrino flux.

79 ASTRONOMY AND ASTROPHYSICS↗

Impact of Domain Knowledge on the Property Prediction of Specialized Machine Learning Models

Developing transferable machine learning models is trending in data-driven materials research. However, how to apply such models to a specific research domain remains unclear. Here, in this work, we choose high-entropy materials as a platform with a specialized data set containing 145,323 DFT-relaxed materials. This data set is used to explore the role of domain-specific knowledge in training effective models. Our tests with three representative graph neural network architectures indicate the model complexity has much smaller influence on performance than the data itself. Specifically, the consideration of low-energy atomic ordering, structures with diverse elemental coverage, and high-order interactions significantly influences the model performance. We also find that domain knowledge-driven sampling can greatly enhance unsupervised learning techniques. This research highlights that developing specialized data sets is more beneficial than further complicating deep learning architectures. Additionally, physics-inspired sampling algorithms are crucially needed for better machine learning models for a specific materials research domain.

36 MATERIALS SCIENCE↗

Neural Network‐Based Methods for Ocean Surface Wave Measurement Using Submarine Distributed Acoustic Sensing (DAS)

Two new data-driven models for estimating ocean surface waves from distributed acoustic sensing (DAS) submarine cable strain rate are developed using supervised machine learning on a 10-day data set collected offshore of Oliktok Point, Alaska. The new models were trained on target data from seafloor pressure moorings at three sites spaced evenly along 27.1 km of cable and were benchmarked against an empirical transfer function method previously used to estimate waves from DAS. A model which uses convolutional neural networks to transform 2-km frequency-wavenumber strain spectra to seafloor pressure spectra outperforms the benchmark in wave height prediction (RMSE of 0.15 vs. 0.41 m) and period prediction (0.29 vs. 0.37 s) when evaluated on a held-out test data set. When applied to a DAS data set collected on the same cable 2 years prior, the CNN-based model maintained similar significant wave height performance (RMSE = 0.23 m) relative to available satellite altimetry data. A two-hidden-layer, fully connected neural network which transforms 1-D strain spectra to seafloor pressure spectra also outperforms the benchmark in wave height prediction (RMSE of 0.19 vs. 0.41 m), but does not generalize as well to the prior data. Regression-based machine learning is useful for estimating waves from DAS data when the pressure-strain relationship varies temporally and spatially across different wave conditions. Models can be applied to DAS data to measure waves with higher spatial resolution and longer temporal coverage than traditional methods, which often measure waves only at a single point.

Davis, Jacob R. [Univ. of Washington, Seattle, WA ↗

A machine-learning-driven data labeling pipeline for scientific analysis in MLExchange

This study introduces a novel labeling pipeline to accelerate the labeling process of scientific data sets by using artificial intelligence (AI)-guided tagging techniques. This pipeline includes a set of interconnected web-based graphical user interfaces (GUIs), where Data Clinic and MLCoach enable the preparation of machine learning (ML) models for data reduction and classification, respectively, while Label Maker is used for label assignment. Throughout this pipeline, data can be accessed through a direct connection to a file system or through Tiled for access through Hypertext Transfer Protocol (HTTP). Our experimental results present three use cases where this labeling pipeline has been instrumental for the study of large X-ray scattering data sets in the area of pattern recognition, the remote analysis of resonant soft X-ray scattering data and the fine-tuning process of foundation models. These use cases highlight the labeling capabilities of this pipeline, including the ability to label large data sets in a short period of time, to perform remote data analysis while minimizing data movement and to enhance the fine-tuning process of complex ML models with human involvement.

Chavez, Tanny (ORCID:0000000193172896)↗

Meteorological Services Annual Data Report for 2024

This document presents the meteorological data collected at Brookhaven National Laboratory (BNL) by Meteorological Services (Met Services) for the calendar year 2024. The purpose is to publicize the data sets available to emergency personnel, researchers and facility operations. Met services has been collecting data at BNL since 1949. Data from 1994 to the present is available in digital format. Data is presented in monthly plots of one-minute data. This allows the reader the ability to peruse the data for trends or anomalies that may be of interest to them. Full data sets are available to BNL personnel and to a limited degree outside researchers. The full data sets allow plotting the data on expanded time scales to obtain greater details (e.g., daily solar variability, inversions, etc.).

54 ENVIRONMENTAL SCIENCES↗

ComStock Measure Documentation: Variable-Speed Pumps

Building on the 3-year End-Use Load Profiles project to calibrate and validate the U.S. Department of Energy's ResStock and ComStock models, this work produces national data sets that enable cities, states, utilities, and other stakeholders to answer a broad range of questions regarding their commercial building stock. ComStock is a highly granular, bottom-up model that uses various data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the commercial building stock across the United States. The "baseline" model intends to represent the U.S. commercial building stock as it existed in 2018. The methodology of the baseline model is discussed in the ComStock Reference Documentation. The goal of this work is to develop energy efficiency and demand flexibility measures that cover market-ready technologies and study their mass adoption impact on the baseline building stock. "Measures" refers to various "what-if" scenarios that can be applied to buildings. The results for the baseline and measure scenario simulations are published in public data sets that provide insights into building stock characteristics, operational behaviors, utility bill impacts, and annual and sub-hourly energy usage by fuel type and end use. This report describes the modeling methodology for a single ComStock measure scenario - variable speed pumps - and briefly introduces key results. The full public data set can be accessed on the ComStock data lake or via the Data Viewer at comstock.nlr.gov. The public data set enables users to create custom aggregations of results for their use case (e.g., filter to a specific county or building type).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

ComStock Measure Documentation: High-Efficiency Rooftop Unit

Building on the 3-year End-Use Load Profiles project to calibrate and validate the U.S. Department of Energy's ResStock and ComStock models, this work produces national data sets that enable cities, states, utilities, and other stakeholders to answer a broad range of questions regarding their commercial building stock. ComStock is a highly granular, bottom-up model that uses various data sources, statistical sampling methods, and advanced building energy simulations to estimate the annual subhourly energy consumption of the commercial building stock across the United States. The "baseline" model intends to represent the U.S. commercial building stock as it existed in 2018. The methodology of the baseline model is discussed in the ComStock Reference Documentation. The goal of this work is to develop energy efficiency and demand flexibility measures that cover market-ready technologies and study their mass adoption impact on the baseline building stock. "Measures" refers to various "what-if" scenarios that can be applied to buildings. The results for the baseline and measure scenario simulations are published in public data sets that provide insights into building stock characteristics, operational behaviors, utility bill impacts, and annual and sub-hourly energy usage by fuel type and end use. This report describes the modeling methodology for a single ComStock measure scenario - high-efficiency rooftop unit (RTU) - and briefly introduces key results. The full public data set can be accessed on the Comstock data lake or via the Data Viewer at comstock.nlr.gov. The public data set enables users to create custom aggregations of results for their use case (e.g., filter to a specific county or building type).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Snow Distribution Patterns Revisited: A Physics-Based and Machine Learning Hybrid Approach to Snow Distribution Mapping in the Sub-Arctic

Snowpack distribution in Arctic and alpine landscapes often occurs in repeating, year-to-year patterns due to local topographic, weather, and vegetation characteristics. Previous studies have suggested that with years of observational data, these snow distribution patterns can be statistically integrated into a snow process modeling workflow. Recent advances in snow hydrology and machine learning (ML) have increased our ability to predict snowpack distribution using in-situ observations, remote sensing data sets, and simple landscape characteristics that can be easily obtained for most environments. Here, we propose a hybrid approach to couple a ML snow distribution pattern (MLSDP) map with a physics-based, snow process model. We trained a random forest ML algorithm on tens of thousands of snow survey observations from a subarctic study area on the Seward Peninsula, Alaska, collected during peak snow water equivalent (SWE). We validated hybrid model outputs using in-situ snow depth and SWE observations, as well as a light detection and ranging data set and a distributed temperature profiling sensor data set. When the hybrid results were compared with the physics-based method, the hybrid method more accurately depicted the spatial patterns of the snowpack, areas of drifting snow, and years when no in-situ observations were used in the random forest ML training data set. The hybrid method also showed improvements in root mean squared error at 61% of locations where time-series estimations of snow depth were observed. These results can be applied to any physics-based model to improve the snow distribution patterning to reflect observed conditions in high latitude and high elevation cold region environments.

54 ENVIRONMENTAL SCIENCES↗

Rapid detection of rare events from in situ X-ray diffraction data using machine learning

High-energy X-ray diffraction methods can non-destructively map the 3D microstructure and associated attributes of metallic polycrystalline engineering materials in their bulk form. These methods are often combined with external stimuli such as thermo-mechanical loading to take snapshots of the evolving microstructure and attributes over time. However, the extreme data volumes and the high costs of traditional data acquisition and reduction approaches pose a barrier to quickly extracting actionable insights and improving the temporal resolution of these snapshots. This article presents a fully automated technique capable of rapidly detecting the onset of plasticity in high-energy X-ray microscopy data. The technique is computationally faster by at least 50 times than the traditional approaches and works for data sets that are up to nine times sparser than a full data set. This new technique leverages self-supervised image representation learning and clustering to transform massive data sets into compact, semantic-rich representations of visually salient characteristics ( e.g. peak shapes). These characteristics can rapidly indicate anomalous events, such as changes in diffraction peak shapes. It is anticipated that this technique will provide just-in-time actionable information to drive smarter experiments that effectively deploy multi-modal X-ray diffraction methods spanning many decades of length scales.

Zheng, Weijian↗

Compactly‐Supported Nonstationary Kernels for Computing Exact Gaussian Processes on Big Data

The Gaussian process (GP) is a widely used method for analyzing large-scale data sets, including spatio-temporal measurements of nonlinear processes that are now commonplace in the environmental sciences. Traditional implementations of GPs involve stationary kernels (also termed covariance functions) that limit their flexibility, and exact methods for inference that prevent application to data sets with more than about 10,000 points. Modern approaches to address stationarity assumptions generally fail to accommodate large data sets, while all attempts to address scalability focus on approximating the Gaussian likelihood, which can involve subjectivity and lead to inaccuracies. In this work, we explicitly derive an alternative kernel that can discover and encode both sparsity and nonstationarity. We embed the kernel within a fully Bayesian GP model and leverage high-performance computing resources to enable the analysis of massive data sets. We demonstrate the favorable performance of our novel kernel relative to existing exact and approximate GP methods across a variety of synthetic data examples. Furthermore, we conduct space–time prediction based on more than 1 million measurements of daily maximum temperature and verify that our results outperform state-of-the-art methods in the Earth sciences. More broadly, having access to exact GPs that use ultra-scalable, sparsity-discovering, nonstationary kernels allows GP methods to truly compete with a wide variety of machine learning methods.

Gaussian processes↗

Refining Planetary Boundary Layer Height Retrievals From Micropulse‐Lidar at Multiple ARM Sites Around the World

Abstract Knowledge of the planetary boundary layer height (PBLH) is crucial for various applications in atmospheric and environmental sciences. Lidar measurements are frequently used to monitor the evolution of the PBLH, providing more frequent observations than traditional radiosonde‐based methods. However, lidar‐derived PBLH estimates have substantial uncertainties, contingent upon the retrieval algorithm used. In addressing this, we applied the Different Thermo‐Dynamic Stabilities (DTDS) algorithm to establish a PBLH data set at five separate Department of Energy's Atmospheric Radiation Measurement sites across the globe. Both the PBLH methodology and the products are subject to rigorous assessments in terms of their uncertainties and constraints, juxtaposing them with other products. The DTDS‐derived product consistently aligns with radiosonde PBLH estimates, with correlation coefficients exceeding 0.77 across all sites. This study delves into a detailed examination of the strengths and limitations of PBLH data sets with respect to both radiosonde‐derived and other lidar‐based estimates of the PBLH by exploring their respective errors and uncertainties. It is found that varying techniques and definitions can lead to diverse PBLH retrievals due to the inherent intricacy and variability of the boundary layer. Our DTDS‐derived PBLH data set outperforms existing products derived from ceilometer data, offering a more precise representation of the PBLH. This extensive data set paves the way for advanced studies and an improved understanding of boundary‐layer dynamics, with valuable applications in weather forecasting, climate modeling, and environmental studies.

54 ENVIRONMENTAL SCIENCES↗

Day-Ahead Probabilistic Forecasting of Net-Load and Demand Response Potentials with High Penetration of Behind-the-Meter Solar-plus-Storage

The goal of this project is to develop advanced methods for day-ahead net-load forecasting, by leveraging the state-of-the-art machine learning techniques. The developed models produce both point and probabilistic forecasts for a variety of use cases, and are versatile to work with different types of data sets. The innovation lies in the novel design of the architectures, leveraging the most recent advances in machine learning that have not been explored in power systems, accompanied by techniques in the broader artificial intelligence fields such as fuzzy systems. This project has achieved the following accomplishments: (1) preprocessing of over 10 data sets covering varying geographical regions, time horizons, and system levels, which form a robust foundation for training and evaluating forecasting models across a wide range of realistic grid scenarios; (2) development of an interactive web app that enables exploratory analysis of load and generation data, and supports better understanding of data trends, anomalies, and correlations, facilitating model development and stakeholder engagement; (3) implementation of over 10 benchmark models for point and probabilistic forecasting, which include a mix of conventional machine learning methods and state-of-the-art deep learning approaches, providing a comprehensive baseline for performance comparison and validation of the proposed models; (4) development of a fuzzy system based gradient boosting model, tailored for small (less than 3 years) data sets, which achieves a mean absolute percentage error (MAPE) of 4% for point forecasting and a 20% improvement in average pinball loss for probabilistic forecasting; (5) development of a Transformer (a state-of-the-art deep learning architecture) based neural network model, tailored for large (3 years or more) data sets, which achieves a MAPE of 2% for point forecasting and a 20% improvement in average pinball loss for probabilistic forecasting; (6) development of a methodology for quantifying DR potential, and extensions of the previous models for multi-target forecasting of net load and DR potential, which achieve a MAPE of 10% for DR potential.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Physical discovery in representation learning via conditioning on prior knowledge

Recent advances in electron, scanning probe, optical, and chemical imaging and spectroscopy yield bespoke data sets containing the information of structure and functionality of complex systems. In many cases, the resulting data sets are underpinned by low-dimensional simple representations encoding the factors of variability within the data. The representation learning methods seek to discover these factors of variability, ideally further connecting them with relevant physical mechanisms. However, generally, the task of identifying the latent variables corresponding to actual physical mechanisms is extremely complex. Here, we present an empirical study of an approach based on conditioning the data on the known (continuous) physical parameters and systematically compare it with the previously introduced approach based on the invariant variational autoencoders. The conditional variational autoencoder (cVAE) approach does not rely on the existence of the invariant transforms and hence allows for much greater flexibility and applicability. Interestingly, cVAE allows for limited extrapolation outside of the original domain of the conditional variable. However, this extrapolation is limited compared to the cases when true physical mechanisms are known, and the physical factor of variability can be disentangled in full. We further show that introducing the known conditioning results in the simplification of the latent distribution if the conditioning vector is correlated with the factor of variability in the data, thus allowing us to separate relevant physical factors. We initially demonstrate this approach using 1D and 2D examples on a synthetic data set and then extend it to the analysis of experimental data on ferroelectric domain dynamics visualized via piezoresponse force microscopy.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

2025 Annual INMM Graph and Tables for High Purity Germanium Detector Normalization Presentation

The data set includes gamma spectroscopy peak data for measurements taken with two different high purity germanium detectors using a mixed nuclide source and a U-235 fuel rod. There are a total of 5 specific energy peaks that were analyzed for the mixed nuclide source stemming from Am-241, Cd-109, Cs-137, and Co-60. There are a total of 3 specific energy peaks that were analyzed for the U-235 fuel rod. The data set includes the calculations and results from using a linear correction factor, absolute efficiency curve, and relative efficiency curve to compare the net peaks counts from two different detectors.

Drumm, Natalie Daphne [Sandia National Laboratorie↗

Comparison of Global Aboveground Biomass Estimates From Satellite Observations and Dynamic Global Vegetation Models

The global forest carbon stocks represent the amount of carbon stored in woody vegetation and are important for quantifying the ability of the global forests to sequester atmospheric CO 2 and to provide ecosystem services (e.g., timber) under climate change. The forest ecosystem carbon pool estimates are highly variable and poorly quantified in areas lacking forest inventory estimates. Here, we compare and analyze aboveground biomass (AGB) estimates from five satellite-based global data sets and nine dynamic global vegetation models (DVGMs). We find that across the data sets, mean AGB exhibits the largest variability around the tropical area. In addition, AGB shows a similar latitudinal trend but large variability among the data sets. Satellite-based AGB estimates are lower than those simulated by DVGMs. The divergence among the satellite-based AGB estimates can be driven by the methodology, input satellite products, and the forested areas used to estimate AGB. The modeled NPP, autotrophic respiration, and carbon allocation mostly drive the variability of AGB simulated by DGVMs. The future availability of a high-quality global forest area map is anticipated to improve AGB estimate accuracy and to reduce the discrepancies among different satellite- and model-based AGB estimates. Furthermore, we suggest the carbon-modeling community reexamine the methodology used to estimate AGB and forested areas for a more robust global forest carbon stock estimation.

54 ENVIRONMENTAL SCIENCES↗