Search NASA⌕ Search

SEARCH · Search NASA

Results for “Dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

A co-registered in-situ and ex-situ dataset from wire arc additive manufacturing process

Recent progress in sensing techniques and data analytics tools have significantly accelerated the development of Wire Arc Additive Manufacturing (WAAM) systems. This data-centric approach emphasizes leveraging sensor data available throughout the production process to optimize performance. Integration of extensive data analysis provides opportunities for improving precision, reducing waste, and enhancing the quality of produced parts. This method relies on AI/ML models and optimization techniques, which are developed using the data collected from various sources, including in-situ sensors, ex-situ imaging, and manufacturing process parameters. The quality and diversity of this data, along with the alignment between different data streams (achieved through spatiotemporal registration) are critical for the successful development of AI/ML and optimization models. In this work, we present a spatiotemporally registered dataset generated during the WAAM process of deposition of a rectangular block. The dataset includes a comprehensive description of the deposition process, process parameters, welding characteristics and acoustic data collected in-situ, and X-Ray Computed Tomography data of the build.

42 ENGINEERING↗

An open retail boundary dataset for South Korea using open data and computer vision technique

Although delineating retail boundaries is important to explore and comprehend the dynamics of the retail sector, it is hard to find studies specifically addressing it in the South Korean context. This study fills this gap by proposing new retail boundaries across South Korea. To achieve this goal, we employed a variety of retailers and building datasets and proposed a unique computer vision-based framework with a deep ensemble voting technique. As a result, we delineated 6,636 distinct retail boundaries that were validated against existing reference retail boundaries. These newly delineated retail boundaries provide valuable insights for researchers, governments, and other relevant stakeholders by enhancing their understanding of retail geography. This dataset can be used as a foundational resource for analyses on topics such as pandemic recovery, retail gentrification, and the resilience of retail spaces in response to e-commerce growth, ultimately contributing to more robust retail sector research in South Korea.

97 MATHEMATICS AND COMPUTING↗

Quantum mechanical dataset of 836k neutral closed-shell molecules with up to 5 heavy atoms from C, N, O, F, Si, P, S, Cl, Br

Abstract We introduce the Vector-QM24 (VQM24) dataset comprehensively covering all possible neutral closed-shell small organic and inorganic molecules with up to five heavy (p-block) atoms: C, N, O, F, Si, P, S, Cl, Br. All valid stoichiometries, Lewis-rule-consistent graphs, and stable conformers (identified via GFN2-xTB) were enumerated combinatorially, yielding 577k conformational isomers spanning 258k constitutional isomers and 5,599 unique stoichiometries. DFT (ωB97X-D3/cc-pVDZ) optimizations were performed for all, and diffusion quantum Monte Carlo (DMC@PBE0(ccECP/cc-pVQZ)) energies are provided for 10,793 lowest-energy conformers with up to 4 heavy atoms. VQM24 includes structures, vibrational modes, rotational constants, thermodynamic properties (Gibbs free energies, enthalpies, ZPVEs, entropies, heat capacities), and electronic properties such as atomization, electron interaction, exchange-correlation, dispersion energies, multipole moments (dipole to hexadecapole), alchemical potentials, Mulliken charges, and wavefunctions. Machine learning models of atomization energies on this dataset reveal significantly higher complexity than QM9, with none achieving chemical accuracy. VQM24 offers a rigorous, high-fidelity benchmark for evaluating quantum machine learning models.

Science & Technology - Other Topics↗

Experimental validation of a collision-radiation dataset for molecular hydrogen in plasmas

Quantitative spectroscopy of molecular hydrogen has generated substantial demand, leading to the accumulation of diverse elementary process data encompassing radiative transitions, electron-impact transitions, predissociations, and quenching. However, their rates currently available are still sparse, and there are inconsistencies among those proposed by different authors. In this study, we demonstrate an experimental validation of such a molecular dataset by composing a collisional-radiative model (CRM) for molecular hydrogen and comparing experimentally obtained vibronic populations across multiple levels. From the population kinetics of molecular hydrogen, the importance of each elementary process in various parameter space is studied. In low-density plasmas (electron density ne≲1017 m−3) the excitation rates from the ground states and radiative decay rates, both of which have been reported previously, determine the excited state population. The inconsistency in the excitation rates affects the population distribution the most significantly in this parameter space. However, in higher density plasmas (ne≳1018 m−3), the excitation rates from excited states become important, which have never been reported in the literature, and may need to be approximated in some way. In order to validate these molecular datasets and approximated rates, we carried out experimental observations for two different hydrogen plasmas; a low-density radio frequency heated plasma (ne≈1016 m−3) and the Large Helical Device (LHD) divertor plasma (ne≳1018 m−3). The visible emission lines from EF1Σg+, HH¯1Σg+, D1Πu±, GK1Σg+, I1Πg±, J1Δg±, h3Σg+, e3Σu+, d3Πu±,g3Σg+, i3Πg±, and j3Δg± states were observed simultaneously and their population distributions were obtained from their intensities. We compared the observed population distributions with the CRM prediction, in particular the CRM with the rates compiled by Janev et al., Miles et al., and those calculated with the molecular convergent close-coupling (MCCC) method. The MCCC prediction gives the best agreement with the experiment, particularly for the emission from the low-density plasma. However, the population distribution in the LHD divertor shows a worse agreement with the CRM than those from low-density plasma, indicating the necessity of the precise excitation rates from excited states. We also found that the rates for the electron attachment is inconsistent with experimental results. This requires further investigation.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Two datasets are better than one: method of double moments for 3D reconstruction in cryo-EM

Cryo-electron microscopy is a powerful imaging technique for reconstructing three-dimensional molecular structures from noisy tomographic projection images of randomly oriented particles. We introduce a new data fusion framework, termed the method of double moments, which reconstructs molecular structures from two instances of the second-order moment of projection images obtained under distinct orientation distributions: one uniform, the other non-uniform and unknown. We prove that these moments generically uniquely determine the underlying structure, up to a global rotation and reflection, and we develop a convex-relaxation-based algorithm that achieves accurate recovery using only second-order statistics. Our results demonstrate the advantage of collecting and modeling multiple datasets under different experimental conditions, illustrating that leveraging dataset diversity can substantially enhance reconstruction quality in computational imaging tasks.

Kam’s method↗

Validation of the DESI DR2 Ly⁢ 𝛼 BAO analysis using synthetic datasets

The second data release (DR2) of the Dark Energy Spectroscopic Instrument (DESI), containing data from the first three years of observations, doubles the number of Lyman-α (Ly α) forest spectra in DR1 and it provides the largest dataset of its kind. To ensure a robust validation of the baryonic acoustic oscillation (BAO) analysis using Ly α forests, we have made significant updates compared to DR1 to both the mocks and the analysis framework used in the validation. In particular, we present CoLoRe-QL, a new set of Lyα mocks that use a quasilinear input power spectrum to incorporate the nonlinear broadening of the BAO peak. Here, we have also increased the number of realizations used in the validation to 400, compared to the 150 realizations used in DR1. Finally, we present a detailed study of the impact of quasar redshift errors on the BAO measurement, and we compare different strategies to mask damped Lyman-α absorbers in our spectra. The BAO measurement from the Ly α dataset of DESI DR2 is presented in a companion publication.

Casas, L. [Institut de Física d’Altes Energies (IF↗

SPT clusters with DES and HST weak lensing. I. Cluster lensing and Bayesian population modeling of multiwavelength cluster datasets

We present a Bayesian population modeling method to analyze the abundance of galaxy clusters identified by the South Pole Telescope (SPT) with a simultaneous mass calibration using weak gravitational lensing data from the Dark Energy Survey (DES) and the Hubble Space Telescope (HST). We discuss and validate the modeling choices with a particular focus on a robust, weak-lensing-based mass calibration using DES data. For the DES Year 3 data, we report a systematic uncertainty in weak-lensing mass calibration that increases from 1% at z = 0.25 to 10% at z = 0.95 , to which we add 2% in quadrature to account for uncertainties in the impact of baryonic effects. We implement an analysis pipeline that joins the cluster abundance likelihood with a multiobservable likelihood for the Sunyaev-Zel’dovich effect, optical richness, and weak-lensing measurements for each individual cluster. We validate that our analysis pipeline can recover unbiased cosmological constraints by analyzing mocks that closely resemble the cluster sample extracted from the SPT-SZ, SPTpol ECS, and SPTpol 500d surveys and the DES Year 3 and HST-39 weak-lensing datasets. This work represents a crucial prerequisite for the subsequent cosmological analysis of the real dataset.

79 ASTRONOMY AND ASTROPHYSICS↗

Descriptor: Infrastructure Perception and Control: Multi-Sensor Object Tracking Dataset (IPC-MSOT)

Traffic intersections are crucial and challenging nodes in transportation networks where multiple lanes of vehicles and pedestrians converge. Traffic accidents often occur at traffic intersections, including a large proportion of traffic fatalities and about one-half of all traffic injuries in the United States. Object detection data were collected in 2024 across three intersections in Colorado Springs, CO, USA, over the course of multiple days and various times to induce a heterogeneous mix of traffic conditions and behaviors. The purpose of the data collection exercises was to learn various attributes about infrastructure sensors and to build a repository of high-resolution, object-level data that can be used for research and development (e.g., to develop multisensor data fusion algorithms). The Infrastructure Perception and Control:Multi-Sensor Object tracking (IPC-MSOT) dataset was collected as part of the U.S. Department of Transportation's Strengthening Mobility and Revolutionizing Transportation (SMART) project, where the city of Colorado Springs, Colorado, and the National Renewable Energy Laboratory collaborated to collect object-level trajectory data from road users using multiple types of infrastructure sensors deployed at different intersections. This dataset allows for testing of late-stage sensor fusion algorithms and their ability to ingest multimodal sensor data, and it can be utilized by traffic engineers to design and evaluate trajectory-based signal control strategies.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

From Data to Insights: A Covariate Analysis of the IARPA BRIAR Dataset for Multimodal Biometric Recognition Algorithms at Altitude and Range

This paper examines covariate effects on fused whole body biometrics performance in the IARPA BRIAR dataset, specifically focusing on UAV platforms, elevated positions, and distances up to 1000 meters. The dataset includes outdoor videos compared with indoor images and controlled gait recordings. Normalized raw fusion scores relate directly to predicted false accept rates (FAR), offering an intuitive means for interpreting model results. A linear model is developed to predict biometric algorithm scores, analyzing their performance to identify the most influential covariates on accuracy at altitude and range. Weather factors like temperature, wind speed, solar loading, and turbulence are also investigated in this analysis. The study found that resolution and camera distance best predicted accuracy and findings can guide future research and development efforts in long-range/elevated/UAV biometrics and support the creation of more reliable and robust systems for national security and other critical domains.

Bolme, David↗

Impact Study of Thunderstorms on the US Power Grid Using Publicly Available Datasets

This work analyzes the impact of thunderstorms on the US power grid based on publicly available data. Since thunderstorms can bring lightning, heavy precipitation, and wind storms, analyzing their impact on the power system provides a combined correlation of lightning strikes, floods, and wind storms on power outages. This paper leverages publicly available thunderstorm datasets from the National Weather Service (NWS) and power outage datasets from Oak Ridge National Laboratory’s Environment for Analysis of Geo-Located Energy Information (EAGLE-I) to study the correlation between thunderstorms and power outages. This work is analyzing the patterns of thunderstorms from 2013-2022, which shows that the thunderstorms are not slowing down and will seem to continue their impact on human life in the future. This work also analyzes the monthly and yearly pattern of the impact of thunderstorms on power systems at the national, state, and county level.

Bhusal, Narayan↗

Curation and Dissemination of Complex Multi-Modal Datasets for Radiation Detection, Localization, and Tracking

The PANDAWN sensor network in Chicago, IL, is a state-of-the-art testbed for networked, multi-modal sensing. It integrates AI/data science methods into its operation, from data acquisition to automated data labeling and curation workflows. The curation and dissemination of diverse multi-modal datasets will enable the development of new radiological/nuclear (R/N) detection, localization, and tracking algorithms and methods relevant across the nonproliferation mission space. This article first introduces the PANDAWN sensor network and the features that make it stand out from previous multi-modal data acquisition efforts. We then review the various data streams acquired on the PANDAWN nodes and present the implementation of an automated data curation pipeline that includes the labeling of radiation and contextual data streams. Here, we finally provide a short overview of different studies that leveraged the curated datasets.

Data curation↗

Unveiling the transferability of PLSR models for leaf trait estimation: lessons from a comprehensive analysis with a novel global dataset

Leaf traits are essential for understanding many physiological and ecological processes. Partial least squares regression (PLSR) models with leaf spectroscopy are widely applied for trait estimation, but their transferability across space, time, and plant functional types (PFTs) remains unclear. We compiled a novel dataset of paired leaf traits and spectra, with 47 393 records for >700 species and eight PFTs at 101 globally distributed locations across multiple seasons. Using this dataset, we conducted an unprecedented comprehensive analysis to assess the transferability of PLSR models in estimating leaf traits. While PLSR models demonstrate commendable performance in predicting chlorophyll content, carotenoid, leaf water, and leaf mass per area prediction within their training data space, their efficacy diminishes when extrapolating to new contexts. Specifically, extrapolating to locations, seasons, and PFTs beyond the training data leads to reduced R 2 (0.12–0.49, 0.15–0.42, and 0.25–0.56) and increased NRMSE (3.58–18.24%, 6.27–11.55%, and 7.0–33.12%) compared with nonspatial random cross-validation. The results underscore the importance of incorporating greater spectral diversity in model training to boost its transferability. These findings highlight potential errors in estimating leaf traits across large spatial domains, diverse PFTs, and time due to biased validation schemes, and provide guidance for future field sampling strategies and remote sensing applications.

59 BASIC BIOLOGICAL SCIENCES↗

Distribution System Dataset Generator for AI Applications [SWR-24-75]

This software is a simple, light-weight python package to generate pytorch compatible machine learning graph dataset representing electric power distribution system. User is able to use these graph datasets to test their graph generation artificial intelligence (AI) models, link prediction AI models, graph classification AI models and so much more. This package uses grid-data-models (https://github.com/NREL-Distribution-Suites/grid-data-models) as input data format for power distribution system. NREL-Ditto (https://github.com/NREL-Distribution-Suites/ditto) tool can be leveraged to transform popular distribution system file formats such as opendss, cyme and synergi to grid-data-models.

Duwadi, Kapil↗

Open-Datasets for WEC Simulation

SAND2025-11468O The Open-Datasets for WEC Simulation is a tool that uses datasets and simulation configuration files to conduct OpenFOAM wave energy converter (WEC) simulations. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Chartrand, Chris [Sandia National Lab. (SNL-CA), L↗

Dataset: Breaking the barrier of human-annotated training data for machine-learning-aided plant research using aerial imagery

This dataset supports the implementation described in the manuscript "Breaking the Barrier of Human-Annotated Training Data for Machine-Learning-Aided Biological Research Using Aerial Imagery." It comprises UAV aerial imagery used to execute the code available at https://github.com/pixelvar79/GAN-Flowering-Detection-paper. For detailed information on dataset usage and instructions for implementing the code to reproduce the study, please refer to the GitHub repository.

generative and adversarial learning↗

Putting error bars on density functional theory dataset

This dataset contains submission files and raw output files from high-throughput DFT simulations to analyze the systemic errors in lattice constant, bulk moduli and formation energy predictions for a range of binary and ternary oxides using four exchange correlation functionals (LDA, PBE, PBEsol and vdW-DF-C09). This data was then used as the basis for employing materials informatics methods to predict the expected errors in the lattice constants of the studied compounds. Predicted errors were also used to better the DFT-predicted lattice parameters. Our results emphasize the link between the computed errors and the electron density and hybridization errors of a functional. In essence, these results provide “error bars” for choosing a functional for the creation of high-accuracy, high-throughput datasets as well as avenues for the development of XC functionals with enhanced performance, thereby enabling the accelerated discovery and design of new materials.

36 MATERIALS SCIENCE↗

Li1−xNiO2 Many-body DMC Benchmark Dataset

The dataset contains all numerical data generated in support of the manuscript “Many‑body Benchmark of Electronic Charge and Spin Densities for Li1–xNiO2​” (Journal of Chemical Theory and Computation, DOI: 10.1021/acs.jctc.5c02097, URL: https://pubs.acs.org/doi/10.1021/acs.jctc.5c02097). The materials included in this repository are: 1. Data files used to produce all figures and tables in the main manuscript and supporting information. 2. Benchmark density‑functional theory (DFT) datasets used for the charge‑ and spin‑density analyses. 3. Reference many‑body diffusion Monte Carlo (DMC) calculations and associated input/output files.

36 MATERIALS SCIENCE↗

Dataset for "A primer on forest structure measurement with lidar for ecologists"

This repository includes data and code accompanying the case study included in the manuscript "A primer on forest structure measurement with lidar for ecologists" (submitted to Ecosphere). We compiled lidar datasets from multiple platforms in a common area to: 1. Demonstrate how differences in sensor characteristics influence density and resolution of lidar data. 2. Provide open-source, co-located datasets for users to further inspect differences in lidar data. 3. Provide example code to perform basic lidar analysis. This case study is meant to allow readers to get hands-on experience with real-world data from different platforms. This case study is not meant to be a rigorous comparison of derived ecological metrics among all sensors; such comparisons can be found throughout other publications referenced throughout the main manuscript. Code includes basic functions in R commonly used to visualize and manipulate lidar data accessible with a normal laptop computer; more sophisticated algorithms for advanced users are also referenced throughout the main manuscript. Terrestrial laser scanning (TLS), mobile laser scanning (MLS), UAS laser scanning (ULS), airborne laser scanning (ALS), and spaceborne laser scanning (SLS) data were collected within the Smithsonian Environmental Research Center (SERC) forest dynamics plot in Maryland, USA. TLS, MLS, and ALS data were collected within 1 month of the 2021 growing season; ULS data were collected in November 2020 (“leaf-off” data) and July 2022 (“leaf-on” data).

54 ENVIRONMENTAL SCIENCES↗