Search NASA⌕ Search

SEARCH · Search NASA

Results for “Dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

From Data to Insights: A Covariate Analysis of the IARPA BRIAR Dataset for Multimodal Biometric Recognition Algorithms at Altitude and Range

This paper examines covariate effects on fused whole body biometrics performance in the IARPA BRIAR dataset, specifically focusing on UAV platforms, elevated positions, and distances up to 1000 meters. The dataset includes outdoor videos compared with indoor images and controlled gait recordings. Normalized raw fusion scores relate directly to predicted false accept rates (FAR), offering an intuitive means for interpreting model results. A linear model is developed to predict biometric algorithm scores, analyzing their performance to identify the most influential covariates on accuracy at altitude and range. Weather factors like temperature, wind speed, solar loading, and turbulence are also investigated in this analysis. The study found that resolution and camera distance best predicted accuracy and findings can guide future research and development efforts in long-range/elevated/UAV biometrics and support the creation of more reliable and robust systems for national security and other critical domains.

Bolme, David↗

Impact Study of Thunderstorms on the US Power Grid Using Publicly Available Datasets

This work analyzes the impact of thunderstorms on the US power grid based on publicly available data. Since thunderstorms can bring lightning, heavy precipitation, and wind storms, analyzing their impact on the power system provides a combined correlation of lightning strikes, floods, and wind storms on power outages. This paper leverages publicly available thunderstorm datasets from the National Weather Service (NWS) and power outage datasets from Oak Ridge National Laboratory’s Environment for Analysis of Geo-Located Energy Information (EAGLE-I) to study the correlation between thunderstorms and power outages. This work is analyzing the patterns of thunderstorms from 2013-2022, which shows that the thunderstorms are not slowing down and will seem to continue their impact on human life in the future. This work also analyzes the monthly and yearly pattern of the impact of thunderstorms on power systems at the national, state, and county level.

Bhusal, Narayan↗

Curation and Dissemination of Complex Multi-Modal Datasets for Radiation Detection, Localization, and Tracking

The PANDAWN sensor network in Chicago, IL, is a state-of-the-art testbed for networked, multi-modal sensing. It integrates AI/data science methods into its operation, from data acquisition to automated data labeling and curation workflows. The curation and dissemination of diverse multi-modal datasets will enable the development of new radiological/nuclear (R/N) detection, localization, and tracking algorithms and methods relevant across the nonproliferation mission space. This article first introduces the PANDAWN sensor network and the features that make it stand out from previous multi-modal data acquisition efforts. We then review the various data streams acquired on the PANDAWN nodes and present the implementation of an automated data curation pipeline that includes the labeling of radiation and contextual data streams. Here, we finally provide a short overview of different studies that leveraged the curated datasets.

Data curation↗

Unveiling the transferability of PLSR models for leaf trait estimation: lessons from a comprehensive analysis with a novel global dataset

Leaf traits are essential for understanding many physiological and ecological processes. Partial least squares regression (PLSR) models with leaf spectroscopy are widely applied for trait estimation, but their transferability across space, time, and plant functional types (PFTs) remains unclear. We compiled a novel dataset of paired leaf traits and spectra, with 47 393 records for >700 species and eight PFTs at 101 globally distributed locations across multiple seasons. Using this dataset, we conducted an unprecedented comprehensive analysis to assess the transferability of PLSR models in estimating leaf traits. While PLSR models demonstrate commendable performance in predicting chlorophyll content, carotenoid, leaf water, and leaf mass per area prediction within their training data space, their efficacy diminishes when extrapolating to new contexts. Specifically, extrapolating to locations, seasons, and PFTs beyond the training data leads to reduced R 2 (0.12–0.49, 0.15–0.42, and 0.25–0.56) and increased NRMSE (3.58–18.24%, 6.27–11.55%, and 7.0–33.12%) compared with nonspatial random cross-validation. The results underscore the importance of incorporating greater spectral diversity in model training to boost its transferability. These findings highlight potential errors in estimating leaf traits across large spatial domains, diverse PFTs, and time due to biased validation schemes, and provide guidance for future field sampling strategies and remote sensing applications.

59 BASIC BIOLOGICAL SCIENCES↗

Distribution System Dataset Generator for AI Applications [SWR-24-75]

This software is a simple, light-weight python package to generate pytorch compatible machine learning graph dataset representing electric power distribution system. User is able to use these graph datasets to test their graph generation artificial intelligence (AI) models, link prediction AI models, graph classification AI models and so much more. This package uses grid-data-models (https://github.com/NREL-Distribution-Suites/grid-data-models) as input data format for power distribution system. NREL-Ditto (https://github.com/NREL-Distribution-Suites/ditto) tool can be leveraged to transform popular distribution system file formats such as opendss, cyme and synergi to grid-data-models.

Duwadi, Kapil↗

Open-Datasets for WEC Simulation

SAND2025-11468O The Open-Datasets for WEC Simulation is a tool that uses datasets and simulation configuration files to conduct OpenFOAM wave energy converter (WEC) simulations. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Chartrand, Chris [Sandia National Lab. (SNL-CA), L↗

Dataset: Breaking the barrier of human-annotated training data for machine-learning-aided plant research using aerial imagery

This dataset supports the implementation described in the manuscript "Breaking the Barrier of Human-Annotated Training Data for Machine-Learning-Aided Biological Research Using Aerial Imagery." It comprises UAV aerial imagery used to execute the code available at https://github.com/pixelvar79/GAN-Flowering-Detection-paper. For detailed information on dataset usage and instructions for implementing the code to reproduce the study, please refer to the GitHub repository.

generative and adversarial learning↗

Putting error bars on density functional theory dataset

This dataset contains submission files and raw output files from high-throughput DFT simulations to analyze the systemic errors in lattice constant, bulk moduli and formation energy predictions for a range of binary and ternary oxides using four exchange correlation functionals (LDA, PBE, PBEsol and vdW-DF-C09). This data was then used as the basis for employing materials informatics methods to predict the expected errors in the lattice constants of the studied compounds. Predicted errors were also used to better the DFT-predicted lattice parameters. Our results emphasize the link between the computed errors and the electron density and hybridization errors of a functional. In essence, these results provide “error bars” for choosing a functional for the creation of high-accuracy, high-throughput datasets as well as avenues for the development of XC functionals with enhanced performance, thereby enabling the accelerated discovery and design of new materials.

36 MATERIALS SCIENCE↗

Li1−xNiO2 Many-body DMC Benchmark Dataset

The dataset contains all numerical data generated in support of the manuscript “Many‑body Benchmark of Electronic Charge and Spin Densities for Li1–xNiO2​” (Journal of Chemical Theory and Computation, DOI: 10.1021/acs.jctc.5c02097, URL: https://pubs.acs.org/doi/10.1021/acs.jctc.5c02097). The materials included in this repository are: 1. Data files used to produce all figures and tables in the main manuscript and supporting information. 2. Benchmark density‑functional theory (DFT) datasets used for the charge‑ and spin‑density analyses. 3. Reference many‑body diffusion Monte Carlo (DMC) calculations and associated input/output files.

36 MATERIALS SCIENCE↗

Dataset for "A primer on forest structure measurement with lidar for ecologists"

This repository includes data and code accompanying the case study included in the manuscript "A primer on forest structure measurement with lidar for ecologists" (submitted to Ecosphere). We compiled lidar datasets from multiple platforms in a common area to: 1. Demonstrate how differences in sensor characteristics influence density and resolution of lidar data. 2. Provide open-source, co-located datasets for users to further inspect differences in lidar data. 3. Provide example code to perform basic lidar analysis. This case study is meant to allow readers to get hands-on experience with real-world data from different platforms. This case study is not meant to be a rigorous comparison of derived ecological metrics among all sensors; such comparisons can be found throughout other publications referenced throughout the main manuscript. Code includes basic functions in R commonly used to visualize and manipulate lidar data accessible with a normal laptop computer; more sophisticated algorithms for advanced users are also referenced throughout the main manuscript. Terrestrial laser scanning (TLS), mobile laser scanning (MLS), UAS laser scanning (ULS), airborne laser scanning (ALS), and spaceborne laser scanning (SLS) data were collected within the Smithsonian Environmental Research Center (SERC) forest dynamics plot in Maryland, USA. TLS, MLS, and ALS data were collected within 1 month of the 2021 growing season; ULS data were collected in November 2020 (“leaf-off” data) and July 2022 (“leaf-on” data).

54 ENVIRONMENTAL SCIENCES↗

Global Multimodal Dataset for Nighttime Light Super-Resolution

The dataset is a collection of spatially and temporally registered high-resolution and low-resolution nighttime light (NTL) images, high-resolution land-use binary masks, and high-resolution road density images from around the world. The NTL images are sourced from the NASA Black Marble product VNP46A2 and the LuoJia1-01 satellite. The land-use binary masks are derived from Google's and the World Resources Institute's DynamicWorld dataset, and the road density images are sourced from OpenStreetMap.

Nighttime lights↗

EAGLE-I County Customer Dataset Fall 2025

This dataset provides a combination of modeled and collected county-level electric customer counts derived from 2023 EIA-861 utility customer data, 2021 HIFLD electric retail service territory boundaries, 2021 LandScan population estimates, and 2025 EAGLE-I customer outages. The dataset details county FIPS code, number of customers, and customer type (modeled, collected, mixed). Outage data in included for all 50 U.S. states, Puerto Rico, and the District of Columbia (excluding other U.S. territories).

24 POWER TRANSMISSION AND DISTRIBUTION↗

FleetREDI Insight: Intrastate Coach Bus Dataset

Capturing real-world data is critical to improving efficiency and supporting technology advancements in commercial vehicles. FleetREDI’s insights provide detailed duty cycle information and highlight unique aspects of the given dataset. Each insight delivers a quick look at the collected data by summarizing the operation and identifying key findings of the initial analysis. This FleetREDI insight explores coach buses operating in Colorado. Coach buses are a primary mover for intrastate transit and are primarily used for longer trips with more comfortable seats and a restroom. All Aboard America! Holdings Inc. offers various fixed-service and charter routes across Colorado on its Bustang fleet out of its depot in Golden, Colorado. NLR installed logging devices and collected operational data on nine 40-foot Bustang motorcoaches operating on fixed routes from May through August 2022. Using NLR’s FleetREDI data platform, this dataset provides a summary of daily operation to help understand duty cycle characteristics. This includes daily distance, fuel use, and estimated engine-produced energy consumption for nine motorcoaches that operated more than 33,000 miles. These vehicles primarily operated on Interstate 25 and Interstate 70. ![FleetREDI interstate bus](FleetREDI-interstate-bus.jpg)

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Dataset for scientific paper "Simulated plant‑mediated oxygen input has strong impacts on fine‑scale porewater biogeochemistry and weak impacts on integrated methane fluxes in coastal wetlands", a modeling study based on field observation at the tidal salt marshes of the Parker River Estuary, Massachusetts, United States

This dataset is the raw and processed data for the paper "Simulated plant ‑ mediated oxygen input has strong impacts on fine ‑ scale porewater biogeochemistry and weak impacts on integrated methane fluxes in coastal wetlands". This study investigated how plant-mediated oxygen input affects subsurface biogeochemical reactions of organic carbon degradation and the resulting methane emissions of coastal wetlands by model simulation. We used the subsurface geochemical simulator PFLOTRAN for the modeling, which produced the simulated changes in porewater chemical substances and methane emissions over 10 days under different scenarios of plant-mediated oxygen input.Specifically, this dataset contains: 1) the input files for PFLOTRAN of all simulation runs conducted in this study. Those files are with an extension of ".in", containing information of the biogeochemical reaction network (stoichiometry, reaction rate, Monod constants, etc), fluid flow rate and oxygen concentration in the fluid which together simulated the plant-mediated oxygen input, the configuration of artificial reactions that simulated the methane fluxes, etc. The PFLOTRAN input files are text files, which can be opened by NotePad, but running these input files will require proper installation of PFLOTRAN (instruction: https://documentation.pflotran.org/user_guide/how_to/installation/installation.html). 2) the raw and processed model output from PFLOTRAN of all simulation runs, and 3) the python scripts used to process the raw model output, including random allocation of root cells, converting raw data into organized formats, calculating the methane fluxes based on the model output, data visualization, etc. The raw and processed model output from PFLOTRAN are in .spydata format, which can be viewed with Python. and 3) the python scripts for data processing and analysis are programming scripts, which can be opened with Python.This modeling work, in particular the model parameterization of root density and initial conditions of porewater concentrations of biogeochemical substances, was based on field measurements at the salt marsh of the Upper Parker River Estuary, Massachusetts, United States.

54 ENVIRONMENTAL SCIENCES↗

Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification" Willard et al. (2025).

This data release provides all data and code used in the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025)" to model stream temperature, evaluate, and assess results. The associated manuscript explores the effect of different ensemble construction techniques across different common machine learning (ML) architectures for predictions in unmonitored basins. Modeling was done using long short-term memory (LSTM), gated recurrent unit (GRU), temporal convolution network (TCN), and extreme gradient boosting (XGBoost) models, and stream site coverage spans 1362 locations across the conterminous United States. The ensemble construction techniques investigated include ensemble by random weight initialization, differing hyperparameters, different random subsets of training data, different subselections of input features, different architectures, and Monte Carlo Dropout. The data is organized into these items items:Code repository and data for the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025).Code: stream_temp_ml_regionalization.zip contains the code repositoryData to run the code:- data_dir.zip -- contains all files that should be moved to the "DATA_DIR" variable defined in the "set_env_vars.sh" script in the code repository- metadata_dir.zip -- contains all files that should be moved to the "METADATA_DIR" variable defined in the "set_env_vars.sh" script in the code repositoryData produced by the code and used in the paper:- outputs_dir.zip - contains model output and results (outputs_dir/results), model weights (outputs_dir/models), and all other outputs used for the paper including feature importances.To cite this code, please use the following BibTeX or MLA entries:bibtex:@misc{willard2025streamensembles,author = {Jared Willard and Charuleka Varadharajan},title = {Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification"},year = {2024},doi = {10.15485/2527393},publisher = {ESS-DIVE Repository},url = {https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2527393}}MLA: Willard, Jared, et al. Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification". 2025. ESS-DIVE Repository, doi:10.15485/2448016.

54 ENVIRONMENTAL SCIENCES↗

CROCUS Air Quality Dataset from the University of Illinois Chicago (UIC), July 2024

This dataset was collected by the measurement system in the Atmosphere, Climate, and Ecosystems (ACE) Lab at the University of Illinois Chicago (UIC) from July 12 to July 31, 2024, as part of the Community Research on Climate and Urban Science (CROCUS) Urban Integrated Field Laboratory (UIFL) project, led by Argonne National Laboratory.To enhance understanding of urban air quality dynamics in Chicago, and as part of the CROCUS 2024 Urban Canyon Intensive Observation Period (IOP), several instruments were set up to provide continuous measurements of air quality parameters in Chicago during July 2024. These measurements cover both aerosols and gas-phase species. It focuses on particle size distribution (2.5–478 nm) measured by two Scanning Mobility Particle Sizers (SMPS) at a 4-min resolution, total particle number concentrations at a 1-s resolution, and chemical composition from a High-Resolution Time-of-Flight Aerosol Mass Spectrometer (AMS) at a 1-min resolution. Key gas-phase species, including NO, NO₂, SO₂, and O₃, are measured at a 1-min resolution, along with high-resolution NO and dimethyl sulfide (DMS) data from a Chemical Ionization Mass Spectrometer (CIMS). Volatile organic compound (VOC) data for toluene, isoprene, and benzene are provided by a GC-PID with a time resolution of 25 minutes.The data are formatted as NetCDF (.nc) files, making them easily accessible using common software such as MATLAB, R, and Python. Each parameter is stored in an individual dataset, which includes detailed instrument information in the header, as well as the corresponding sample start time and concentration/distribution data for each sample.

54 ENVIRONMENTAL SCIENCES↗

The Foundational Industrial Energy Dataset (FIED): Open-Source Data on Industrial Facilities

The state of data on industrial energy use has co-evolved over several decades with the demands of industrial energy analysis. The most recent development - analysis in support of decarbonizing the industrial sector - has changed the characteristics of industrial data that are useful for analysts and model developers. Although data and its collection processes may be cast from a conventional viewpoint as objective and free from the influence of social dynamics, this provides an incomplete picture of not only the processes by which information is generated, but also the limitations and opportunities of data to be useful for analysis. The foundational industry energy data set (FIED) is a result of the confluence of trends in open data and the demand for higher resolution industrial energy analysis. The general approach to compiling the FIED involves accessing, filtering, and formatting data published by federal organizations on the Internet for public use. Unlike most industrial energy datasets, which are published by the U.S. Energy Information Administration (EIA), the FIED relies on core datasets from the U.S. Environmental Protection Agency (EPA). The FIED addresses several of the areas of growing disconnect between the demands of industrial energy analysis and the state of industrial energy data by providing unit-level characterization - including estimates of energy use, greenhouse gas emissions, and design capacities - for facilities that are identified by latitude and longitude. This enables local-level analysis of existing combustion equipment, as well as regional comparisons with traditional industrial energy data estimates. The report summarizes the general logic behind compiling the FIED. The FIED itself and its Python code are available from OpenEI and GitHub, respectively.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗