Search NASA⌕ Search

SEARCH · Search NASA

Results for “Dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Dataset for "A primer on forest structure measurement with lidar for ecologists"

This repository includes data and code accompanying the case study included in the manuscript "A primer on forest structure measurement with lidar for ecologists" (submitted to Ecosphere). We compiled lidar datasets from multiple platforms in a common area to: 1. Demonstrate how differences in sensor characteristics influence density and resolution of lidar data. 2. Provide open-source, co-located datasets for users to further inspect differences in lidar data. 3. Provide example code to perform basic lidar analysis. This case study is meant to allow readers to get hands-on experience with real-world data from different platforms. This case study is not meant to be a rigorous comparison of derived ecological metrics among all sensors; such comparisons can be found throughout other publications referenced throughout the main manuscript. Code includes basic functions in R commonly used to visualize and manipulate lidar data accessible with a normal laptop computer; more sophisticated algorithms for advanced users are also referenced throughout the main manuscript. Terrestrial laser scanning (TLS), mobile laser scanning (MLS), UAS laser scanning (ULS), airborne laser scanning (ALS), and spaceborne laser scanning (SLS) data were collected within the Smithsonian Environmental Research Center (SERC) forest dynamics plot in Maryland, USA. TLS, MLS, and ALS data were collected within 1 month of the 2021 growing season; ULS data were collected in November 2020 (“leaf-off” data) and July 2022 (“leaf-on” data).

54 ENVIRONMENTAL SCIENCES↗

Global Multimodal Dataset for Nighttime Light Super-Resolution

The dataset is a collection of spatially and temporally registered high-resolution and low-resolution nighttime light (NTL) images, high-resolution land-use binary masks, and high-resolution road density images from around the world. The NTL images are sourced from the NASA Black Marble product VNP46A2 and the LuoJia1-01 satellite. The land-use binary masks are derived from Google's and the World Resources Institute's DynamicWorld dataset, and the road density images are sourced from OpenStreetMap.

Nighttime lights↗

EAGLE-I County Customer Dataset Fall 2025

This dataset provides a combination of modeled and collected county-level electric customer counts derived from 2023 EIA-861 utility customer data, 2021 HIFLD electric retail service territory boundaries, 2021 LandScan population estimates, and 2025 EAGLE-I customer outages. The dataset details county FIPS code, number of customers, and customer type (modeled, collected, mixed). Outage data in included for all 50 U.S. states, Puerto Rico, and the District of Columbia (excluding other U.S. territories).

24 POWER TRANSMISSION AND DISTRIBUTION↗

FleetREDI Insight: Intrastate Coach Bus Dataset

Capturing real-world data is critical to improving efficiency and supporting technology advancements in commercial vehicles. FleetREDI’s insights provide detailed duty cycle information and highlight unique aspects of the given dataset. Each insight delivers a quick look at the collected data by summarizing the operation and identifying key findings of the initial analysis. This FleetREDI insight explores coach buses operating in Colorado. Coach buses are a primary mover for intrastate transit and are primarily used for longer trips with more comfortable seats and a restroom. All Aboard America! Holdings Inc. offers various fixed-service and charter routes across Colorado on its Bustang fleet out of its depot in Golden, Colorado. NLR installed logging devices and collected operational data on nine 40-foot Bustang motorcoaches operating on fixed routes from May through August 2022. Using NLR’s FleetREDI data platform, this dataset provides a summary of daily operation to help understand duty cycle characteristics. This includes daily distance, fuel use, and estimated engine-produced energy consumption for nine motorcoaches that operated more than 33,000 miles. These vehicles primarily operated on Interstate 25 and Interstate 70. ![FleetREDI interstate bus](FleetREDI-interstate-bus.jpg)

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Dataset for scientific paper "Simulated plant‑mediated oxygen input has strong impacts on fine‑scale porewater biogeochemistry and weak impacts on integrated methane fluxes in coastal wetlands", a modeling study based on field observation at the tidal salt marshes of the Parker River Estuary, Massachusetts, United States

This dataset is the raw and processed data for the paper "Simulated plant ‑ mediated oxygen input has strong impacts on fine ‑ scale porewater biogeochemistry and weak impacts on integrated methane fluxes in coastal wetlands". This study investigated how plant-mediated oxygen input affects subsurface biogeochemical reactions of organic carbon degradation and the resulting methane emissions of coastal wetlands by model simulation. We used the subsurface geochemical simulator PFLOTRAN for the modeling, which produced the simulated changes in porewater chemical substances and methane emissions over 10 days under different scenarios of plant-mediated oxygen input.Specifically, this dataset contains: 1) the input files for PFLOTRAN of all simulation runs conducted in this study. Those files are with an extension of ".in", containing information of the biogeochemical reaction network (stoichiometry, reaction rate, Monod constants, etc), fluid flow rate and oxygen concentration in the fluid which together simulated the plant-mediated oxygen input, the configuration of artificial reactions that simulated the methane fluxes, etc. The PFLOTRAN input files are text files, which can be opened by NotePad, but running these input files will require proper installation of PFLOTRAN (instruction: https://documentation.pflotran.org/user_guide/how_to/installation/installation.html). 2) the raw and processed model output from PFLOTRAN of all simulation runs, and 3) the python scripts used to process the raw model output, including random allocation of root cells, converting raw data into organized formats, calculating the methane fluxes based on the model output, data visualization, etc. The raw and processed model output from PFLOTRAN are in .spydata format, which can be viewed with Python. and 3) the python scripts for data processing and analysis are programming scripts, which can be opened with Python.This modeling work, in particular the model parameterization of root density and initial conditions of porewater concentrations of biogeochemical substances, was based on field measurements at the salt marsh of the Upper Parker River Estuary, Massachusetts, United States.

54 ENVIRONMENTAL SCIENCES↗

Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification" Willard et al. (2025).

This data release provides all data and code used in the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025)" to model stream temperature, evaluate, and assess results. The associated manuscript explores the effect of different ensemble construction techniques across different common machine learning (ML) architectures for predictions in unmonitored basins. Modeling was done using long short-term memory (LSTM), gated recurrent unit (GRU), temporal convolution network (TCN), and extreme gradient boosting (XGBoost) models, and stream site coverage spans 1362 locations across the conterminous United States. The ensemble construction techniques investigated include ensemble by random weight initialization, differing hyperparameters, different random subsets of training data, different subselections of input features, different architectures, and Monte Carlo Dropout. The data is organized into these items items:Code repository and data for the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025).Code: stream_temp_ml_regionalization.zip contains the code repositoryData to run the code:- data_dir.zip -- contains all files that should be moved to the "DATA_DIR" variable defined in the "set_env_vars.sh" script in the code repository- metadata_dir.zip -- contains all files that should be moved to the "METADATA_DIR" variable defined in the "set_env_vars.sh" script in the code repositoryData produced by the code and used in the paper:- outputs_dir.zip - contains model output and results (outputs_dir/results), model weights (outputs_dir/models), and all other outputs used for the paper including feature importances.To cite this code, please use the following BibTeX or MLA entries:bibtex:@misc{willard2025streamensembles,author = {Jared Willard and Charuleka Varadharajan},title = {Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification"},year = {2024},doi = {10.15485/2527393},publisher = {ESS-DIVE Repository},url = {https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2527393}}MLA: Willard, Jared, et al. Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification". 2025. ESS-DIVE Repository, doi:10.15485/2448016.

54 ENVIRONMENTAL SCIENCES↗

CROCUS Air Quality Dataset from the University of Illinois Chicago (UIC), July 2024

This dataset was collected by the measurement system in the Atmosphere, Climate, and Ecosystems (ACE) Lab at the University of Illinois Chicago (UIC) from July 12 to July 31, 2024, as part of the Community Research on Climate and Urban Science (CROCUS) Urban Integrated Field Laboratory (UIFL) project, led by Argonne National Laboratory.To enhance understanding of urban air quality dynamics in Chicago, and as part of the CROCUS 2024 Urban Canyon Intensive Observation Period (IOP), several instruments were set up to provide continuous measurements of air quality parameters in Chicago during July 2024. These measurements cover both aerosols and gas-phase species. It focuses on particle size distribution (2.5–478 nm) measured by two Scanning Mobility Particle Sizers (SMPS) at a 4-min resolution, total particle number concentrations at a 1-s resolution, and chemical composition from a High-Resolution Time-of-Flight Aerosol Mass Spectrometer (AMS) at a 1-min resolution. Key gas-phase species, including NO, NO₂, SO₂, and O₃, are measured at a 1-min resolution, along with high-resolution NO and dimethyl sulfide (DMS) data from a Chemical Ionization Mass Spectrometer (CIMS). Volatile organic compound (VOC) data for toluene, isoprene, and benzene are provided by a GC-PID with a time resolution of 25 minutes.The data are formatted as NetCDF (.nc) files, making them easily accessible using common software such as MATLAB, R, and Python. Each parameter is stored in an individual dataset, which includes detailed instrument information in the header, as well as the corresponding sample start time and concentration/distribution data for each sample.

54 ENVIRONMENTAL SCIENCES↗

Solar, Wind, and Load Forecasting Dataset for MISO, NYISO, and SPP Balancing Areas

The Performance-based Energy Resource Feedback, Optimization, and Risk Management (PERFORM) program is an initiative intended to foster "a fundamental shift in grid management rooted in an understanding of asset risk and system risk" (ARPA-E 2020). Launched by the Advanced Research Projects Agency-Energy (ARPA-E), the program supports efforts to incorporate uncertainty in electric power decision making. In support of PERFORM, the National Renewable Energy Laboratory (NREL) has produced a set of time-coincident forecasts of solar, wind, and load profiles. As part of Phase I of the PERFORM effort, NREL created a dataset that consists of one year of time-coincident load, wind, and solar actuals and probabilistic forecasts based on data from the Electric Reliability Council of Texas (ERCOT) (Bryce et al. 2023). In Phase II, NREL developed similar datasets for three other U.S. Independent System Operators (ISO): the Midcontinent Independent System Operator (MISO), the New York Independent System Operator (NYISO), and the Southwest Power Pool (SPP).

24 POWER TRANSMISSION AND DISTRIBUTION↗

The Foundational Industrial Energy Dataset (FIED): Open-Source Data on Industrial Facilities

The state of data on industrial energy use has co-evolved over several decades with the demands of industrial energy analysis. The most recent development - analysis in support of decarbonizing the industrial sector - has changed the characteristics of industrial data that are useful for analysts and model developers. Although data and its collection processes may be cast from a conventional viewpoint as objective and free from the influence of social dynamics, this provides an incomplete picture of not only the processes by which information is generated, but also the limitations and opportunities of data to be useful for analysis. The foundational industry energy data set (FIED) is a result of the confluence of trends in open data and the demand for higher resolution industrial energy analysis. The general approach to compiling the FIED involves accessing, filtering, and formatting data published by federal organizations on the Internet for public use. Unlike most industrial energy datasets, which are published by the U.S. Energy Information Administration (EIA), the FIED relies on core datasets from the U.S. Environmental Protection Agency (EPA). The FIED addresses several of the areas of growing disconnect between the demands of industrial energy analysis and the state of industrial energy data by providing unit-level characterization - including estimates of energy use, greenhouse gas emissions, and design capacities - for facilities that are identified by latitude and longitude. This enables local-level analysis of existing combustion equipment, as well as regional comparisons with traditional industrial energy data estimates. The report summarizes the general logic behind compiling the FIED. The FIED itself and its Python code are available from OpenEI and GitHub, respectively.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

LUCID Thrust 1 - Dataset Identification and Biodata Catalog Creation

The LUCID DOE consortium, part of the Department of Energy’s Biological and Environmental Research (BER) program, advances Low Dose Radiation (LDR) research through multidisciplinary efforts across seven key thrusts. This document focuses on Thrust 1, which centers on the creation of curated multimodal population health datasets and supports broader efforts within the LUCID program, including AI-based hypothesis generation, experimental design, and the study of LDR-induced health risks. Specifically, it describes the identification and cataloging of Thrust 1’s curated LDR datasets and biodata, emphasizing their critical role in supporting various research thrusts within the consortium, with potential applications in healthcare and public policy. In addition, the document includes an evaluation of three Large Language Models (LLMs)—GPT-4, SOLAR-10B, and Mixtral-8x7B—based on their ability to extract features from 25 LDR studies. The results indicate that GPT-4 performed the best, while Mixtral-8x7B demonstrated limited knowledge. Overall, this work advances understanding in radiation protection, risk assessment, and medical treatments, while providing valuable resources for researchers, educators, and policymakers.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Basin-Scale Structural Features Database: Spatial Datasets to Support Carbon Storage Resource Assessments

Presentation slides on "Basin-Scale Structural Features Database: Spatial Datasets to Support Carbon Storage Resource Assessments" for CCUS 2025 Annual Meeting. The Basin-Scale Structural Features database contains a series of basin-scale spatial datasets representing structural features, including faults, fractures, folds, and earthquakes. Designed to support carbon storage feasibility and resources assessments for Carbon Capture and Storage (CCS) projects, the database leverages publicly available data resources from authoritative sources (e.g. US Geological Survey, State Geologic Surveys), and aims to help users better understand basin-scale structural features, as well as potential data gaps in areas with sparse information.

basin scale↗

An Exploratory Data Mining Investigation for Constructing a Publicly Sourced Dataset of Foreign Hypersonic Tests

This document details a data mining exercise that resulted in an exploratory dataset of publicly reported foreign (non-US) hypersonic vehicle test events. Using a combination of targeted English language searches and country-specific queries, the study aggregates information from digital news media, official press releases, and social media posts. The resulting list of events captures the publicly available accounts of foreign hypersonic tests, although it does not represent an exhaustive record. Limitations such as inconsistent reporting, translation challenges, and the inherently provisional nature of open-source data are acknowledged. This dataset serves as an initial reference point for further inquiries into high-speed atmospheric phenomena and may facilitate future efforts to correlate these events with geophysical measurements.

33 ADVANCED PROPULSION SYSTEMS↗

Hourly Load Profile Dataset for Federal, State, and Municipal Electric Vehicle Fleets in the United States

The electrification of U.S. federal, state, and municipal fleets is accelerating rapidly, driven by an increased availability of competitive electric vehicle (EV) options and supportive policies and targets. The dataset described in this report, accessible at data.nrel.gov/submissions/280, provides a critical foundation for identifying fleet electricity demand, projecting these future demands, and developing actionable strategies to support the widespread electrification of government fleets. The dataset incorporates available fleet data, including 54% of federal agency vehicles approved for analysis (notably, the U.S. Postal Service is absent). Additionally, it includes data from 50,000 state government vehicles and 94,000 local government vehicles. While this represents a small fraction of the 4.4 million vehicles owned by state and local governments reported by the Federal Highway Administration (2022), the framework supports future expansion as more fleet inventory data become available.

33 ADVANCED PROPULSION SYSTEMS↗

Foundational Dataset for Developing Large-Sample Stream Temperature Models in the Conterminous United States

This dataset provides inputs, evaluation results, and trained weights from a large-sample Long Short-Term Memory (LSTM) model designed to predict daily stream temperatures across unregulated river reaches in the conterminous United States (CONUS). It includes dynamic meteorological and hydrologic forcings, static physiographic attributes, and model outputs from cross-validation experiments spanning 300 basins. It supports reproducible modeling, direct application for new basins, and provides data suitable for integration with reservoir and river simulations under current and future climates. It contains two .zip files described below · RQ-AI_runs.zip: Model outputs from 10-fold cross-validation experiments, including observed and predicted daily stream temperatures, along with test performance metrics for water years 2017–2019. Two versions are included: 1. Model trained and validated using subbasin-area weighted dynamic features. 2. Model trained and validated using whole-basin area weighted dynamic features. · RQ-AI_inputs.zip: Collection of all formatted dynamic and static predictor datasets (meteorological, hydrologic, and physiographic features) used in model training and analysis. Detailed instructions and data structure is held at the following GitLab repository: https://code.ornl.gov/tempwise/training.

Gomez-Velez, Jesus [Oak Ridge National Laboratory ↗

Hydrological dataset for reservoir sedimentation in Texas

This dataset provides comprehensive hydrological information about reservoir sedimentation in Texas, including observed and remotely sensed sediment concentration, river discharge, watershed boundary, lake geometry, reservoir capacity, land use and land cover (LULC), and population data. In-depth interpretation of the dataset is elaborated in in a journal article, entitled " Sedimentation and nonlinear trapping in Texas reservoirs identified using remote sensing and bathymetric survey records (will be accepted soon at Water Resources Research)."

Lee, Jiyong [Oak Ridge National Laboratory (ORNL),↗

National Park Air Quality Index Dataset

The National Park Air Quality Index dataset (NPS-AQI) consists of webcam images taken from the National Park Service's publicly available air quality web cameras and associated measurements for air pollutants, AQI, and meteorological data obtained via the publicly available NPS Gaseous Pollutant Monitoring Program. The full dataset is a collection of 146,822 images paired with air quality measurements. The specific measurements reported are: ozone ppm, 8-hour running average ozone ppm, so2 ppm, AQI (derived from ozone), temperature, and humidity. The images are 1500X1000 pixel PNG files arranged into folders by NPS site and named according to the time and date the image was taken. There are three CSV files (representing "training", "validation", and "testing" images splits) containing image names and associated NPS site names, air pollutant measurements, and meteorlogical data.

Svinth, Christian N↗

Datasets for Custom-trained Machine-learning Interatomic Potentials: Nitric Acid Aqueous Solution

This dataset was generated using an iterative active learning strategy with the ArcaNN software package (https://github.com/arcann-chem/arcann_training) to train machine-learning interatomic potentials (MLIPs) for aqueous nitric acid. Each active-learning cycle consisted of three stages: (1) training, (2) exploration, and (3) labeling. The initial training set comprised approximately 800 randomly selected configurations from a previous study by Lewis et al. (https://doi.org/10.1021/jp205510q), which investigated nitric acid solutions at 2, 3, 4, and 5 mol/L. For all configurations, single-point calculations of atomic forces and total energies were performed at the quantum density functional theory BLYP-D2 and PBE-D3 levels of theory using the CP2K Quickstep module. Valence electrons were treated explicitly, while core electrons on all atoms were represented by norm-conserving Goedecker–Teter–Hutter (GTH) pseudopotentials. Long-range dispersion interactions were accounted for using Grimme dispersion corrections. Wave functions were expanded in a mixed Gaussian-and-plane-wave scheme using TZV2P-MOLOPT basis sets for all elements and an 800 Ry auxiliary plane-wave cutoff for the electron density. Self-consistent field convergence was accelerated using orbital transformation and Direct Inversion in the Iterative Subspace, with a convergence threshold of 10^{-6}. All single-point calculations were carried out in periodic orthorhombic cells whose dimensions match those of the molecular configurations sampled from earlier trajectories. The CELL_REF keyword in CP2K was used to define a fixed reference cell, ensuring consistency in the reference data used for MLIP training, particularly when cell fluctuations are present in NpT simulations. The resulting high-fidelity energies and forces constitute the ground-truth labels used to train the MLIPs contained in this dataset.

Dinpajooh, Mohammadhasan [Pacific Northwest Nation↗