Search NASA⌕ Search

SEARCH · Search NASA

Results for “Dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

LUCID Thrust 1 - Dataset Identification and Biodata Catalog Creation

The LUCID DOE consortium, part of the Department of Energy’s Biological and Environmental Research (BER) program, advances Low Dose Radiation (LDR) research through multidisciplinary efforts across seven key thrusts. This document focuses on Thrust 1, which centers on the creation of curated multimodal population health datasets and supports broader efforts within the LUCID program, including AI-based hypothesis generation, experimental design, and the study of LDR-induced health risks. Specifically, it describes the identification and cataloging of Thrust 1’s curated LDR datasets and biodata, emphasizing their critical role in supporting various research thrusts within the consortium, with potential applications in healthcare and public policy. In addition, the document includes an evaluation of three Large Language Models (LLMs)—GPT-4, SOLAR-10B, and Mixtral-8x7B—based on their ability to extract features from 25 LDR studies. The results indicate that GPT-4 performed the best, while Mixtral-8x7B demonstrated limited knowledge. Overall, this work advances understanding in radiation protection, risk assessment, and medical treatments, while providing valuable resources for researchers, educators, and policymakers.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Basin-Scale Structural Features Database: Spatial Datasets to Support Carbon Storage Resource Assessments

Presentation slides on "Basin-Scale Structural Features Database: Spatial Datasets to Support Carbon Storage Resource Assessments" for CCUS 2025 Annual Meeting. The Basin-Scale Structural Features database contains a series of basin-scale spatial datasets representing structural features, including faults, fractures, folds, and earthquakes. Designed to support carbon storage feasibility and resources assessments for Carbon Capture and Storage (CCS) projects, the database leverages publicly available data resources from authoritative sources (e.g. US Geological Survey, State Geologic Surveys), and aims to help users better understand basin-scale structural features, as well as potential data gaps in areas with sparse information.

basin scale↗

An Exploratory Data Mining Investigation for Constructing a Publicly Sourced Dataset of Foreign Hypersonic Tests

This document details a data mining exercise that resulted in an exploratory dataset of publicly reported foreign (non-US) hypersonic vehicle test events. Using a combination of targeted English language searches and country-specific queries, the study aggregates information from digital news media, official press releases, and social media posts. The resulting list of events captures the publicly available accounts of foreign hypersonic tests, although it does not represent an exhaustive record. Limitations such as inconsistent reporting, translation challenges, and the inherently provisional nature of open-source data are acknowledged. This dataset serves as an initial reference point for further inquiries into high-speed atmospheric phenomena and may facilitate future efforts to correlate these events with geophysical measurements.

33 ADVANCED PROPULSION SYSTEMS↗

Hourly Load Profile Dataset for Federal, State, and Municipal Electric Vehicle Fleets in the United States

The electrification of U.S. federal, state, and municipal fleets is accelerating rapidly, driven by an increased availability of competitive electric vehicle (EV) options and supportive policies and targets. The dataset described in this report, accessible at data.nrel.gov/submissions/280, provides a critical foundation for identifying fleet electricity demand, projecting these future demands, and developing actionable strategies to support the widespread electrification of government fleets. The dataset incorporates available fleet data, including 54% of federal agency vehicles approved for analysis (notably, the U.S. Postal Service is absent). Additionally, it includes data from 50,000 state government vehicles and 94,000 local government vehicles. While this represents a small fraction of the 4.4 million vehicles owned by state and local governments reported by the Federal Highway Administration (2022), the framework supports future expansion as more fleet inventory data become available.

33 ADVANCED PROPULSION SYSTEMS↗

Foundational Dataset for Developing Large-Sample Stream Temperature Models in the Conterminous United States

This dataset provides inputs, evaluation results, and trained weights from a large-sample Long Short-Term Memory (LSTM) model designed to predict daily stream temperatures across unregulated river reaches in the conterminous United States (CONUS). It includes dynamic meteorological and hydrologic forcings, static physiographic attributes, and model outputs from cross-validation experiments spanning 300 basins. It supports reproducible modeling, direct application for new basins, and provides data suitable for integration with reservoir and river simulations under current and future climates. It contains two .zip files described below · RQ-AI_runs.zip: Model outputs from 10-fold cross-validation experiments, including observed and predicted daily stream temperatures, along with test performance metrics for water years 2017–2019. Two versions are included: 1. Model trained and validated using subbasin-area weighted dynamic features. 2. Model trained and validated using whole-basin area weighted dynamic features. · RQ-AI_inputs.zip: Collection of all formatted dynamic and static predictor datasets (meteorological, hydrologic, and physiographic features) used in model training and analysis. Detailed instructions and data structure is held at the following GitLab repository: https://code.ornl.gov/tempwise/training.

Gomez-Velez, Jesus [Oak Ridge National Laboratory ↗

Hydrological dataset for reservoir sedimentation in Texas

This dataset provides comprehensive hydrological information about reservoir sedimentation in Texas, including observed and remotely sensed sediment concentration, river discharge, watershed boundary, lake geometry, reservoir capacity, land use and land cover (LULC), and population data. In-depth interpretation of the dataset is elaborated in in a journal article, entitled " Sedimentation and nonlinear trapping in Texas reservoirs identified using remote sensing and bathymetric survey records (will be accepted soon at Water Resources Research)."

Lee, Jiyong [Oak Ridge National Laboratory (ORNL),↗

National Park Air Quality Index Dataset

The National Park Air Quality Index dataset (NPS-AQI) consists of webcam images taken from the National Park Service's publicly available air quality web cameras and associated measurements for air pollutants, AQI, and meteorological data obtained via the publicly available NPS Gaseous Pollutant Monitoring Program. The full dataset is a collection of 146,822 images paired with air quality measurements. The specific measurements reported are: ozone ppm, 8-hour running average ozone ppm, so2 ppm, AQI (derived from ozone), temperature, and humidity. The images are 1500X1000 pixel PNG files arranged into folders by NPS site and named according to the time and date the image was taken. There are three CSV files (representing "training", "validation", and "testing" images splits) containing image names and associated NPS site names, air pollutant measurements, and meteorlogical data.

Svinth, Christian N↗

Datasets for Custom-trained Machine-learning Interatomic Potentials: Nitric Acid Aqueous Solution

This dataset was generated using an iterative active learning strategy with the ArcaNN software package (https://github.com/arcann-chem/arcann_training) to train machine-learning interatomic potentials (MLIPs) for aqueous nitric acid. Each active-learning cycle consisted of three stages: (1) training, (2) exploration, and (3) labeling. The initial training set comprised approximately 800 randomly selected configurations from a previous study by Lewis et al. (https://doi.org/10.1021/jp205510q), which investigated nitric acid solutions at 2, 3, 4, and 5 mol/L. For all configurations, single-point calculations of atomic forces and total energies were performed at the quantum density functional theory BLYP-D2 and PBE-D3 levels of theory using the CP2K Quickstep module. Valence electrons were treated explicitly, while core electrons on all atoms were represented by norm-conserving Goedecker–Teter–Hutter (GTH) pseudopotentials. Long-range dispersion interactions were accounted for using Grimme dispersion corrections. Wave functions were expanded in a mixed Gaussian-and-plane-wave scheme using TZV2P-MOLOPT basis sets for all elements and an 800 Ry auxiliary plane-wave cutoff for the electron density. Self-consistent field convergence was accelerated using orbital transformation and Direct Inversion in the Iterative Subspace, with a convergence threshold of 10^{-6}. All single-point calculations were carried out in periodic orthorhombic cells whose dimensions match those of the molecular configurations sampled from earlier trajectories. The CELL_REF keyword in CP2K was used to define a fixed reference cell, ensuring consistency in the reference data used for MLIP training, particularly when cell fluctuations are present in NpT simulations. The resulting high-fidelity energies and forces constitute the ground-truth labels used to train the MLIPs contained in this dataset.

Dinpajooh, Mohammadhasan [Pacific Northwest Nation↗

APPL Hyperspectral_Imaging_Dataset_for_Heritability_Analysis_in_Populus_trichocarpa

This dataset contains hyperspectral imaging data collected at the Advanced Plant Phenotyping Laboratory (APPL) at Oak Ridge National Laboratory. Natural variants of Populus trichocarpa were imaged using a high-throughput hyperspectral phenotyping pipeline to quantify spectral reflectance traits for downstream quantitative genetics analyses. The dataset includes hyperspectral image files and derived reflectance data products suitable for extracting spectral features across the measured wavelength range (e.g., VNIR and/or SWIR, depending on instrument configuration), along with associated sample metadata (e.g., genotype identifiers, experimental design factors, and imaging run identifiers). These data were generated to support analyses of broad-sense heritability of hyperspectral traits and their relationships with biochemical phenotypes (including lignin traits from Py-MBMS).

APPL↗

Dataset_for_Conserved_macromolecular_architecture_of_Poplar_secondary_cell_walls_revealed_by_ssNMR_and_atomistic_modeling

This dataset contains solid-state 13C NMR data and atomistic molecular dynamics simulation files supporting the study of nanoscale secondary cell wall architecture across 13 genetically diverse Populus trichocarpa genotypes grown under uniform greenhouse conditions in 13C-enriched CO2 atmospheres (~89% 13C enrichment).The dataset contains two collections of solid-state 13C NMR data. (1) 200 MHz data (Bruker Avance III HD, 4 mm HX probe, 10 kHz MAS): raw Bruker TopSpin experiment folders and DMFIT-exported ascii spectra for selective and non-selective 1D 13C-13C spin diffusion experiments (3000 ms mixing) used to quantify inter-polymer spatial proximities, and short-mixing (1 ms) reference spectra used for polymeric abundance quantification by spectral deconvolution. (2) 600 MHz data (Bruker Avance III, 1.6 mm PhoenixNMR HXY probe, 30 kHz MAS): raw Bruker TopSpin experiment folders containing 2D CORD, 2D CP-INADEQUATE, and 13C/1H relaxation (T1, T1rho) experiments for all 13 genotypes, with processed Excel workbooks per experiment type. Molecular dynamics simulation code, coordinate files, and analysis scripts (NAMD/CHARMM/Python) for six atomistic cell wall models are included. Summarized ssNMR data are compiled into a single excel file and subjected to statistical analysis. Multivariate analysis code (PCA, Pearson correlation) and summary data are provided as excel worksheets and Jupyter notebooks (Python 3).

09 BIOMASS FUELS↗

TCR-H: explainable machine learning prediction of T-cell receptor epitope binding on unseen datasets

Artificial-intelligence and machine-learning (AI/ML) approaches to predicting T-cell receptor (TCR)-epitope specificity achieve high performance metrics on test datasets which include sequences that are also part of the training set but fail to generalize to test sets consisting of epitopes and TCRs that are absent from the training set, i.e., are ‘unseen’ during training of the ML model. We present TCR-H, a supervised classification Support Vector Machines model using physicochemical features trained on the largest dataset available to date using only experimentally validated non-binders as negative datapoints. TCR-H exhibits an area under the curve of the receiver-operator characteristic (AUC of ROC) of 0.87 for epitope ‘hard splitting’ (i.e., on test sets with all epitopes unseen during ML training), 0.92 for TCR hard splitting and 0.89 for ‘strict splitting’ in which neither the epitopes nor the TCRs in the test set are seen in the training data. Furthermore, we employ the SHAP (Shapley additive explanations) eXplainable AI (XAI) method for post hoc interrogation to interpret the models trained with different hard splits, shedding light on the key physiochemical features driving model predictions. TCR-H thus represents a significant step towards general applicability and explainability of epitope:TCR specificity prediction.

60 APPLIED LIFE SCIENCES↗

Heterogeneous Multi-Domain Dataset Synthesis to Facilitate Privacy and Risk Assessments in Smart City IoT

The emergence of the Smart Cities paradigm and the rapid expansion and integration of Internet of Things (IoT) technologies within this context have created unprecedented opportunities for high-resolution behavioral analytics, urban optimization, and context-aware services. However, this same proliferation intensifies privacy risks, particularly those arising from cross-modal data linkage across heterogeneous sensing platforms. To address these challenges, this paper introduces a comprehensive, statistically grounded framework for generating synthetic, multimodal IoT datasets tailored to Smart City research. The framework produces behaviorally plausible synthetic data suitable for preliminary privacy risk assessment and as a benchmark for future re-identification studies, as well as for evaluating algorithms in mobility modeling, urban informatics, and privacy-enhancing technologies. As part of our approach, we formalize probabilistic methods for synthesizing three heterogeneous and operationally relevant data streams—cellular mobility traces, payment terminal transaction logs, and Smart Retail nutrition records—capturing the behaviors of a large number of synthetically generated urban residents over a 12-week period. The framework integrates spatially explicit merchant selection using K-Dimensional (KD)-tree nearest-neighbor algorithms, temporally correlated anchor-based mobility simulation reflective of daily urban rhythms, and dietary-constraint filtering to preserve ecological validity in consumption patterns. In total, the system generates approximately 116 million mobility pings, 5.4 million transactions, and 1.9 million itemized purchases, yielding a reproducible benchmark for evaluating multimodal analytics, privacy-preserving computation, and secure IoT data-sharing protocols. To show the validity of this dataset, the underlying distributions of these residents were successfully validated against reported distributions in published research. We present preliminary uniqueness and cross-modal linkage indicators; comprehensive re-identification benchmarking against specific attack algorithms is planned as future work. This framework can be easily adapted to various scenarios of interest in Smart Cities and other IoT applications. By aligning methodological rigor with the operational needs of Smart City ecosystems, this work fills critical gaps in synthetic data generation for privacy-sensitive domains, including intelligent transportation systems, urban health informatics, and next-generation digital commerce infrastructures.

IoT↗

Dark Energy Survey Year 6 Results: Photometric Dataset for Cosmology

We describe the photometric dataset assembled from the full 6 yr of observations by the Dark Energy Survey (DES) in support of static-sky cosmology analyses. DES Y6 Gold is a curated dataset derived from DES Data Release 2 (DR2) that incorporates improved measurement, photometric calibration, object classification and value-added information. Y6 Gold comprises nearly 5000 deg$^{2}$ of grizY imaging in the south Galactic cap and includes 669 million objects with a depth of i$_{AB}$ ∼ 23.4 mag at a signal-to-noise ratio ∼ 10 for extended objects and a top-of-the-atmosphere photometric uniformity <2 mmag. Y6 Gold augments DES DR2 with simultaneous fits to multiepoch photometry for more robust galaxy shapes, colors, and photometric redshift estimates. Y6 Gold features improved morphological star–galaxy classification with an efficiency of 98.6% and a contamination of 0.8% for galaxies with 17.5 < i$_{AB}$ < 22.5. Additionally, it includes per-object quality information, and accompanying maps of the footprint coverage, masked regions, imaging depth, survey conditions, and astrophysical foregrounds that are used for cosmology analyses. After quality selections, benchmark samples contain 448 million galaxies and 120 million stars. This publication is complemented by data access and documentation.

79 ASTRONOMY AND ASTROPHYSICS↗

Validation of the DESI DR2 Ly$\alpha$ BAO analysis using synthetic datasets

The second data release (DR2) of the Dark Energy Spectroscopic Instrument (DESI), containing data from the first three years of observations, doubles the number of Lyman-$\alpha$ (Ly$\alpha$) forest spectra in DR1 and it provides the largest dataset of its kind. To ensure a robust validation of the Baryonic Acoustic Oscillation (BAO) analysis using Ly$\alpha$ forests, we have made significant updates compared to DR1 to both the mocks and the analysis framework used in the validation. In particular, we present CoLoRe-QL, a new set of Ly$\alpha$ mocks that use a quasi-linear input power spectrum to incorporate the non-linear broadening of the BAO peak. We have also increased the number of realisations used in the validation to 400, compared to the 150 realisations used in DR1. Finally, we present a detailed study of the impact of quasar redshift errors on the BAO measurement, and we compare different strategies to mask Damped Lyman-$\alpha$ Absorbers (DLAs) in our spectra. The BAO measurement from the Ly$\alpha$ dataset of DESI DR2 is presented in a companion publication.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

ARM SGP PBLH and MLH datasets from Raman lidar and Doppler lidar

The planetary boundary layer (PBL) plays a critical role in the atmosphere by transferring heat, moisture, and momentum. The warm PBL has a distinct diurnal cycle including the daytime convective mixing layer (ML) and nighttime residual layer developments. Thus, simultaneous determinations of PBL height (PBLH) and ML height (MLH) are necessary for studying PBL characterization and processes. Here, new approaches are developed to provide reliable PBLH and MLH estimates to characterize warm PBL evolution. The approaches use Raman lidar (RL) water vapor mixing ratio (WVMR) and Doppler lidar (DL) vertical velocity measurements at the Southern Great Plains (SGP) atmospheric observatory, which was established by the Atmospheric Radiation Measurement (ARM) User Facility. Compared to widely used lidar aerosol measurements for PBLH, WVMR is a better tracer for PBL vertical mixing. For PBLH, the approach classifies PBL water vapor structures into a few general patterns, then uses a slope method and dynamic threshold method to determine PBLH. For MLH, wavelet analysis is used to reconstruct 2D variance from DL vertical wind velocity measurements according to the turbulence eddy size to minimize the impacts of gravity wave and eddy size on variance calculations; then, a dynamic threshold method is used to determine MLH. Remotely-sensed PBLHs and MLHs are compared with radiosonde measurements based on the Richardson number method. Good agreements between them confirm that the proposed new algorithms are reliable for PBLH and MLH characterization. The algorithms are applied to warm-season RL and ML measurements at the SGP site for five years to study warm-season PBL structure and processes. The weekly composited diurnal evolutions of PBLHs and MLHs in a warm climate were provided to illustrate diurnal and seasonal PBL evolutions. This reliable data set of PBLH and MLH values will be valuable for studying PBL processes, model evolution, and PBL parameterization improvements. The MLH dataset includes the MLH in values of km above ground level. The PBLH dataset includes the PBLH in values of km above ground level, along with a flag ("situation_PBLH") to determine the state of the PBL (1 = Cloudy Condition, 2 = Stable Layer, 3 = Multi-layer WVMR structure, 4 = Well-Mixed PBL, 5 = A de-coupled layer, 6 = Other).

mixing layer height↗

IM3 Data Center Driven Grid Stress Dataset for the U.S. Western Interconnection

This dataset provides projected grid stress and reliability results (including all model inputs and outputs from an open-source grid operations modeling framework - GO), for the Integrated Multisector, Multiscale Modeling (IM3) project, under varying levels of data center demand growth between 2025 and 2035 in the U.S. Western Interconnection. The scenarios and sensitivity experiments are combinations of different data center demand growth rates and energy, weather, population and economic pathways. Data center demand growth projections were sourced from the Electric Power Research Institute (EPRI). The data center demand growth projection names are: Low (3.71% annual data center demand growth) Moderate (5% annual data center demand growth) High (10% annual data center demand growth) Higher (15% annual data center demand growth) Energy, weather, population and economic pathways are informed by two Shared Socioeconomic Pathways (SSP3 and SSP5) and two Representative Concentration Pathways (RCP4.5 and RCP8.5) following the hotter general circulation model (GCM) forcing group from a set of perturbed thermodynamics simulations. The resulting pathway names are: rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 The main scenarios and sensitivity experiments are detailed below. Reference scenario: The projected grid stress and reliability results for the U.S. Western Interconnection from a previous study. This scenario does not consider data center demand growth explicitly. Data center scenario: Building on the reference scenario, this scenario considers various data center growth rates and how they impact the U.S. Western Interconnection. Data center loads are modeled as flat 8760-hr profiles. This scenario does not consider new generation and transmission capacities specifically designed to meet the new data center demands. The related folder is named "flat". Delayed generator retirements sensitivity experiment: Building on the data center scenario, this experiment explores the impact of different levels of natural gas and nuclear generator retirement delays. The resulting scenario names are: (1) postponing 100% nuclear retirements; (2) postponing 100% nuclear and 25% natural gas retirements; (3) postponing only 50% natural gas retirements; (4) postponing 100% nuclear and 50% natural gas retirements; (5) postponing 100% nuclear and 75% natural gas retirements; and (6) postponing 100% nuclear and 100% natural gas retirements. The related folder names are: no_gen_retire_0_gas, no_gen_retire_25_gas, no_gen_retire_50_gas, no_gen_retire_50_gas_only, no_gen_retire_75_gas, and no_gen_retire_100_gas. Demand response through curtailment sensitivity experiment: Building on the data center scenario, this experiment explores the impact of different participation and compensation levels of data center demand response. The resulting scenario names are: (1) 5% demand available for curtailment with 750 $/MWh compensation; (2) 5% demand available for curtailment with 500 $/MWh compensation; (3) 5% demand available for curtailment with 250 $/MWh compensation; (4) 15% demand available for curtailment with 750 $/MWh compensation; (5) 15% demand available for curtailment with 500 $/MWh compensation; and (6) 15% demand available for curtailment with 250 $/MWh compensation. The related folder names are: dr_cost_250_drup_0_drdown_5, dr_cost_250_drup_0_drdown_15, dr_cost_500_drup_0_drdown_5, dr_cost_500_drup_0_drdown_15, dr_cost_750_drup_0_drdown_5, and dr_cost_750_drup_0_drdown_15. Combination of delayed generator retirements and demand response through curtailment sensitivity experiment: The impact of combining postponing 100% nuclear and 25% natural gas retirements with 5% demand available for curtailment with 750 $/MWh compensation is simulated. The related folder is named "dr_cost_750_drup_0_drdown_5_nuc_100_gas_25". Please refer to the README file for a detailed description of the dataset including individual files and references.

Artificial Intelligence↗

Interface between astrophysical datasets and distributed database management systems (DAVID)

This is a status report on the progress of the DAVID (Distributed Access View Integrated Database Management System) project being carried out at Louisiana State University, Baton Rouge, Louisiana. The objective is to implement an interface between Astrophysical datasets and DAVID. Discussed are design details and implementation specifics between DAVID and astrophysical datasets.

Iyengar, S. S.↗

Clementine: Anticipated scientific datasets from the Moon and Geographos

The Clementine spacecraft mission is designed to test the performance of new lightweight and low-power detectors developed at the Lawrence Livermore National Laboratory (LLNL) for the Strategic Defense Initiative Office (SDIO). A secondary objective of the mission is to acquire useful scientific data, principally of the Moon and the near-Earth asteroid Geographos. The spacecraft will be in an elliptical polar orbit about the Moon for about 2 months beginning in February of 1994 and it will fly by Geographos on August 31. Clementine will carry seven detectors each weighing less than about 1 kg: two Star Trackers wide-angle uv/vis wide-angle Short Wavelength IR (SWIR) Long-Wavelength IR (LWIR) and LIDAR (Laser Image Detection And Ranging) narrow-angle imaging and ranging. Additional presentations about the mission detectors and related science issues are in this volume. If fully successful Clementine will return about 3 million lunar images, a dataset with nearly as many bits of data (uncompressed) as the first cycle of Magellan and more than 5000 images of Geographos. The complete and efficient analysis of such large data sets requires systematic processing efforts. Described below are concepts for two such efforts for the Clementine mission: global multispectral imaging of the Moon and videos of the Geographos flyby. Other anticipated datasets for which systematic processing might be desirable include multispectral observations of Earth; LIDAR altimetry of the Moon with high-resolution imaging along each ground track; high-resolution LIDAR color along each lunar ground track which could be used to identify potential titanium-rich deposits at scales of a few meters; and thermal IR imaging along each lunar ground track (including nighttime observations near the poles).

Mcewen, A. S.↗