Search NASA⌕ Search

SEARCH · Search NASA

Results for “Dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

220EV Box Truck Telemetry Dataset

This dataset contains telemetry data from the 220EV box truck as provided by the manufacturer, Dana. Data are recorded in a time series interval of 1 second. Recorded data contain odometer reading (kilometers), axle speed, state of charge of the battery, and cabin heater setting.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Detailed Simulation Datasets Quantifying U.S. DOE VTO/HFTO R&D Benefits Across Light- to Heavy-Duty Vehicles

For more than 20 years, Argonne National Laboratory’s Vehicle & Mobility Systems Department has assessed how R&D investments by the U.S. Department of Energy’s Transportation Technologies Office and Alternative Fuels and Feedstocks Office affect vehicle energy use and cost. The analyses are performed using Autonomie, Argonne’s full-vehicle simulation tool for energy consumption, performance, and cost. The study covers five time frames ranging from present day through 2050, with more than 30 vehicle classes and applications (10 light duty and >20 medium and heavy duty), as well as six powertrain configurations (conventional, start-stop, hybrid electric vehicle, plug-in hybrid electric vehicle, battery-electric vehicle, and fuel cell electric vehicle) and five fuels (gasoline, diesel, natural gas, hydrogen, and electricity). Low and high technology uncertainty scenarios have been considered to capture a realistic range of outcomes. The resulting datasets include the assumptions used (i.e., efficiency, $/kWh), vehicle-level data (power, energy, weight, and cost), and outputs such as energy consumption, manufacturer’s suggested retail price, and total cost of ownership. These data are critical to stakeholders working in transportation, technology assessment, and long-term R&D planning.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Autonomie Simulation Datasets in Support of U.S. DOT-NHTSA Advanced Vehicle Technology Research

Understanding how new vehicle technologies affect fuel economy and energy use is critical to the regulatory work performed by the U.S. Department of Transportation’s National Highway Traffic Safety Administration (NHTSA), which sets Corporate Average Fuel Economy (CAFE) standards under the Energy Policy and Conservation Act of 1975. In order to support this work, Argonne National Laboratory uses Autonomie, a full-vehicle simulation tool, to evaluate advanced powertrain architectures and their effects on vehicle energy consumption and performance. A wide range of vehicle classes has been assessed (i.e., internal combustion engine vehicles, hybrid electric vehicles, plug-in hybrid electric vehicles, battery-electric vehicles, and fuel cell electric vehicles), as well as the effects of various technology improvements such as lightweighting, aerodynamic refinements, and low-rolling-resistance tires. Simulations have been run across multiple drive cycles to capture fuel and electricity use under realistic operating conditions. The resulting datasets include detailed vehicle-level results, model assumptions, and validation reports, all of which have been made publicly available through NHTSA in support of the 2023 notice of proposed rulemaking covering light-duty vehicles for model years 2027 to 2035. These data are critical to stakeholders working in fuel economy regulation, vehicle technology assessment, and energy policy analysis.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Dataset: "Widespread Drought-driven Declines in Streamflows and Water quality in the Upper Colorado River Basin (1998-2022)"

This data package contains the associated data and scripts for Nagamoto, E., Ombadi, M., Ciulla, F. et al. Widespread drought-driven declines in streamflows and water quality in the Upper Colorado River Basin during 1998-2022. Commun Earth Environ 7, 734 (2026). https://doi.org/10.1038/s43247-026-03890-5. This purpose of this study was to investigate the impact of the 21st century drought on water quantity and quality at catchments throughout the Upper Colorado River Basin (UCRB). We used stream flow, water temperature, specific conductance, air temperature, precipitation, and catchment attribute data for over 200 sites in the UCRB, collected from the National Water Information System using Basin3D (Varadharajan, 2023), GAGESII (Falcone, 2010), and the Google Earth Engine. We identified years of severe drought between 1998 and 2022 using the Standardized Precipitation Evaporation Index (SPEI), then calculated the relative change percentage of the stream flow, water temperature, and specific conductance from drought versus non-drought years. We used the attribute information from GAGESII to investigate what physical traits of catchments are associated streamflow vulnerability (greater relative change) or resilience to drought. We used land cover data from the National Land Cover Database (USGS, 2024) to assess any changes to physical attributes that may not be represented in the static attributes information in GAGESII. To increase data availability, we modeled stream temperature using methods from Willard, 2023. While the study period is water years 1998 to 2022, the raw water quantity and quality data extends to 1950 and the meteorological data extends to 1980. The data and code can be downloaded via the UCRB_drought.zip. Within the zip, the files are organized as follows: - INPUTS: Contains all input data used in UCRB_Drought_Workflow.ipynb - OUTPUTS: Contains all intermediate data created from UCRB_Drought_Workflow.ipynb as well as final products including the calculated Standardized Evapotranspiration Index (SPEI) - climatic_variables: The code used to collect meteorologic data from Google Earth Engine - feature_importance: The code used for the catchment attributes analysis - preprocessing: Code used in UCRB_Drought_Workflow_Preprocessing.ipynb - pyeto: Code used in UCRB_Drought_Workflow_Preprocessing.ipynb - calculations: Code used in UCRB_Drought_Workflow_Impacts.ipynb - plotting: Code used in UCRB_Drought_Workflow_Impacts.ipynb - README.md - UCRB_Drought_Workflow_Preprocessing.ipynb: The code used to prep raw data for the analysis - UCRB_Drought_Workflow_Impact.ipynb: The code which uses the prepped raw data for analysis, and plots all figures - requirements_ucrb-drought_v2.yml: The requirements file to create a virtual environment and Jupyter Lab kernel to run the code The INPUTS folder is organized into the following major directories and sub-directories. The "RDC_WT_SC_RAW" folder contains raw data for streamflow, water temperature, and specific conductance in a ".h5" file. The "NLCD_RAW" folder contains ".csv" files with annual land cover percentages for counties within the UCRB. The "MET_RAW" folder contains a ".csv" file with monthly meteorological data (air temperature and precipitation) for the sites in the UCRB which was obtained from code in the climatic_variables folder. The "GAGESII" folder contains ".csv" files with physical catchment attribute variables for catchments across the country. The "WT_LSTM_data" folder contains ".csv" files with calculated WT (Willard, 2023) and the associated RMSEs. The "Upper_Colorado_River_Basin_Boundary" folder contains geographic data including a shapefile for plotting in the UCRB_Drought_Workflow.ipynb. The "RESERVOIRS_RAW" folder contains ".csv" files for each reservoir in the UCRB with daily reservoir storage. There are also two files in the INPUTS folder that have combined reservoir storage data and reservoir metadata. The OUTPUTS folder is organized into the following major directories and sub-directories. The "RDC_WT_SC_data" folder contains a folder "Water_year" with the associated cleaned data, metadata, and data availability information in ".csv" files, a folder "Median_Relchange" with the relative change comparing drought to non-drought years in ".csv" files, and a folder "Peak95_Min5_Relchange" that has ".csv" files for the relative change in peak (95th %) and minimum (5th %) variables. The "NLCD_data" folder contains the difference in land cover from the beginning to end of the study period and the percentage of the county that is within UCRB bounds can be found in Nagamoto et al (2025)). The "MET_data" folder contains separated monthly air temperature and precipitation data and the calculated PET in ".csv" files. The "SPEI_data" folder contains ".csv" files with calculated SPEI values (one restricted to the study period and the other with information from the entire MET data period). The "Paper_Tables" folder contains two ".csv" files containing site information and data availability and information about the GAGESII trait aggregated categories. The base directory includes the file “flmd.csv” for a list and description of all files and the file “dd.csv” for data dictionaries. Scripts for preprocessing, analysis, and figure generation are located in the associated GitHub repository found at [https://github.com/iNAIADS/drought-impacts/tree/develop/UCRB-drought]. UPDATE 1: Title and code file updated to match submitted manuscript 10-15-2025. UPDATE 2: Code and data files updated to match revised manuscript 3-4-2026. UPDATE 3: Code and data files updated to match revised manuscript 6-7-2026. ** NOTE: DD and FLMD have not been updated yet. UPDATE 4: Added associated Manuscript information and DD and FLMD have been updated. To cite this code, please use the following BibTeX: @misc{nagamoto2025drought, author = {Emily Nagamoto and Fabio Ciulla and Mohammad Ombadi and Jared Willard and Rosemary Carroll and Charuleka Varadharajan}, title = {Dataset: "Widespread Drought-driven Declines in Streamflows and Water quality in the Upper Colorado River Basin (1998-2022)"}, year = {2025}, doi = {10.15485/2551894}, publisher = {ESS-DIVE Repository}, url = {https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2551894} }

54 ENVIRONMENTAL SCIENCES↗

Historic climate, cosmogenic 10Be, denudation-rate, and geospatial datasets from the Pikes Peak region, Colorado, USA

This data package contains geographic information system (GIS) layers and tabular datasets associated with the study of elevation-dependent denudation rates on Pikes Peak in the Front Range of the Rocky Mountains, Colorado, USA. The package includes GIS layers used to produce the study-area map, including sample locations, sample watershed boundaries, the Pikes Peak batholith, Pleistocene glacier extent, weather station locations, and elevation and hillshade rasters, together with comma-separated value (CSV) tables and matching CSV data dictionaries. These mapped layers provide the geographic framework for interpreting denudation patterns across the Pikes Peak region and for relating sample locations to watershed geometry, bedrock setting, glacial history, and nearby climate stations. The first group of tables reports climate and geospatial context for the study area. These files include station-based temperature and precipitation data used to characterize elevational gradients in mean annual climate and monthly climate seasonality, sample locations, denudation-rate and topographic metrics, fixed frost-cracking model parameters, frost-cracking intensity and precipitation-frequency metrics, and stream-power inversion results. Together, these data provide the basis for evaluating how denudation varies with elevation, climate, and landscape form across sampled catchments on Pikes Peak. The second group of tables reports cosmogenic nuclide and erosion-model results used in the denudation analysis. Included files contain accelerator mass spectrometry (AMS) measurements for in situ-produced cosmogenic beryllium-10 (10Be), including sample identifiers, measured 10Be:9Be ratios, analytical uncertainties, carrier mass, quartz mass, blank corrections, blank-group statistics, and calculated 10Be concentrations and uncertainties. Additional tables summarize stream-power-law inversion results for sampled catchments, including optimized model parameters, predicted erosion rates, residual metrics, channel-pixel counts, and convergence status, as well as regression equations and summary statistics used to evaluate relationships among elevation, climate, frost cracking, precipitation forcing, and denudation rate. The package contains GIS files, comma-separated value files (.csv), Microsoft Excel files (.xlsx), CSV data dictionaries, a file-level metadata table, and a readme text file.

10Be cosmogenic nuclides↗

Gold-Thiolate Nanocluster Dynamics Dataset and Supplementary Files

Supplementary information accompanying the article “Gold-Thiolate Nanocluster Dynamics and Intercluster Reactions Enabled by a Machine Learned Interatomic Potential” including the training/testing dataset, potential files, and simulation results.

36 MATERIALS SCIENCE↗

CaloFlow for CaloChallenge dataset 1

CALOFLOW is a new and promising approach to fast calorimeter simulation based on normalizing flows. Applying CALOFLOW to the photon and charged pion ≥ant showers of Dataset 1 of the Fast Calorimeter Simulation Challenge 2022, we show how it can produce high-fidelity samples with a sampling time that is several orders of magnitude faster than ≥ant. We demonstrate the fidelity of the samples using calorimeter shower images, histograms of high level features, and aggregate metrics such as a classifier trained to distinguish CALOFLOW from ≥ant samples.

Physics↗

Understanding Event Trajectories Across Massive Temporal Datasets with Word Embeddings and Visualization

In collaboration with researchers from Virginia Tech, Savannah River National Laboratory has continued development of a natural language processing pipeline to identify and extract events of interest from massive open data sources in the domain of worldwide state-sponsored civil nuclear energy. The foundation of the pipeline is built on compass aligned temporal word embedding models, whereby contextual shifts are automatically identified by comparing keyword embedding vectors across successive time windows. Within the approach, a contextual shift indicates the occurrence of a potential event of interest. However, in such a broad topical domain that captures events at a global scale, across various life cycle stages, and across numerous different technology types, a user that is monitoring events may have broad interests in capturing many different event types with varying degrees of signal. As such, the quantity of information that may be returned from an automated event extraction pipeline can be substantial, requiring manual effort to sift through the information to identify any relevant bits of information. Therefore, a more streamlined workflow that aids in directing a user toward specific information at different points in time is necessary. The workflow presented here has been developed with this concept in mind, built on top of the initial prototype event extraction pipeline, whereby a user can analyze temporal text-based data sources at multiple different contextual levels to isolate key points in time and key subdomains captured within a data corpus. Using multiple corpuses that consist of approximately 7 million Tweets and 7 million news articles, the team has extended compass aligned temporal word embedding models to establish an interconnected and hierarchical structure that relates known key words of interest to documents, local topics (i.e., within a time window), and global topics across the corpuses. All of this information is packaged into a visual analytics system that is linked to the information extraction pipeline and enables a user to identify contextual information that describes the evolution of a high dimensional embedding space across time to isolate changes of interest and explore associated events. This report demonstrates the use of these analytics and a means to fuse information across multiple datasets.

97 MATHEMATICS AND COMPUTING↗

Deep Learning-based Surrogate Model for Efficient Reservoir Simulation in Large-scale Geological Carbon Storage: Application in IBDP Dataset

This project introduces an advanced deep learning (DL)-based surrogate modeling approach to enhance the efficiency and accuracy of large-scale geological carbon storage (GCS) simulations. Using the Illinois Basin Decatur Project (IBDP) dataset as training data, the study employs a residual U-Net architecture to predict critical state variables such as pressure and CO₂ saturation, as well as CO₂ plume migration. By incorporating key geological parameters (e.g., porosity, permeability, and rock facies) and physics-informed inputs like the diffusive time of flight and time step, the DL model effectively reduces computational complexity while maintaining robust physical constraints. Compared to traditional simulators like Eclipse, the DL model achieves remarkable accuracy, with a root mean square error (RMSE) of 1.57 psi for pressure and 0.007 for saturation, and dramatically reduces computational time from hours to just 69.9 seconds for 50-step simulations. These results demonstrate the potential of innovative DL methodologies to improve the predictivity and operational efficiency of GCS simulations, providing a reliable foundation for decision-making in CCS operations. Supported by the SMART initiative, this project underscores the success of leveraging computational innovations to advance CCS technologies.

advanced deep learning↗

Multi-fidelity equations of state and transport coefficient datasets for pulsed-power applications

Reliably simulating experiments relevant to the National Nuclear Security Administration (NNSA) requires a detailed description of material properties across a wide range of conditions. Such properties include the equations of state, charged-particle transport coefficients, and optical properties like the opacity. Together, these properties make up the material models used in radiation-magnetohydrodynamic simulations of nuclear fusion experiments. Many of these models do not incorporate uncertainties in the data used to produce them. It is unknown whether these uncertainties significantly impact the interpretation of simulation results and diagnostics. The purpose of this work is to quantify how such uncertainties impact simulations of pulsed-power experiments. We accomplished this task by first assessing discrepancies between approaches used to generate the data. This included bringing together members of the high-energy-density community spanning the three NNSA laboratories and multiple universities. Then, using these data, we developed a general framework that systematically incorporates physical uncertainties within the material models suitable for uncertainty quantification analyses. The framework utilizes machine learning, Bayesian inference, and incorporates multi-fidelity datasets. We demonstrated the framework by quantifying the impact that material model uncertainties have on simulations of pulsed-power experiments underway on Z at Sandia National Laboratories. As a result of this work, we discovered that modest uncertainties in material models (roughly 20%) correspond to significant uncertainties in the outputs from simulations. Our framework has enabled rapid construction of material models through an automated procedure and allows for the generation of material models of interest to the NNSA.

36 MATERIALS SCIENCE↗

Building Datasets and Training Methods for ML Based Magnet Quench Detection

Detecting quenches in superconducting (SC) magnets during training is a challenging process that involves capturing physical events that occur at different frequencies and appear as various signal features. These events may be correlated across instrumentation type, thermal cycle, and ramp. These events together build a more complete picture of continuous processes occurring in the magnet, and may allow us to flag potential precursors for quench detection. We present our work on building an automatic machine learning (ML) based quench detection system. We build upon our existing work on unsupervised auto-encoders for acoustic sensors and quench antenna (QA) by first establishing a supervised ML training pipeline. We show the results of an event tagging, analysis, and simulation framework on our QA and acoustic data which are used concurrently to build a training dataset for a supervised implementation. We then show how this supervised training can be used as a prior in a semi-supervised framework and compare this to the unsupervised neural network auto-encoder performance.This allows us to have a more concrete understanding of the performance of our algorithms relative to physical events occurring in the magnet, and also provides a baseline software tool to generically evaluate our quench prediction autoencoders under completely unsupervised, supervised, and semi-supervised training conditions.

Khan, Maira [Fermilab]↗

Bridging the Gap Between Astronomical Datasets: From Proof-of-Concept to AI Model Deployment with Domain Adaptation

Artificial Intelligence is transforming astrophysics, from studying stars and galaxies to analyzing cosmic large-scale structures. However, a critical challenge arises when AI models trained on simulations or past observational data are applied to new observation— leading to domain shifts, reduced robustness, and increased uncertainty of model predictions. This talk will explore these issues, highlighting examples such as galaxy morphology classification and cosmological parameter inference, where AI struggles to adapt across different datasets. We will discuss domain adaptation as a strategy to improve model generalization and mitigate biases—essential for making AI-driven discoveries reliable. Notably, these challenges extend beyond astrophysics, affecting AI applications across physics and other scientific domains. Addressing them is essential for maximizing AI’s impact in advancing scientific research.

Ćiprijanović, Aleksandra [Fermilab]↗

Hourly Load Profile Dataset for Electric Airport Ground Support Equipment in the United States

Currently, there are limited data on the magnitude and timing of electricity demand from electric ground support equipment (eGSE) across U.S. airports. To address this gap, this study presents a modeling approach for estimating hourly annual electricity demand from eGSE at the 50 largest U.S. commercial airports. These datasets, accessible at data.nrel.gov/submissions/279, provide critical insights into the potential grid impacts and electricity demand associated with eGSE adoption.

33 ADVANCED PROPULSION SYSTEMS↗

Hourly Load Profile Dataset for Electric Transit Bus Depots in the United States

Transit buses operate primarily in dense urban areas, where nearby populations face increased exposure to fine particulates, nitrogen oxides, and other harmful pollutants. Electrifying transit buses presents a clear opportunity to reduce greenhouse gas emissions and improve urban air quality. However, widespread adoption may pose significant energy and infrastructure challenges, which can be mitigated through proactive planning and investment. This report presents a robust modeling framework and an initial estimation of the hourly electricity demand at transit bus depots across the United States. The resulting depot-level dataset, available at data.nrel.gov/submissions/282, provides valuable insights for infrastructure planning and electricity demand forecasting, supporting the scalable electrification of transit bus fleets nationwide.

33 ADVANCED PROPULSION SYSTEMS↗

Hourly Load Profile Dataset for Electric Port Cargo Handling Equipment in the United States

Historically, ports have relied on fossil fuels, particularly diesel, as their primary energy source. Transitioning to electric cargo handling equipment (eCHE) offers a promising solution, as this relatively mature technology eliminates tailpipe emissions, reduces harmful airborne particulates, lowers noise pollution, and supports decarbonization. This study develops an initial estimation of hourly electricity demand for eCHE at the top 25 container ports (by tonnage) in the United States. These datasets, accessible at data.nrel.gov/submissions/281, provide valuable insights into the electricity demand patterns and potential grid impacts associated with widespread eCHE adoption, forming a foundation for future refinement based on stakeholder feedback.

33 ADVANCED PROPULSION SYSTEMS↗

A Scalable Hardware-and-Human-in-the-Loop Grid-interactive Efficient Building Equipment Performance Dataset

This project developed a publicly available, high-fidelity dataset about the interactions among humans, homes, and heat pumps supporting grid interactive efficient buildings to balance demand on the grid with comfort for occupants. Laboratory measurements and simulations of the hardware capture the second-scale electric power dynamics of heat pumps providing grid services like load shifting and load shedding. Field measurements, behavior tracking, and qualitative surveys of people in their homes over multiple years—including experimentally adjusting the heating and cooling system to provide grid services—to capture the reciprocal effect of human behavior on grid services, and grid services on human comfort. Taken together, these data capture the complete Hardware and Human in the loop system for residential heat pumps, reducing large uncertainties in simulation for design, and models for control of heat pumps, and grid-interactive buildings.

24 POWER TRANSMISSION AND DISTRIBUTION↗