Search NASA⌕ Search

SEARCH · Search NASA

Results for “geospatial data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

LandScan Mosaic

The LandScan program at Oak Ridge National Laboratory (ORNL), in collaboration with the National Geospatial-Intelligence Agency (NGA), continues to deliver the most accurate and up to date global, high resolution gridded population data. Additionally, the latest advancements in the LandScan HD methodology led to reduced latency in development of rapid updates for geopolitical events. With momentum towards reporting more up to date population estimates, feedback from the user community expressed interest in reporting population estimates in ranges - whether to express a level of uncertainty or confirm to leadership and stakeholders the modeled data are estimates. Building upon the need to understand uncertainty or confidence in the modeled data and report ranges at the global scale, LandScan Mosaic was developed. LandScan Mosaic represents the next generation of high-resolution population modeling, building upon the established success of previous LandScan HD iterations. While LandScan HD employed a deterministic big data fusion approach, LandScan Mosaic enhances this methodology by integrating advanced machine learning techniques to impute missing, yet crucial, population model parameters. This advancement allows for probabilistic modeling of building occupancy and population distribution, incorporating uncertainty quantification through Monte Carlo sampling methods. By combining big data fusion with machine learning-driven imputation and stochastic modeling, LandScan Mosaic provides a more comprehensive and robust representation of population dynamics. LandScan Mosaic will be following the in the footsteps of its longstanding counterpart LandScan Global and releasing a global gridded population raster, at the 3-arcsecond resolution. This technical report documents the current stage of development of LandScan Mosaic, detailing the methodologies and data sources behind the modeling. Stakeholders are encouraged to use this document as an authoritative reference for insight into Mosaic’s data development processes. However, readers should note that LandScan Mosaic remains in a late-stage research and development phase, and methodologies and data presented here are subject to refinements ahead of the anticipated global release in Summer 2025. Feedback and inquiries from users and stakeholders are welcomed as we continue to refine and enhance this important population resource.

97 MATHEMATICS AND COMPUTING↗

Custom surface reflectance, shade mask, and equivalent water thickness maps for the Colorado Headwaters Ecological Spectroscopy Study (2025)

This dataset contains land surface reflectance estimates and additional derived products generated from NEON Imaging Spectrometer (NIS) data collected in the Upper Gunnison river basin during June and July of 2025. Data was collected over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). These products were derived from radiance and LiDAR data collected by the NEON Airborne Observation Platform (AOP) campaign funded by the Colorado Headwaters Ecological Spectroscopy Study (CHESS) (doi:10.15485/3017965). Products include per-pixel surface reflectance (rfl) and reflectance uncertainty (rfl_unc), observational data (obs), canopy equivalent water thickness (ewt), and shade masks. Atmospheric correction was performed per flightline using the ISOFIT (Imaging Spectrometer Optimal FITting) optimal estimation framework to estimate surface reflectance and the associated per-band reflectance uncertainty. Reflectance retrievals achieved a mean absolute error of 1.5% across diverse validation surfaces (see validation report.pdf). Equivalent water thickness was calculated from surface reflectance using the Beer–Lambert absorption of liquid water. Shade masks were generated based on the geometry between the sun angle, ground surface, and sensor at the time of flight. Data products are provided per-flightline and as mosaics for each domain. Flightline data products are provided as ENVI-formatted binary files (rfl, rfl_unc, ewt) and GeoTIFFs (shade). Reflectance and uncertainty mosaics are provided as tiled NetCDFs, while all other mosaicked products are provided as cloud-optimized GeoTIFFs. These formats are supported by common geospatial software (e.g., QGIS, ArcGIS, ENVI) and programmatic libraries in Python (e.g., rasterio, xarray, spectral, netCDF4) and R (e.g., terra, ncdf4). Processing workflows were designed to be equivalent to those used to generate the 2018 CHESS campaign airborne imaging spectroscopy data products (doi:10.15485/3013527). All outputs were co-registered to a common spatial grid to support time series analyses. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgment: Data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). Computational research was carried out at the Jet Propulsion Laboratory, California Institute of Technology, under a contract with the National Aeronautics and Space Administration (80NM0018D0004) and was funded by EMIT Extended Mission Phase E Science.

2018 NEON and 2025 CHESS Campaigns↗

Bias Correcting NOAA's High-Resolution Rapid Refresh (HRRR) Wind Resource Data for Grid Integration Applications [Slides]

Many weather years of high-quality wind data are widely accepted in the grid integration community to be important for studying wind energy technical potential, energy system operations, and grid resilience. NREL makes high-quality wind and solar resource data available. NREL's Grid-Atmosphere workshop (March 2024) identified NREL National Solar Radiation Database as widely used in grid integration modeling, but there is less agreement on commonly used wind datasets. One important factor identified by ESIG's 2023 report 'Weather Dataset Needs for Planning and Analyzing Modern Power Systems' for gold standard wind data is regular updates. To address the need for regular updates, NREL's team can now process all currently available and regularly updated High-Resolution Rapid Refresh (HRRR) outputs. HRRR is an hourly-updated operational forecast product produced by the National Oceanic and Atmospheric Administration (NOAA) (Dowell et al., 2022). One barrier to NREL using HRRR is systematic bias and consistency with NREL's existing wind datasets (e.g. WIND Toolkit, 'WTK') across weather years. To address this barrier, we show that the HRRR can be interpolated and bias-corrected to be consistent with NRE's existing datasets. We call the new dataset BC-HRRR (bias-corrected HRRR). As with historical datasets like the WTK, BC-HRRR is intended for use in grid integration modeling (e.g., capacity expansion, production cost, and resource adequacy modeling). BC-HRRR's (2015-present) consistency with WTK (2007-2013) allows NREL to extend internal grid integration tooling with 15+ weather years of wind data with low-overhead extensibility to future years as they are made available by NOAA. The rest of this slide deck documents the BC-HRRR processing methods, validation, and its implications for intended use.

17 WIND ENERGY↗

Mauka Energy FEVER Tool Dataset

Mauka Energy’s dataset, developed under the Forestry Electric Vehicle Energy Routing (FEVER) project and funded by the U.S. Department of Energy’s Small Business Innovation Research program, is a high-resolution geospatial resource designed to support energy modeling for electric log trucks in complex forestry environments. The dataset integrates detailed spatial and road network data to enable accurate simulation of vehicle performance across varied terrain. At its core, the dataset incorporates lidar-derived elevation models, road alignments, and surface classifications from Oregon State University’s McDonald-Dunn Research Forest. These data capture fine-scale variations in slope, curvature, and surface conditions across forest road systems, allowing for vehicle-level analysis of energy consumption and recovery. The dataset also includes data collected on the surrounding public and private road networks in Benton County, Oregon, used in real-world haul routes. These connecting segments provide critical context for modeling transitions between forest operations and regional transportation infrastructure, incorporating attributes such as grade profiles, elevation change, and speed constraints. This combined dataset underpins the development of Mauka Energy’s rolldown tool, which quantifies energy use and regenerative braking potential on downhill and variable-grade segments. By leveraging high-resolution terrain and road data, the FEVER project enables more accurate assessment of electric vehicle feasibility and performance in forestry applications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Building Fraction and Mean Height for Los Angeles (2010 and 2020) at 30 m Resolution

This dataset provides 30 m resolution spatial layers of building fraction and mean building height for the Los Angeles metropolitan region for 2010 and 2020. Derived from the Model America version 2 building dataset, it represents the distribution of total building footprint area and average building height across the metropolitan region, offering a consistent geospatial resource for urban analysis. The dataset is intended to support applications in urban climate studies, land-use assessment, city-scale modeling, and planning by enabling comparison of urban form characteristics over time. Additional dataset details including data processing details are provided in the Readme document (README_LA_BF_MeanHeight_2010_2020 1.txt).

geospatial↗

Multidimensional perspectives of geo-epidemiology: from interdisciplinary learning and research to cost–benefit oriented decision-making

Research typically promotes two types of outcomes (inventions and discoveries), which induce a virtuous cycle: something suspected or desired (not previously demonstrated) may become known or feasible once a new tool or procedure is invented and, later, the use of this invention may discover new knowledge. Research also promotes the opposite sequence—from new knowledge to new inventions. This bidirectional process is observed in geo-referenced epidemiology—a field that relates to but may also differ from spatial epidemiology. Geo-epidemiology encompasses several theories and technologies that promote inter/transdisciplinary knowledge integration, education, and research in population health. Based on visual examples derived from geo-referenced studies on epidemics and epizootics, this report demonstrates that this field may extract more (geographically related) information than simple spatial analyses, which then supports more effective and/or less costly interventions. Actual (not simulated) bio-geo-temporal interactions (never captured before the emergence of technologies that analyze geo-referenced data, such as geographical information systems) can now address research questions that relate to several fields, such as Network Theory. Thus, a new opportunity arises before us, which exceeds research: it also demands knowledge integration across disciplines as well as novel educational programs which, to be biomedically and socially justified, should demonstrate cost-effectiveness. Grounded on many bio-temporal-georeferenced examples, this report reviews the literature that supports this hypothesis: novel educational programs that focus on geo-referenced epidemic data may help generate cost-effective policies that prevent or control disease dissemination.

59 BASIC BIOLOGICAL SCIENCES↗

Flood Susceptibility Mapping Using Machine Learning and Geospatial-Sentinel-1 SAR Integration for Enhanced Early Warning Systems

This study presents a comprehensive framework for flood susceptibility mapping by integrating geospatial factors with both statistical and machine learning models. Thirteen Flood-related factors, including DEM, slope, TWI, NDVI, etc., are extracted as features of models, and historical flood data derived from Sentinel-1 SAR from 2018 to 2023 are used as the target variables of the models. These datasets are analyzed using a frequency-based statistical model and three machine learning models, including Random Forest, XGBoost, and CNN, to generate flood susceptibility maps. The performance of each model is evaluated through AUC; and SHAP scores are separately generated for Machine learning (ML) models to explain each feature contribution in the ML model. The generated susceptibility maps are validated by high-flood-risk locations monitored by flood sensors, BLE inundation models, and flood-prone areas suggested by the Local Community Task Force. The results indicate that the XGBoost model outperforms all other models, with an AUC of 0.92 and demonstrates the highest alignment with recommended high-flood-risk locations, while the frequency-based statistical model showed the weakest performance with an AUC of 0.65. SHAP value graphs highlight the elevation, slope, and TWI as the most influential features across all models. The susceptibility maps generated by the machine learning model show strong agreement with the BLE map and high-flood-risk areas identified by the local Community Task Force.

Google Engine↗

Powered By reV [Slides]

The reV model empowers users to calculate energy capacity, generation, and cost based on geospatial intersection with grid infrastructure and land-use characteristics. The tool can model a single site up to an entire continent at temporal resolutions ranging from five minutes to hourly, spanning a single year or multiple decades. By automating access to resource data at unprecedented scale, fidelity, and flexibility, the reV model integrates formerly disparate analysis frameworks in the fields of resource modeling, technical potential, and energy cost supply curves.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

OReole-FM: successes and challenges toward billion-parameter foundation models for high-resolution satellite imagery

While the pretraining of Foundation Models (FMs) for remote sensing (RS) imagery is on the rise, models remain restricted to a few hundred million parameters. Scaling models to billions of parameters has been shown to yield unprecedented benefits including emergent abilities, but requires data scaling and computing resources typically not available outside industry R&D labs. In this work, we pair high-performance computing resources including Frontier supercomputer, America's first exascale system, and high-resolution optical RS data to pretrain billion-scale FMs. Our study assesses performance of different pretrained variants of vision Transformers across image classification, semantic segmentation and object detection benchmarks, which highlight the importance of data scaling for effective model scaling. Moreover, we discuss construction of a novel TIU pretraining dataset, model initialization, with data and pretrained models intended for public release. By discussing technical challenges and details often lacking in the related literature, this work is intended to offer best practices to the geospatial community toward efficient training and benchmarking of larger FMs.

Ambrozio Dias, Philipe↗

H3 Geospatial Mapping Resolution Recommendations

This report provides a brief overview of the Hexagonal Hierarchical Geospatial Indexing System (H3) and its benefits and use cases; a comparison of population estimates of various H3 resolutions with administrative boundaries (county, zip code, etc.); and H3 resolution recommendations for electric outage data reporting for US states considering optimal balance between accuracy, computational efficiency, and privacy preservation.

99 GENERAL AND MISCELLANEOUS↗

CERF: IM3 Projected Western US Power Plant Locations

Overview The Capacity Expansion Regional Feasibility (CERF) model is an open-source geospatial python package that provides new power plant locations at a 1km resolution. The model ingests U.S. state or regional-scale electricity system capacity expansion plans, such as those produced by the Global Change Analysis Model (GCAM-USA), and identifies feasible, site-specific locations for individual new power plants (renewable and non-renewable). CERF combines high-resolution geospatial suitability analyses with an economic algorithm that selects individual plant siting locations based on grid interconnection costs and the locational marginal value of new generation. The model incorporates a wide range of dynamic constraints and opportunities, such as protected lands, population density, existing infrastructure, and water availability. This dataset provides CERF power plant siting results for IM3 Phase 2 simulations across eight different scenarios for the Western US through 2055. The scenarios include combinations of two Shared Socioeconomic Pathways (SSP3 and SSP5) with four high-resolution climate projections specific to the United States (see, https://tgw-data.msdlive.org/). These climate projections include "hotter" and "cooler" variants for two Representative Concentration Pathways (RCP4.5 and RCP8.5). The resulting eight simulations are: rcp45cooler_ssp3 rcp45cooler_ssp5 rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85cooler_ssp3 rcp85cooler_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 CERF siting results in this dataset correspond to capacity expansion plans in the GCAM-USA IM3 Phase 2 simulation data and are available for each of the above scenarios. Data Details Temporal Range: 2015-2055 in 5-year timesteps. Note that 2015 is the experiment base year and 2020 and beyond represent model simulation years. Spatial Range: Plant locations are provided for the eleven states in the Western US including Arizona, California, Colorado, Idaho, Montana, New Mexico, Nevada, Oregon, Utah, Washington, and Wyoming. Spatial Resolution: 1 km-squared, provided in x and y coordinates Geospatial Projection: Albers Equal Area Conic (ESRI:102003) File Type: csv The dataset contains subdirectories for each of the eight scenarios described in the overview. Each scenario folder contains two subfolders with the following information: 1. Power Plant Data This directory contains a single .csv file of power plant locations for both pre-existing (non-CERF sited plants in operation in 2015) and new (CERF-sited) power plants across the temporal range along with additional CERF model output parameters for CERF-sited plants. Plant with a siting year earlier than 2020 correspond to facilities that are operational leading into the first timestep CERF simulation. For a more detailed description of CERF model output parameters, see the CERF model documentation. Note that the cerf_plant_id parameter is unique within each scenario file but not across scenario files. Parameter Descriptions scenario - Name of scenario cerf_plant_id - Unique siting identifier cerf_sited - If True, indicates that plant was sited by CERF model. If False, indicates pre-existing facility region_name - Name of region (state) tech_id - Technology ID tech_name - Full generation technology name inclusive of cooling type (if applicable) and additional characteristics tech_simple - Simplified generation technology type unit_size_mw - Power plant unit size (MW) xcoord - X coordinate in the default CRS (meters) ycoord - Y coordinate in the default CRS (meters) index - Index position in the flattend 2D array buffer_in_km - Exclusion buffer around site (km) sited_year - Year of siting retirement_year - Year of retirement lmp_zone - Locational marginal price (LMP) zone ID locational_marginal_price_usd_per_mwh - Locational marginal price ($/MWh) generation_mwh_per_year - Generation output (MWh/yr) operating_cost_usd_per_year - Cost of plant operations ($/yr) net_operational_value - Net operational value based on LMP and and operating costs ($/yr) interconnection_cost - Cost of interconnection for transmission & gas pipeline (if applicable) net_locational_cost -- Difference of interconnection cost and operating value ($/yr) capacity_factor_fraction - Capacity factor (fraction) carbon_capture_rate_fraction - Carbon capture rate (fraction) fuel_co2_content_tons_per_btu - Fuel CO2 content (tons/Btu) fuel_price_usd_per_mmbtu - Fuel price ($/MMBtu) fuel_price_esc_rate_fraction - Fuel price escalation rate (fraction) heat_rate_btu_per_kWh - Heat rate (Btu/kWh) lifetime_yrs - Technology lifetime for annuity (years) operational_life_yrs - Operational lifetime for retirement (years) variable_om_usd_per_mwh - Variable operation and maintenance costs of yearly capacity use ($/MWh) variable_om_esc_rate_fraction - Variable operation and maintenance costs escalation rate (fraction) carbon_tax_usd_per_ton - Carbon tax ($/ton) carbon_tax_esc_rate_fraction - Carbon tax escalation rate (fraction) 2. Storage Data This directory contains information on new and pre-existing energy storage facilities operational in each timestep along with various storage operational parameters. The 2015 timestep provides pre-existing energy storage data and corresponds with facilities that are operational leading into the first model simulation timestep. Note that coordinates in the storage files correspond to the interconnection point on the grid (substation location), not individual energy storage locations. Energy storage is added in a cumulative process at each given interconnection point. That is, each individual file provides the total operational storage capacity interconnected to the specified substation for the given timestep, inclusive of previously installed storage at that location and new storage installed in that timestep at that location. Parameters scenario - Name of scenario timestep - Simulation timestep name - Unique storage identifier s_typ - Type of energy storage technology (battery or pumped storage hydro) s_node - Node ID of interconnecting substation xcoord - X coordinate in the default CRS (meters) ycoord - Y coordinate in the default CRS (meters) charge_rate - Maximum charge rate (power capacity) of storage system (MW) discharge_rate - Maximum discharge rate (power capacity) of storage system (MW) duration - Duration of storage system (hours) max_SoC - Allowed maximum state of charge (energy capacity) of storage system (MWh) min_SoC -Allowed minimum state of charge (energy capacity) of storage system (MWh) charge_eff - Efficiency of charge (fraction between 0 and 1) discharge_eff - Efficiency of discharge (fraction between 0 and 1) Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program.

CERF↗

Techno-Economic Simulation Results Using dGeo for EGS-Based District Heating in the Northeastern United States

This dataset presents the results of techno-economic simulations performed using the Distributed Geothermal Market Demand Model (dGeo) to evaluate the feasibility of Enhanced Geothermal Systems (EGS)-based district heating in the Northeastern United States. Developed by the National Renewable Energy Laboratory (NREL), dGeo is a geospatially resolved, bottom-up modeling framework designed to explore the deployment potential of geothermal distributed energy resources. The dataset, created as part of the Cornell EGS Ground-Truthing Project, provides census tract-level data that includes inputs and outputs such as thermal demand, road length, energy prices, geothermal system sizing, annual energy contributions from geothermal and natural gas peaking boilers, system capital costs (CAPEX), operation and maintenance costs (OPEX), and the levelized cost of heat (LCOH). Key simulation parameters include geothermal gradients, measured well depths, production temperatures, and district heating piping lengths based on S1400 neighborhood road lengths. The simulations assume a target bottom hole temperature of 80C and the development of new district heating networks in each census tract.

15 GEOTHERMAL ENERGY↗

Tool-Based Case Studies on Strategic Deployment of Untapped Micro-Pumped Hydro Storage in Michigan

With most classical hydropower sites already utilized and the global push for rapid integration of renewable energy sources accelerating, there is a critical need to identify alternative energy storage solutions. Pumped hydro energy storage, which accounts for the vast majority of global grid-scale storage, remains one of the most cost-effective and long-duration storage technologies available. Hence, this study presents a novel tool designed to assess the untapped potential of inland lakes and reservoirs for micro-PSH, using Michigan’s relatively flat landscape as a case study due to its extensive but underutilized water infrastructure. To ensure accuracy and reliability, the tool incorporates extensive data gathered from authorized sources, covering more than 420 water facilities and potential reservoirs in the state. The tool evaluates key parameters such as horizontal and vertical distances, volume, and the total storage capacity of each reservoir. Its robust assessment framework integrates these metrics to evaluate each site’s potential. The tool’s intuitive interface and geospatial visualizations support actionable insights for planners and scalable deployment of distributed storage infrastructure.

13 HYDRO ENERGY↗

Geothermal Power Systems Analysis: Outcome of Industry Stakeholders Workshop: Preprint

Geothermal cost and performance evaluation implemented via technoeconomic assessment (TEA) modeling is critical for the Department of Energy (DOE) and other geothermal industry stakeholders in assessing the current state of geothermal technologies and to identify existing hurdles to commercially viable geothermal development. The Geothermal Electricity Technology Evaluation Model (GETEM) is a major TEA tool used in estimating the economic feasibility and levelized cost of energy (LCOE) of conventional hydrothermal systems and enhanced geothermal systems (EGS). Since 2021, GETEM has been transitioning from an intricate spreadsheet model to a user-friendly tool within the System Advisor Model (SAM) developed by the National Renewable Energy Laboratory (NREL). Apart from enabling an expanded visibility of the geothermal model among other renewable resources, having GETEM in SAM has the advantage of simulation automation, better usability, updates tracking, active user inputs/feedback, and extended financial modeling. GETEM is used in developing supply curves for the Annual Technology Baseline (ATB). The ATB data are inputs to the Renewable Energy Potential (reV) and the Regional Energy Deployment System (ReEDS) models. The geothermal module in NREL’s reV model assesses the geothermal energy potential in the conterminous United States by defining the geospatial intersection of geothermal resources with existing grid infrastructure within the constraint of land use characteristics. The ReEDS model is a capacity expansion model used for simulating the long-term build-out and operation of the US generation and transmission system based on current energy costs and policies. To ensure enhanced representation of current industry trends in our model transitions and development, we organized a two-day virtual workshop to elicit geothermal industry stakeholder input and recommendations on our current approaches and assumptions on technoeconomic, resource assessment, and deployment scenarios modeling of geothermal technologies. Participants included developers, operators, investors, regulatory agencies, system modelers, national laboratory researchers, consultants, and other stakeholders. In this workshop, we gained stakeholder insights on current geothermal plant performance (i.e., capacity factors), updated drilling costs and learning curves, and next generation technologies such as closed loop and superhot rock geothermal. Other outcomes from this workshop and its impact on future geothermal development feasibility, resource availability, and capacity expansion studies are compiled and discussed.

Annual Technology Baseline↗

Applying a Multisector Scenario Framework to Evaluate Past and Future Public Surface Water Supply Infrastructure Strategies in Texas

Datasets supporting the index model and scenario analysis used in evaluating surface water supply strategies across different water system types in Texas. These data underpin the scenario development and application of five key indicators: Water Availability Index (WAI), Water Quality Index (WQI), Energy Requirement Index (ERI), Water Treatment Cost (WTC), and Water Infrastructure Cost (WIC). The datasets are organized by system type—stream reaches (flowlines), waterbodies, and reservoirs—and include both raw and standardized index values. The integrated datasets also provide scenario classifications (original and adjusted) based on infrastructure and planning priorities, enabling comparison across Shared Socioeconomic Pathways (SSPs). Additional strategy-level data are included to support evaluation of state-level new reservoir projects in relation to cost and availability tradeoffs. Please refer to the README file provided in Files for more details. Descriptions of the datasets are provided below. Dataset(s) Descriptions Folder: Index_model_database.zip Subfolder: Stream_reach.zip Fl_wf.csv, Fl_wq.csv, Fl_er.csv, Fl_wf_wtcUV.csv, Fl_wf_wtcnoUV.csv, Fl_allfac_wic1.csv, Fl_allfac_wic2.csvDatasets for computing WAI, WQI, ERI, WTC, and WIC for surface water systems classified as stream reaches (flowlines). Subfolder: Waterbody.zip Wb_wf.csv, Wb_wq.csv, Wb_er.csv, Wb_wf_wtcUV.csv, Wb_wf_wtcnoUV.csv, Wb_allfac_wic1.csv, Wb_allfac_wic2.csvEquivalent index model datasets for waterbodies, reflecting hydrologic and infrastructure attributes specific to impounded natural systems. Subfolder: Reservoir.zip Rs_wf.csv, Rs_wq.csv, Rs_er.csv, Rs_wf_wtcUV.csv, Rs_wf_wtcnoUV.csv, Rs_allfac_wic1.csv, Rs_allfac_wic2.csvIndex model datasets specific to regulated reservoir systems, incorporating both resource indicators and cost parameters. Folder: Integrated data.zip combined_merged_data.csv, combined_merged_data_scenario.csvDatasets integrating index model indicators (both raw and scaled) with scenario classifications, including adjustments reflecting SSP-aligned transitions and planning shifts. Folder: Additional data.zip wai_supplystrat_wic_merged.csvCurated dataset capturing proposed major reservoir-based municipal water supply strategies in Texas. Integrates site-level planning data with estimated capital infrastructure costs and water availability scores for comparative assessment.

geospatial↗

Renewable Energy Potential Model: Hawaii Geothermal Supply Curves

This dataset extends the development of the Renewable Energy Potential (reV) model to include geothermal energy, with a specific focus on Hawaii. Provided here are the results of two scenarios that were modeled for geothermal energy in Hawaii: binary enhanced geothermal systems (EGS) at a depth of 2.5 km and hydrothermal binary systems at a depth of 1.5 km. The resource data for both scenarios were derived from Lautze and Haskins (2024) using an exponential method. The PFA probability of heat map was used as a look up table for which temperature gradient to use (Lautze and Haskins, 2024). The dataset provides geospatial and techno-economic details for evaluating geothermal energy potential. It includes spatial coordinates, estimated capacity factors, developable area, resource potential, and annual energy production metrics. Economic details such as levelized cost of electricity (LCOE), site development costs, transmission costs, and fixed-charge rates are also included. The reV model, originally developed for wind and solar energy, incorporates these variables to evaluate deployment constraints related to land use, environmental and cultural factors, and grid integration.

15 GEOTHERMAL ENERGY↗

Existing Hydropower Assets (EHA) Capacity Plant Database, 2005-2024

Existing Hydropower Asset (EHA) Annual Capacity is a geospatial point-level dataset containing annual capacity over the years (2005-2024) and key characteristics of operational U.S. hydropower plants with 1 megawatt or greater of nameplate capacity. EIA form 860 and EHA are the primary sources of the derived data.

Johnson, Megan [ORNL] (ORCID:0000000290141741)↗