SEARCH · Search NASA
Results for “Data Science Model”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Kamodo: Simplifying Model Data Access and Utilization
To address the lack of user-friendly software needed to simplify the utilization of model data across Heliophysics, the Community Coordinated Modeling Center (CCMC) at NASA’s Goddard Space Flight Center has developed a model-agnostic method via Kamodo for users to easily access and utilize model data in their workflows. By abstracting away the broad range of file formats and the intricacies of interpolation on specialized grids, this approach significantly lowers the barrier to model data access and utilization for the community while adding exciting new capabilities to their tool boxes. This paper describes the direct interfaces to the model data, called model readers, and a basic introduction on how to use them. Additionally, we detail the planned approach for including custom interpolation codes, and include current progress on specialized visualization developments. The CCMC is maintaining Kamodo as an official NASA open-sourced software to enable and encourage community collaboration.
Reanalysis of Rat Data from Spacelab Life Sciences 2 (SLS-2) to Reveal Research Gaps in Spaceflight Data
Using and analyzing the legacy data obtained in space life sciences missions has the potential to provide researchers a complete picture of the molecular changes associated with space without further experimentation. This project’s objective is to extract, filter, organize, and analyze all Rattus norvegicus data and metadata obtained from Columbia’s Spacelab Life Sciences 2 (SLS-2, STS-58) mission to explore the ways that we can compile information from model organisms, in our case rats, to create a reliable model to understand biological mechanisms in response to these space flight changes. By reusing rare space legacy data coupled with data analysis techniques, we can combine individual preexisting datasets with current ones to gain new, comprehensive insights about the effects of spaceflight on our bodies. Our methods can also lead to the creation of a standardized pipeline that could be applied to other space life science datasets for analysis. In this review, every biological experiment conducted on rats in the SLS-2 Mission was studied with our pipeline to create a new biological library and model that could be used by scientists from around the world to make novel discoveries and develop new hypotheses from this priceless information without the limitation of the costs of spaceflight experimentation.
Model scripts associated with “Revisiting controls on hyporheic respiration with knowledge-guided machine learning at continental scale”
NOTE: The manuscript associated with this data package is currently in review. The data/scripts may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final scripts and additional metadata. This data package is associated with the publication “Revisiting controls on hyporheic respiration with knowledge-guided machine learning at continental scale” submitted to Environmental Science & Technology (Zheng et al. 2026). The project combines mechanistic process modeling with knowledge-guided machine learning (KGML) to evaluate how organic matter chemistry, microbial biomass, and physical substrate accessibility regulate realized respiration rates across river corridors. All data used in this paper have been previously published and can be accessed at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1729719 (Goldman et al., 2020). This data package contains 3 R-markdown (Rmd) preprocessing scripts for the previously published data and subsequent modelling workflows. The full workflow with input and output data can be found in the associated GitHub repository at https://github.com/jianqiuz/KGML-WHONDRS.
Angular Correlation Date Measurements with the GeRMAC system
Advanced modeling and simulation efforts have improved at Idaho National Laboratory in recent years with a solid foundation of experimental results. Current computational methods represent significant modeling capabilities but are limited by the accuracy and availability of nuclear data. The creation of pre- and post-processing software tools to address these limitations is fundamental to the improvement of nuclear science modeling capacities. One aspect of predictive modeling tools deals with gamma-rays emitted from radionuclides, including fissile or fissionable material, fission products, or activation products, produced in reactor experiments or other neutron environments. The resulting radionuclides decay in unique ways, providing complications upon measurement as a result of random and cascade, or true, coincidence summing. These effects are not easily quantified during modeling efforts of gamma-ray source terms., The germanium rotational measurements for angular correlation (GeRMAC) system was built to quantify the relative angles for gamma rays emitted by radionuclides of interest to investigate true coincidence, or cascade, summing as well as the nuclear energy levels of decay schemes of interest. Proof of concept studies utilize a series of laboratory check sources to provide validity, and it will soon be used to perform the same measurements for fission products of interest. The resulting data can be used to implement into a Monte Carlo code, such as Geant4, to provide more precise gamma-ray source terms following irradiations of materials.
CERES Monthly TOA and SRB Averages (SRBAVG) data in HDF-EOS Grid (CER_SRBAVG_TRMM-PFM-VIRS_Edition2B)
The Monthly TOA/Surface Averages (SRBAVG) product contains a month of space and time averaged Clouds and the Earth's Radiant Energy System (CERES) data for a single scanner instrument. The SRBAVG is also produced for combinations of scanner instruments. The monthly average regional flux is estimated using diurnal models and the 1-degree regional fluxes at the hour of observation from the CERES SFC product. A second set of monthly average fluxes are estimated using concurrent diurnal information from geostationary satellites. These fluxes are given for both clear-sky and total-sky scenes and are spatially averaged from 1-degree regions to 1-degree zonal averages and a global average. For each region, the SRBAVG also contains hourly average fluxes for the month and an overall monthly average. The cloud properties from SFC are column averaged and are included on the SRBAVG. [Location=GLOBAL] [Temporal_Coverage: Start_Date=1998-02-01; Stop_Date=2000-03-31] [Spatial_Coverage: Southernmost_Latitude=-90; Northernmost_Latitude=90; Westernmost_Longitude=-180; Easternmost_Longitude=180] [Data_Resolution: Latitude_Resolution=1 degree; Longitude_Resolution=1 degree; Horizontal_Resolution_Range=100 km - < 250 km or approximately 1 degree - < 2.5 degrees; Temporal_Resolution=1 month; Temporal_Resolution_Range=Monthly - < Annual].
Modern chemical graph theory
Abstract Graph theory has a long history in chemistry. Yet as the breadth and variety of chemical data is rapidly changing, so too do graph encoding methods and analyses that yield qualitative and quantitative insights. Using illustrative cases within a basic mathematical framework, we showcase modern chemical graph theory's utility in Chemists' analysis and model development toolkit. The encoding of both experimental and simulation data is discussed at various levels of granularity of information. This is followed by a discussion of the two major classes of graph theoretical analyses: identifying connectivity patterns and partitioning methods. Measures, metrics, descriptors, and topological indices are then introduced with an emphasis upon enhancing interpretability and incorporation into physical models. Challenging data cases are described that include strategies for studying time dependence. Throughout, we incorporate recent advancements in computer science and applied mathematics that are propelling chemical graph theory into new domains of chemical study. This article is categorized under: Molecular and Statistical Mechanics > Molecular Dynamics and Monte‐Carlo Methods Structure and Mechanism > Computational Materials Science Structure and Mechanism > Molecular Structures
GBaTSv2: a revised synthesis of the likely basal thermal state of the Greenland Ice Sheet
The basal thermal state (frozen or thawed) of the Greenland Ice Sheet is under-constrained due to few direct measurements, yet knowledge of this state is becoming increasingly important to interpret modern changes in ice flow. The first synthesis of this state relied on inferences from widespread airborne and satellite observations and numerical models, for which most of the underlying datasets have since been updated. Further, new and independent constraints on the basal thermal state have been developed from analysis of basal and englacial reflections observed by airborne radar sounding. Here we synthesize constraints on the Greenland Ice Sheet's basal thermal state from boreholes, thermomechanical ice-flow models that participated in the Ice Sheet Model Intercomparison Project for CMIP6 (ISMIP6; Coupled Model Intercomparison Project Phase 6), IceBridge BedMachine Greenland v4 bed topography, Making Earth Science Data Records for Use in Research Environments (MEaSUREs) Multi-Year Greenland Ice Sheet Velocity Mosaic v1 and multiple inferences of a thawed bed from airborne radar sounding. Most constraints can only identify where the bed is likely thawed rather than where it is frozen. This revised synthesis of the Greenland likely Basal Thermal State version 2 (GBaTSv2) indicates that 33 % of the ice sheet's bed is likely thawed, 40 % is likely frozen and the remainder (28 %) is too uncertain to specify. The spatial pattern of GBaTSv2 is broadly similar to the previous synthesis, including a scalloped frozen core and thawed outlet-glacier systems. Although the likely basal thermal state of nearly half (46 %) of the ice sheet changed designation, the assigned state changed from likely frozen to likely thawed (or vice versa) for less than 6 % of the ice sheet. This revised synthesis suggests that more of northern Greenland is likely thawed at its bed and conversely that more of southern Greenland is likely frozen, both of which influence interpretation of the ice sheet's present subglacial hydrology and models of its future evolution. The GBaTSv2 dataset, including both code that performed the analysis and the resulting datasets, is freely available at https://doi.org/10.5281/zenodo.6759384 (MacGregor, 2022).
Increasing the Reproducibility and Replicability of Supervised AI/ML in the Earth Systems Science by Leveraging Social Science Methods
Artificial intelligence (AI) and machine learning (ML) pose a challenge for achieving science that is both reproducible and replicable. The challenge is compounded in supervised models that depend on manually labeled training data, as they introduce additional decision-making and processes that require thorough documentation and reporting. We address these limitations by providing an approach to hand labeling training data for supervised ML that integrates quantitative content analysis (QCA)—a method from social science research. The QCA approach provides a rigorous and well-documented hand labeling procedure to improve the replicability and reproducibility of supervised ML applications in Earth systems science (ESS), as well as the ability to evaluate them. Specifically, the approach requires (a) the articulation and documentation of the exact decision-making process used for assigning hand labels in a “codebook” and (b) an empirical evaluation of the reliability” of the hand labelers. In this paper, we outline the contributions of QCA to the field, along with an overview of the general approach. We then provide a case study to further demonstrate how this framework has and can be applied when developing supervised ML models for applications in ESS. With this approach, we provide an actionable path forward for addressing ethical considerations and goals outlined by recent AGU work on ML ethics in ESS.
Package Data for CERF-Data Centers
This dataset contains sample input 100m resolution raster files for running the CERF-DC python package (see https://github.com/IMMM-SFA/cerf_data_centers) at the state level across the CONUS. Due to data availability constraints, some of the items included in this dataset are proxies or assumptions for siting factors used in the model. These are individually noted in the item descriptions and can be exchanged with more detailed information upon availability. Data Descriptions The following raster files are included in the data download: state_siting_region.tif — State areas identified by state FIPS code composite_siting_suitability.tif — Value of 1 indicates suitable siting location, 0 otherwise. The following areas are excluded from siting: Areas within 300m of a federal airport runway Waterbodies Areas with slope >16% Areas susceptible to sinkholes High coastal or inland flood risk areas Local, state, and federal parks, leisure areas, and cemeteries Areas >2 km away from electric substations Areas >5 km away from a municipal water supplier service area Areas >2 km away from high-speed fiber provider service territory Protected Areas Database of the United States (PAD-US) areas Railroads, major roadways, and minor roadways Military areas and training grounds Developed lands Areas >0.8 km (0.5 miles) from developed lands land_value_dollar_per_sqft.tif — USD per square foot (sqft) derived from USDA $/acre land cost personal_property_tax_rate.tif — Personal property tax rate by state. Uses an assumed 0.0125 personal property tax rate for states with personal property tax, 0 for states without personal property tax. real_property_tax_rate.tif — Real property tax rate. Based on county level residential real estate property tax rates. sales_tax_rate.tif — Sales tax rate by state. mechanical_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through mechanical processes based on local water stress and humidity levels. water_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through evaporative (water cooled) processes based on local water stress and humidity levels. distance_to_substation.tif — Distance to nearest substation in hundreds of meters (i.e., value of 1 equals a distance of 100m). Offshore areas have a value of 0. industrial_electricity_rates_dollar_per_kwh.tif — USD/kWh industrial electricity rates. Represents the average industrial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. commercial_electricity_rates_dollar_per_kwh.tif — USD/kWh commercial electricity rates. Represents the average commercial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. data_center_market_locations.tif — Grid cells with positive values represent the centroid of existing data center market clusters. The value of non-zero grid cells represents the number of data centers in the market cluster. All other grid cells have a value of 0. Geospatial Metadata CRS: Albers Equal Area Conic (ESRI:102003) Extent: -2415585.0000000023283064,-1441981.2605773280374706 : 2384414.9999999976716936,1708018.7394226719625294 Dimensions: X: 48000 Y: 31500 Bands: 1 Origin: -2415585.0000000023283064,1708018.7394226719625294 Pixel Size: 100,-100 Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4.0 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830
Package Data for CERF-Data Centers
This dataset contains sample input 100m resolution raster files for running the CERF-DC python package (see https://github.com/IMMM-SFA/cerf_data_centers) at the state level across the CONUS. Due to data availability constraints, some of the items included in this dataset are proxies or assumptions for siting factors used in the model. These are individually noted in the item descriptions and can be exchanged with more detailed information upon availability. Data Descriptions The following raster files are included in the data download: state_siting_region.tif — State areas identified by state FIPS code composite_siting_suitability.tif — Value of 1 indicates suitable siting location, 0 otherwise. The following areas are excluded from siting: Areas within 300 m of a federal airport runway or within an airport area boundary Waterbodies Areas with slope >16% Areas susceptible to sinkholes High coastal or inland flood risk areas Local, state, and federal parks, leisure areas, and cemeteries Areas >2 km away from electric substations Areas >5 km away from a municipal water supplier service area Areas >2 km away from high-speed fiber provider service territory USGS Protected Areas Database of the United States (PAD-US) GAP status 1, 2, or 3 areas US National Parks Wetlands USFWS critical habitats BIA land areas Railroads, major roadways, and minor roadways Military areas and training grounds NLCD developed lands Areas >0.8 km (0.5 miles) from NLCD developed lands land_value_dollar_per_sqft.tif — USD per square foot (sqft) derived from USDA $/acre land cost personal_property_tax_rate.tif — Personal property tax rate by state. Uses an assumed 0.0125 personal property tax rate for states with personal property tax, 0 for states without personal property tax. real_property_tax_rate.tif — Real property tax rate. Based on county level residential real estate property tax rates. sales_tax_rate.tif — Sales tax rate by state. mechanical_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through mechanical processes based on local water stress and humidity levels. water_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through evaporative (water cooled) processes based on local water stress and humidity levels. distance_to_substation.tif — Distance to nearest substation in hundreds of meters (i.e., value of 1 equals a distance of 100m). Offshore areas have a value of 0. industrial_electricity_rates_dollar_per_kwh.tif — USD/kWh industrial electricity rates. Represents the average industrial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. commercial_electricity_rates_dollar_per_kwh.tif — USD/kWh commercial electricity rates. Represents the average commercial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. data_center_market_locations.tif — Grid cells with positive values represent the centroid of existing data center market clusters. The value of non-zero grid cells represents the number of data centers in the market cluster. All other grid cells have a value of 0. Geospatial Metadata CRS: Albers Equal Area Conic (ESRI:102003) Extent: -2415585.0000000023283064,-1441981.2605773280374706 : 2384414.9999999976716936,1708018.7394226719625294 Dimensions: X: 48000 Y: 31500 Bands: 1 Origin: -2415585.0000000023283064,1708018.7394226719625294 Pixel Size: 100,-100 Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4.0 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830
Enhancing NASA Earth Science Data Discovery from Scientific Publications
Earth observations from space borne instruments have evolved explosively in the past decades. Following closely are reanalysis systems assimilating model and observational data, yielding even longer records and larger number of variables. Thanks to advances in internet technology, it is now easier than ever to visualize and analyze these data using web interfaces. On the other hand, it also becomes an increasingly daunting task to build upon the existing knowledge published in various peer reviewed sources, and navigate toward the most relevant data, analysis, and visualization. We present an analysis of a subset of publications that utilized a popular visualization web interface at the NASA Goddard Earth Science Data and Information Services Center. Known as "Giovanni", it allows researchers from wide backgrounds to work with hundreds of variables from space observations and assimilation systems. Since coming online more than a decade ago, Giovanni has been credited in more than 100 papers per year, and the total count now is estimated to be nearly 1,500. Many of these papers contain valuable information about when, where and how Giovanni has been used, and hence forge an opportunity to learn and share the knowledge of which variables were used for what research projects. The purpose of our work is to retrieve the information from the papers and organize it as a knowledge repository which links together datasets, variables, places, dates and phenomena all of which reflect the essence of the published research. Since the publications are unstructured texts, we use natural language processing along with machine learning methods in the retrieval process. One of the challenges is deciphering the dataset names, because in many cases researchers refer to variables, rather than the datasets containing them. To constrain the number of terms, we deploy Earth Science ontologies as dictionaries for the term extraction. We demonstrate that storing these terms and underlying ontologies, along with datasets, variables and papers in the knowledge graph database, enables various linkages between all these entities facilitating the data discovery. Thus, we are setting a qualitatively new stage in improvements of web data interfaces, where machine learning techniques are used to establish and optimize usage-based discovery of data.
Confronting Future Models with Future Satellite Observations of Clouds and Aerosols
The NASA Aerosol, Clouds, Convection and Precipitation (ACCP) Study convened a workshop in November 2020 to understand the future of modeling aerosols, clouds, convection and precipitation, and how satellite data can contribute to that future. ACCP is a project to define a satellite mission to be launched late in the 2020’s to advance cloud and aerosol science, following the recommendations of the latest NASA Decadal Survey. The ACCP modeling workshop goal was to answer the following questions: 1. What will be the critical science questions for clouds and aerosols in 10 years? 2. Where will simulations of clouds and aerosols across scales of space (process models to global) and time (nowcasting to climate prediction) be in 10 years? 3. What data will be available from space? What data would provide the most benefit? 4. What are the state of the art methods for confronting models with cloud and aerosol observations, including assimilation and climatological analysis techniques? The virtual workshop was anchored by a series of pre-recorded talks. Two days of synchronous sessions focused on discussion of the talks, along with small group breakout exercises. After an introduction to the ACCP concept came a panel discussing the future of modeling clouds and aerosols across scales. Participants were then asked to contribute their ideas. On the second day, there were two panel discussions. First came a discussion of the future of satellite observing systems. Second was a discussion of model-data synthesis methods. Finally, participants were asked to develop their own model-data synthesis proposal. The meeting began with an overview of the ACCP mission concept by Graeme Stephens (NASA-JPL). ACCP is a satellite mission for clouds and aerosols, likely anchored by advanced active lidar and radar systems in space, designed to observe detailed aerosol profile and type information, as well as cloud microphysics and dynamics. ACCP will integrate across sensors and observational types to get multi-spectral views of the same scenes, with better resolution than is available today. Launch is scheduled for 2027 or 2029. ACCP is being thought of as a comprehensive mission that may comprise more than one platform and more than one orbit plane (i.e., inclined and polar), with multiple combinations of sensors.
NASA Giovanni: Analyze, Compare, and Visualize 2000+ Earth Satellite and Model Variables Without Downloading Data and Software
Over vast oceans and remote continents, observations are often scarce and discontinuous. Satellite and model data play a critical role in research and applications. However, finding and accessing satellite and model data can be a daunting task for many, especially those outside the community. The NASA Goddard Earth Sciences (GES) Data and Information Services Center (DISC), one of 12 NASA Science Mission Directorate Data Centers, provides Earth science data, information, and services to everyone such as researchers, application users, educators, and students. GES DISC archives and supports datasets applicable to several NASA Earth Science Focus Areas including Atmospheric Composition, Water & Energy Cycles, Carbon Cycle & Ecosystem, and Climate Variability. To facilitate data discovery, evaluation, and exploration, GES DISC has developed the Geospatial Interactive Online Visualization ANd aNalysis Infrastructure (Giovanni), an online tool to analyze and visualize NASA remote sensing and model data without downloading data and software. As of this writing, over 2000 Earth satellite and model variables are available in Giovanni, including several wellknown NASA satellite missions (e.g., TRMM, GPM) and projects (e.g., MERRA-2, GPCP). Giovanni provides twenty-two plots that can be used to analyze, compare, and explore Earth data across different disciplines. Results can be shared with colleagues and downloaded for further analysis. Over the years, Giovanni has helped publish over 3000 referral papers. In this presentation, we will showcase key variables and plot types in Giovanni with examples. In particular, we will present several popular precipitation products from GPM and CPCP for evaluation and comparison.
Improving GES Disc Data Search and Discovery Through AI Metadata Augmentation
NASA’s Goddard Earth Science (GES) Data and Information Services Center (DISC) is one of twelve data centers in NASA's Science Mission Directorate (SMD), providing vital earth science data to a diverse user base. To enhance the discoverability of this data, GES DISC employs a keyword search system, which leverages scientific keywords embedded in dataset metadata. However, the evolving nature of scientific applications of our data necessitates regular review and augmentation of these keywords. To address this, we developed a service to automatically predict missing science keywords in the metadata. This service constructs a knowledge graph from the latest GES DISC metadata within NASA’s Common Metadata Repository (CMR). Using an open-source library, we trained a machine learning model to predict absent science keywords in the metadata. Our preliminary results indicate that the model has high levels of accuracy at predicting science keywords in the dataset metadata when exposed to data not included in its training. These predicted keywords were then evaluated by GES DISC data curation scientists and compared against other AI tools for metadata augmentation. We aim to enhance the overall usability and accessibility of NASA’s earth science data by implementing this tool in our data curation processes.
A Web-based Collaborative Tool for Mars Analog Data Exploration
Solving today's complex research and modeling challenges are dependent on our ability to discover, access, integrate, and share information from multiple sources. The planetary sciences community is no exception'; over the last few years, the need for data mining and exploration tools that can expedite comparative studies between Martian and terrestrial analogs sites and aid the interpretation of Mars data sets has become evident. Data sharing maximizes scientific return from studies and data sets.
IM3 Projected US Data Center Locations
IM3 Projected US Data Center Locations This dataset contains model projections of new data center facilities in the contiguous United States (CONUS) through 2035 using the CERF – Data Centers model. Data center locations are modeled across four data center electricity demand growth scenarios (low, moderate, high, higher) and five market gravity scenarios (0%, 25%, 50%, 75%, 100%). Projected locations are intended to be regional representations of feasible siting locations in the future to assess potential grid and water stress impacts. The data center load growth scenarios correspond with the rates outlined in EPRI (2024) and include 3.71%, 5%, 10%, and 15% annual growth of electricity demand for data centers from 2023 values in 37 states across the CONUS. Market gravity scenarios correspond to the relative importance of proximity to data center markets or high population areas compared to locational cost in the siting algorithm. 0% market gravity means that siting decisions were entirely determined by the locational cost in each feasible location. 100% market gravity means that only market proximity was considered when siting. Other scenarios have weight placed on both components where total weight always equals 100%. Locational cost is dependent on facility cooling type and corresponding electricity cost, taxes, and other factors. Facility cooling type is spatially determined where high water stress and/or areas with high summer wet bulb temperatures are assumed to operate with mechanical cooling for a higher fraction of the year rather than evaporative cooling. Feasible data center siting areas are based on geospatial suitability raster data developed with open-source information. The following areas are excluded from siting: Areas within 300 m of a federal airport runway Waterbodies Areas with slope >16% Areas susceptible to sinkholes High coastal or inland flood risk areas Local, state, and federal parks, leisure areas, and cemeteries Areas >2 km away from electric substations Areas >5 km away from a municipal water supplier service area Areas >2 km away from high-speed fiber provider service territory Protected Areas Database of the United States (PAD-US) areas Railroads, major roadways, and minor roadways Military areas and training grounds NLCD developed lands Areas >0.8 km (0.5 miles) from NLCD developed lands Because we use open-source information, proprietary information that can influence siting decisions such as individual tax agreements with cities, detailed fiber line connectivity, electric grid power capacity agreements, and others, are not currently accounted for in the modeling process. Using specific building locations and footprints in the dataset for local planning purposes is not advised. Technical Information Geospatial data is provided in geojson format using the Albers Equal Area Conic (ESRI:102003) coordinate reference system. The datasets contain the following parameters: id - unique identification number within given scenario file growth_scenario – data center demand growth scenario market_gravity_weight – market gravity weight scenario (%) region – name of region (i.e., US State) total_cost_million_usd – locational siting cost ($million) campus_size_square_ft – total land acquired for data center facility (square ft) data_center_it_power_mw – IT power of data center facility (MW) mechanical_cooling_frac – fraction of year when data center uses mechanical cooling system water_cooling_frac– fraction of year when data center uses evaporative cooling system cooling_energy_demand_mwh – total annual facility energy demand for cooling (MWh) cooling_water_demand_mgy – total annual facility water demand for cooling (MG) cooling_water_consumption_mgy – total annual facility water consumed (MG) normalized_locational_cost – normalized total locational cost score for location normalized_gravity_score – normalized market gravity score for location weighted_siting_score – total weighted siting score of locational cost and gravity score geometry – polygon geometry of facility Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4.0 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830
IM3 Projected US Data Center Locations
IM3 Projected US Data Center Locations This dataset contains model projections of new data center facilities in the contiguous United States (CONUS) through 2035 using the CERF – Data Centers model. Data center locations are modeled across four data center electricity demand growth scenarios (low, moderate, high, higher) and five market gravity scenarios (0%, 25%, 50%, 75%, 100%). Projected locations are intended to be regional representations of feasible siting locations in the future to assess potential grid and water stress impacts. The data center load growth scenarios correspond with the rates outlined in EPRI (2024) and include 3.71%, 5%, 10%, and 15% annual growth of electricity demand for data centers from 2023 values in 37 states across the CONUS. Market gravity scenarios correspond to the relative importance of proximity to data center markets or high population areas compared to locational cost in the siting algorithm. 0% market gravity means that siting decisions were entirely determined by the locational cost in each feasible location. 100% market gravity means that only market proximity was considered when siting. Other scenarios have weight placed on both components where total weight always equals 100%. Locational cost is dependent on facility cooling type and corresponding electricity cost, taxes, and other factors. Facility cooling type is spatially determined where high water stress and/or areas with high summer wet bulb temperatures are assumed to operate with mechanical cooling for a higher fraction of the year rather than evaporative cooling. Feasible data center siting areas are based on geospatial suitability raster data developed with open-source information. The following areas are excluded from siting: Areas within 300 m of a federal airport runway or within an airport area boundary Waterbodies Areas with slope >16% Areas susceptible to sinkholes High coastal or inland flood risk areas Local, state, and federal parks, leisure areas, and cemeteries Areas >2 km away from electric substations Areas >5 km away from a municipal water supplier service area Areas >2 km away from high-speed fiber provider service territory USGS Protected Areas Database of the United States (PAD-US) GAP status 1, 2, or 3 areas US National Parks Wetlands USFWS critical habitats BIA land areas Railroads, major roadways, and minor roadways Military areas and training grounds NLCD developed lands Areas >0.8 km (0.5 miles) from NLCD developed lands Because we use open-source information, proprietary information that can influence siting decisions such as individual tax agreements with cities, detailed fiber line connectivity, electric grid power capacity agreements, and others, are not currently accounted for in the modeling process. Using specific building locations and footprints in the dataset for local planning purposes is not advised. Technical Information Geospatial data is provided in geojson format using the Albers Equal Area Conic (ESRI:102003) coordinate reference system. The datasets contain the following parameters: id - unique identification number within given scenario file growth_scenario – data center demand growth scenario market_gravity_weight – market gravity weight scenario (%) region – name of region (i.e., US State) total_cost_million_usd – locational siting cost ($million) campus_size_square_ft – total land acquired for data center facility (square ft) data_center_it_power_mw – IT power of data center facility (MW) mechanical_cooling_frac – fraction of year when data center uses mechanical cooling system water_cooling_frac– fraction of year when data center uses evaporative cooling system cooling_energy_demand_mwh – total annual facility energy demand for cooling (MWh) cooling_water_demand_mgy – total annual facility water demand for cooling (MG) cooling_water_consumption_mgy – total annual facility water consumed (MG) normalized_locational_cost – normalized total locational cost score for location normalized_gravity_score – normalized market gravity score for location weighted_siting_score – total weighted siting score of locational cost and gravity score geometry – polygon geometry of facility Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4.0 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830