Search NASA⌕ Search

SEARCH · Search NASA

Results for “python”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53

sdt (Solar Data Tools) [SWR-25-130]

Solar Data Tools (sdt) is an open-source Python library for analyzing PV power (and irradiance) time-series data. It was developed to enable analysis of unlabeled PV data, i.e. with no model, no meteorological data, and no performance index required, by taking a statistical signal processing approach in the algorithms used in the package’s main data processing pipeline. Solar Data Tools empowers PV system fleet owners or operators to analyze system performance a hundred times faster even when they only have access to the most basic data stream—power output of the system.

Meyers-Im, Bennet [National Laboratory of the Rock↗

BM3DORNL

BM3DORNL is a high-performance, open-source library for removing streak and ring artifacts from computed-tomography (CT) data, developed for neutron imaging at Oak Ridge National Laboratory's Spallation Neutron Source (VENUS beamline) and applicable to X-ray CT as well. Ring artifacts — concentric rings in reconstructed slices caused by detector pixel-to-pixel response non-uniformities — appear as vertical streaks in the sinogram and degrade both image quality and quantitative analysis. BM3DORNL operates in the sinogram domain using an adaptation of the BM3D (block-matching and 3D collaborative filtering) algorithm (Dabov et al., 2007). It provides a dedicated streak-removal mode, a true multi-scale BM3D variant (after Mäkinen et al., 2021) that suppresses wide streaks single-scale methods miss, and an alternative Fourier–SVD method (~2.6× faster) combining FFT-based energy detection with rank-1 SVD. The computationally intensive core is implemented in Rust with parallel (Rayon) block matching, integral-image pre-screening, and optimized transforms, and is exposed through a simple Python API (with an optional GUI) so it integrates directly into existing tomography reconstruction pipelines. It processes both 2D sinograms and 3D sinogram stacks, is pip-installable for Linux and macOS, and is documented at https://bm3dornl.readthedocs.io.

Zhang, Chen [Oak Ridge National Laboratory (ORNL),↗

NEREIDS

NEREIDS (NEutron REsonance Imaging Diagnostic Suite) is a research-grade toolkit for neutron resonance imaging, providing a Rust physics core, Python bindings, and NeXus/HDF5-compliant data I/O, with a path to fast, standalone GUI workflows for VENUS/MARS.

Zhang, Chen [Oak Ridge National Laboratory (ORNL),↗

Data and Code for: Observation-constrained agroecosystem model inversion reveals continental-scale variation of winter wheat traits

This repository contains the simulation outputs and processing scripts associated with the study of winter wheat traits across the United States, utilizing the Ecosys agroecosystem model. The dataset includes model results for both rainfed and irrigated winter wheat systems, supporting the findings presented in the manuscript titled "Observation-constrained agroecosystem model inversion reveals continental-scale variation of winter wheat traits." Data includes the original Ecosys simulation outputs (archived in .db format within the compressed .zip files) and extracted analysis data (stored in .pkl files for efficient processing). Python code for data processing and figure generation is provided in a Jupyter notebook. External Observational Datasets should refer to the following official repositories for the input and validation data used in this study. The eddy covariance data from the AmeriFlux network (https://ameriflux.lbl.gov/). Climate-forcing data of NLDAS-2 from NASA LDAS (https://ldas.gsfc.nasa.gov/nldas/nldas-2-forcing-data). Soil data from the Gridded Soil Survey Geographic Database (gSSURGO), available at (https://www.nrcs.usda.gov/resources/data-and-reports/gridded-soil-survey-geographic-gssurgo-database). Crop yields, planting and harvest dates from the USDA public databases (https://quickstats.nass.usda.gov/; https://webapp.rma.usda.gov/apps/actuarialinformationbrowser/CropCriteria.aspx). Satellite-derived SLOPE GPP data from ORNL DAAC (https://daac.ornl.gov/cgi-bin/dsviewer.pl?ds_id=1786). Land use and crop progress information from the USDA Crop Data Layer and Crop Progress and Condition Gridded Layers (https://www.nass.usda.gov/Research_and_Science/). The Ecosys model code is available online at https://github.com/jinyun1tang/ECOSYS.

Wheat↗

Nondestructive Modular Leak Detection in 3D Printed 316L Stainless Steel Pipes via Laser Powder Bed Fusion

This research investigates the leak detection features of 316L Stainless Steel pipe structures manufactured via Laser Powder Bed Fusion (LPBF). This work involves the design of a modular sensor system integrating nondestructive evaluation (NDE) methods, including thermal imaging and ultrasonic frequency detection to detect and characterize leaks in components. This aims to improve leak detection sensitivity within medium-pressure gas systems, during continuous operation without halting flow or introducing safety risks. The system could be adaptable for use on unmanned aerial vehicles (UAVs), enabling remote leak detection in active environments. A custom pneumatic system incorporating temperature and pressure sensors was assembled to detect leaks in LPBF-printed 316L SS tee pipes. Experimental results and simulations confirm the system’s effectiveness in leak detection and material evaluation. This research program also integrated a Python-based image recognition platform based on a metallography and optical microscopy to assess the porosity and complement the leak detection data on the printed structures. This allows a detailed analysis of pore distribution and internal leak paths, which could compromise structural integrity, critical for quality control during manufacturing. Findings suggest that the investigated approach holds potential for enhancing leak detection technologies and adapt them for advanced manufactured parts.

36 MATERIALS SCIENCE↗

AWAKEN 2023 CLAMPS2 AERI TROPoe Retrievals

These files contain 24 hour periods of retrieved thermodynamic profiles derived from observations collected by the CLAMPS2 Atmospheric Emitted Radiance Interferometer (AERI). The TROPoe algorithm is a python equivalent to the AERIoe algorithm (see Turner and Loehnert, 2014; Turner and Blumberg 2019). These data were collected during the AWAKEN project. AERI only data is available.

54 ENVIRONMENTAL SCIENCES↗

AWAKEN 2023 CLAMPS2 MWR TROPoe Retrievals

These files contain 24 hour periods of retrieved thermodynamic profiles derived from observations collected by the CLAMPS2 Microwave Radiometer (MWR). The TROPoe algorithm is a python equivalent to the AERIoe algorithm (see Turner and Loehnert, 2014; Turner and Blumberg 2019). These data were collected during the AWAKEN project. MWR only data is available.

54 ENVIRONMENTAL SCIENCES↗

AWAKEN 2023 CLAMPS1 AERI TROPoe Retrievals

These files contain 24 hour periods of retrieved thermodynamic profiles derived from observations collected by the CLAMPS1 Atmospheric Emitted Radiance Interferometer (AERI). The TROPoe algorithm is a python equivalent to the AERIoe algorithm (see Turner and Loehnert 2014; Turner and Blumberg 2019). These data were collected during the AWAKEN project. AERI only data is available.

54 ENVIRONMENTAL SCIENCES↗

CACTI CSAPR2 Taranis Retrievals

Taranis is an end-to-end processing chain for radar data written in Python with C extension for computation performance. Features include: masking for quality control, specific differential phase (Kdp), attenuation correction for reflectivity factor (Z) and differential reflectivity (Zdr) in rain, and additional geophysical retrievals. Retrievals are mostly drawn from literature or open-source software when appropriate, and have been tested, tuned, and modified to work with one another cohesively rather than using isolated off-the-shelf algorithms. Incorporated algorithms include hydrometeor (echo) identification, rain water content, raindrop mass-weighted mean diameter (gamma size distribution assumption), and rainfall rate (QPE). Taranis data sets exist for CSAPR2 PPI, HSRHI, and sector RHI scans. Cartesian-gridded data sets were also produced as well as a near-surface rain rate retrieval. More details can be found in the README.

54 ENVIRONMENTAL SCIENCES↗

SOSAT: Geohazard Risk Assessment in Carbon Sequestration Operations

Like all operations in which fluids are injected into the subsurface, geologic carbon storage (GCS) presents inherit risks that must be assessed and mitigated to successfully deploy such technology in the field. For example, alterations to pore pressure and the stress state due to fluid injection may trigger certain geohazards, such as potential fault activation and induced seismicity or unintentional hydraulic fracturing. Here we present capabilities for assessing these potential risks using the State of Stress Analysis Tool (SOSAT): a Python library and web-based application capable of probabilistically estimating the subsurface state of stress informed by various field observations. Using posterior distributions of principal stress components, SOSAT can assess the risk of fault activation within the injection formation for either a critically oriented fault or a fault with a user-prescribed orientation. Additionally, SOSAT can estimate the probability of unintentional hydraulic fracturing of the intact reservoir rock. Here, we demonstrate these SOSAT capabilities using a hypothetical GCS site. This type of risk assessment can assist with the effective deployment of GCS technology in the field by informing safe design and management practices.

Haagenson, Ryan J.↗

CERF: IM3 Projected Western US Power Plant Locations

Overview The Capacity Expansion Regional Feasibility (CERF) model is an open-source geospatial python package that provides new power plant locations at a 1km resolution. The model ingests U.S. state or regional-scale electricity system capacity expansion plans, such as those produced by the Global Change Analysis Model (GCAM-USA), and identifies feasible, site-specific locations for individual new power plants (renewable and non-renewable). CERF combines high-resolution geospatial suitability analyses with an economic algorithm that selects individual plant siting locations based on grid interconnection costs and the locational marginal value of new generation. The model incorporates a wide range of dynamic constraints and opportunities, such as protected lands, population density, existing infrastructure, and water availability. This dataset provides CERF power plant siting results for IM3 Phase 2 simulations across eight different scenarios for the Western US through 2055. The scenarios include combinations of two Shared Socioeconomic Pathways (SSP3 and SSP5) with four high-resolution climate projections specific to the United States (see, https://tgw-data.msdlive.org/). These climate projections include "hotter" and "cooler" variants for two Representative Concentration Pathways (RCP4.5 and RCP8.5). The resulting eight simulations are: rcp45cooler_ssp3 rcp45cooler_ssp5 rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85cooler_ssp3 rcp85cooler_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 CERF siting results in this dataset correspond to capacity expansion plans in the GCAM-USA IM3 Phase 2 simulation data and are available for each of the above scenarios. Data Details Temporal Range: 2015-2055 in 5-year timesteps. Note that 2015 is the experiment base year and 2020 and beyond represent model simulation years. Spatial Range: Plant locations are provided for the eleven states in the Western US including Arizona, California, Colorado, Idaho, Montana, New Mexico, Nevada, Oregon, Utah, Washington, and Wyoming. Spatial Resolution: 1 km-squared, provided in x and y coordinates Geospatial Projection: Albers Equal Area Conic (ESRI:102003) File Type: csv The dataset contains subdirectories for each of the eight scenarios described in the overview. Each scenario folder contains two subfolders with the following information: 1. Power Plant Data This directory contains a single .csv file of power plant locations for both pre-existing (non-CERF sited plants in operation in 2015) and new (CERF-sited) power plants across the temporal range along with additional CERF model output parameters for CERF-sited plants. Plant with a siting year earlier than 2020 correspond to facilities that are operational leading into the first timestep CERF simulation. For a more detailed description of CERF model output parameters, see the CERF model documentation. Note that the cerf_plant_id parameter is unique within each scenario file but not across scenario files. Parameter Descriptions scenario - Name of scenario cerf_plant_id - Unique siting identifier cerf_sited - If True, indicates that plant was sited by CERF model. If False, indicates pre-existing facility region_name - Name of region (state) tech_id - Technology ID tech_name - Full generation technology name inclusive of cooling type (if applicable) and additional characteristics tech_simple - Simplified generation technology type unit_size_mw - Power plant unit size (MW) xcoord - X coordinate in the default CRS (meters) ycoord - Y coordinate in the default CRS (meters) index - Index position in the flattend 2D array buffer_in_km - Exclusion buffer around site (km) sited_year - Year of siting retirement_year - Year of retirement lmp_zone - Locational marginal price (LMP) zone ID locational_marginal_price_usd_per_mwh - Locational marginal price ($/MWh) generation_mwh_per_year - Generation output (MWh/yr) operating_cost_usd_per_year - Cost of plant operations ($/yr) net_operational_value - Net operational value based on LMP and and operating costs ($/yr) interconnection_cost - Cost of interconnection for transmission & gas pipeline (if applicable) net_locational_cost -- Difference of interconnection cost and operating value ($/yr) capacity_factor_fraction - Capacity factor (fraction) carbon_capture_rate_fraction - Carbon capture rate (fraction) fuel_co2_content_tons_per_btu - Fuel CO2 content (tons/Btu) fuel_price_usd_per_mmbtu - Fuel price ($/MMBtu) fuel_price_esc_rate_fraction - Fuel price escalation rate (fraction) heat_rate_btu_per_kWh - Heat rate (Btu/kWh) lifetime_yrs - Technology lifetime for annuity (years) operational_life_yrs - Operational lifetime for retirement (years) variable_om_usd_per_mwh - Variable operation and maintenance costs of yearly capacity use ($/MWh) variable_om_esc_rate_fraction - Variable operation and maintenance costs escalation rate (fraction) carbon_tax_usd_per_ton - Carbon tax ($/ton) carbon_tax_esc_rate_fraction - Carbon tax escalation rate (fraction) 2. Storage Data This directory contains information on new and pre-existing energy storage facilities operational in each timestep along with various storage operational parameters. The 2015 timestep provides pre-existing energy storage data and corresponds with facilities that are operational leading into the first model simulation timestep. Note that coordinates in the storage files correspond to the interconnection point on the grid (substation location), not individual energy storage locations. Energy storage is added in a cumulative process at each given interconnection point. That is, each individual file provides the total operational storage capacity interconnected to the specified substation for the given timestep, inclusive of previously installed storage at that location and new storage installed in that timestep at that location. Parameters scenario - Name of scenario timestep - Simulation timestep name - Unique storage identifier s_typ - Type of energy storage technology (battery or pumped storage hydro) s_node - Node ID of interconnecting substation xcoord - X coordinate in the default CRS (meters) ycoord - Y coordinate in the default CRS (meters) charge_rate - Maximum charge rate (power capacity) of storage system (MW) discharge_rate - Maximum discharge rate (power capacity) of storage system (MW) duration - Duration of storage system (hours) max_SoC - Allowed maximum state of charge (energy capacity) of storage system (MWh) min_SoC -Allowed minimum state of charge (energy capacity) of storage system (MWh) charge_eff - Efficiency of charge (fraction between 0 and 1) discharge_eff - Efficiency of discharge (fraction between 0 and 1) Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program.

CERF↗

The Capacity Expansion Regional Feasibility (CERF) Model: High-Resolution Power Plant Siting

Abstract This presentation gives an overview of the geospatial power plant siting model CERF. CERF (Capacity Expansion Regional Feasibility) is an open source Python package developed under the Integrated Multisector Multiscale Modeling (IM3) Project at PNNL. This presentation covers an overview of how the CERF model works, walks through various power plant siting analyses, and discusses future research opportunities for the model. The CERF model can be accessed at https://github.com/IMMM-SFA/cerf. PNNL Information Release Number: PNNL-SA-207336 Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program.

Mongird, Kendall [Pacific Northwest National Labor↗

Integrated GW Farm ABM

This Data Repository includes data used for the integrated groundwater- farm ABM model, raw model output from scenario ensemble, and processed outputs that isolate the groundwater storage depletion outcomes for the 35,000 farm cells. Model Inputs: Farm ABM Inputs: This folder contains the input data used by the integrated groundwater - farm ABM modelling script (Python file) used for the high performance computing (HPC) experiments. The sub-folder "data inputs" contains all of the farm attribute data, while the three files in the folder have the hydrogeological data lookup table (NLDAS Cost Curve Attributes.csv), a lookup table (Theis well function table.csv) for the groundwater cost curve function, and the farm indexes and corresponding NLDAS ids for all of the cells run in this experiment (nldas farms subset final.csv). NLDAS Cost curve hydrogeological data: Hydrogeological data aggregated to 1/8 degree resolution and aligned with the NLDAS grid. Parameters include: water depth below ground surface [meters], subsurface porosity [unitless], aquifer depth from ground surface to aquifer bottom [meters], annual average recharge (USGS: mm, Doll: meters), and three different hydraulic conductivity (K) values (meters/day). The three K values represent the mean value from Gleeson et al. (2018), one standard deviation above the mean from Gleeson et al. (2018), and the de Graaf et al. 2020 modifications to certain lithologies. Additional information about these datasets and their processing are documented in the supplement to Yoon et al. 2025 (in review). Output: Raw outputs: This folder contains a .zip file that has model outputs for the entire scenario ensemble. There is one csv for each farm id, using the format "farm farmid cases.csv". The relationship between the farm id and NLDAS id is defined by the "nldas farms subset final.csv" located in the Farm ABM Inputs folder. Each csv has 625 rows, corresponding to 625 combinations of different scenario parameter values. Each row (scenario) represents the outcome of a 100 year simulation. Columns define scenario settings and summary statistics for each scenario. The first four columns define the scenario settings: "hydro ratio," "econ ratio," "K scenario," and "gamma scenario." The hydro and econ ratios are values passed to the modeling script that influence multipliers for other model parameters, as documented in the supplement to Yoon et al. 2025 (in review). The gamma multiplier is a coefficient multiplier applied to the baseline gamma values (values below 1 represent lower unobserved costs compared to baseline, values above 1 represent higher costs). The K scenario names represent K values of: "low": 0.5 m/d, "int 1": 2.5 m/d, "int 2": 10 m/d, "high": 50 m/d, and "gleeson": mean Gleeson K value. "Perc vol depleted" is the fraction of groundwater depleted at the end of the 100 simulation. Processed Output: Derived depletion outcomes from raw outputs: All of the individual csv files from the Raw outputs were aggregated into a single file that has the scenario settings and fraction depletion "Perc vol depleted" for every farm cell, for every scenario. The other two files define relationships between the farm id, NLDAS id, and local and major aquifer units, used for aquifer-level depletion analysis.

Agent based modeling↗

Baltimore Social-Environmental Collaborative (BSEC) Doppler Lidar & Derived Products

This repository contains all processed Doppler‐lidar outputs from the PSU lidar deployed for the Baltimore Social‐Environmental Collaborative (BSEC) project. Vertical Stare Scans (fixed‐beam, vertical profiling): 1 Hz backscatter intensity (m⁻¹ sr⁻¹), signal‐to‐noise ratio (unitless), and Doppler vertical‐velocity (m s⁻¹) on ~30 m range gates, stored as CF-compliant NetCDF. Wind Profiles (horizontal‐wind retrieval): daily NetCDF outputs of retrieved horizontal wind speed (m s⁻¹) and direction (degrees), computed from the angled‐scan returns. Profile Statistics (summary statistics on the vertical velocity): 15 min windows (default) of mean, variance, skewness, kurtosis, high-frequency variance, etc., as a function of height; saved as CF-compliant NetCDF files. Boundary Layer Height (BLH) (fuzzy-logic output): 15 min BLH estimates (m), with lower/upper fuzzy bounds (m) and a quality flag (0–4) indicating data status (e.g., no data, good, below range, ran out of signal, cloud-topped). Cloud Base Height (Haar-gradient detection): 15 min estimates of cloud-base height (m) with a cloud-detection quality flag (0–3: none, low, moderate, high). All five product streams are organized by year and date under their own top-level folders (01_Vertical_Stare_Scans/ through 05_Cloud_Height/). Each folder contains a data_ /YYYY/ subdirectory with daily CF-compliant NetCDF outputs (96 windows per day at 15 min intervals). Global attributes in each file include creation history, version (2.0.0), institution, and source. Instrument & MeasurementsThe PSU Doppler Lidar samples aerosol backscatter (m⁻¹ sr⁻¹), signal-to-noise ratio, and radial velocity at ~1 Hz. Vertical stare scans point the beam straight up; after collecting angled scans through multiple elevation angles, the "Wind Profiles" product contains the fully retrieved horizontal wind speed and direction. Data were collected continuously at ~30 m range resolution, with a typical height ceiling of ~12 km. How to Use Open any NetCDF with Python's xarray, MATLAB, or similar CF-compliant tools. Stare scans and angled-scan retrievals (Wind Profiles) are CF-compliant daily NetCDF files. Profile-Statistics, BLH, and Cloud Height files are daily 15 min summaries (96 time steps per file). Inspect the included variables (e.g., vertical_velocity_variance, wind_speed, BLH, cloud_base_height) for your analyses. Use the quality flags (BLH_flag, cloud_flag) to filter out poor-quality retrievals. For more information or questions about processing methods, please contact:Nicholas E. Prince ⟨nec5299@psu.edu⟩Penn State Department of Meteorology & Atmospheric Science

Air Quality↗

NATURF: Urban Building Parameters for Chicago, Illinois, USA at a 100m resolution

132 Urban parameters based on building physical dimensions and location were generated for the city of Chicago at 100m resolution using the NATURF model. To use the binary file with WRF, the binary file and the index file must be placed in their own directory in WRF_GEOG and accessed in the same way NUDAPT44 would be accessed.

Vernon, Chris R [Pacific Northwest National Labora↗

Package Data for CERF-Data Centers

This dataset contains sample input 100m resolution raster files for running the CERF-DC python package (see https://github.com/IMMM-SFA/cerf_data_centers) at the state level across the CONUS. Due to data availability constraints, some of the items included in this dataset are proxies or assumptions for siting factors used in the model. These are individually noted in the item descriptions and can be exchanged with more detailed information upon availability. Data Descriptions The following raster files are included in the data download: state_siting_region.tif — State areas identified by state FIPS code composite_siting_suitability.tif — Value of 1 indicates suitable siting location, 0 otherwise. The following areas are excluded from siting: Areas within 300m of a federal airport runway Waterbodies Areas with slope >16% Areas susceptible to sinkholes High coastal or inland flood risk areas Local, state, and federal parks, leisure areas, and cemeteries Areas >2 km away from electric substations Areas >5 km away from a municipal water supplier service area Areas >2 km away from high-speed fiber provider service territory Protected Areas Database of the United States (PAD-US) areas Railroads, major roadways, and minor roadways Military areas and training grounds Developed lands Areas >0.8 km (0.5 miles) from developed lands land_value_dollar_per_sqft.tif — USD per square foot (sqft) derived from USDA $/acre land cost personal_property_tax_rate.tif — Personal property tax rate by state. Uses an assumed 0.0125 personal property tax rate for states with personal property tax, 0 for states without personal property tax. real_property_tax_rate.tif — Real property tax rate. Based on county level residential real estate property tax rates. sales_tax_rate.tif — Sales tax rate by state. mechanical_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through mechanical processes based on local water stress and humidity levels. water_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through evaporative (water cooled) processes based on local water stress and humidity levels. distance_to_substation.tif — Distance to nearest substation in hundreds of meters (i.e., value of 1 equals a distance of 100m). Offshore areas have a value of 0. industrial_electricity_rates_dollar_per_kwh.tif — USD/kWh industrial electricity rates. Represents the average industrial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. commercial_electricity_rates_dollar_per_kwh.tif — USD/kWh commercial electricity rates. Represents the average commercial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. data_center_market_locations.tif — Grid cells with positive values represent the centroid of existing data center market clusters. The value of non-zero grid cells represents the number of data centers in the market cluster. All other grid cells have a value of 0. Geospatial Metadata CRS: Albers Equal Area Conic (ESRI:102003) Extent: -2415585.0000000023283064,-1441981.2605773280374706 : 2384414.9999999976716936,1708018.7394226719625294 Dimensions: X: 48000 Y: 31500 Bands: 1 Origin: -2415585.0000000023283064,1708018.7394226719625294 Pixel Size: 100,-100 Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4.0 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall↗

Package Data for CERF-Data Centers

This dataset contains sample input 100m resolution raster files for running the CERF-DC python package (see https://github.com/IMMM-SFA/cerf_data_centers) at the state level across the CONUS. Due to data availability constraints, some of the items included in this dataset are proxies or assumptions for siting factors used in the model. These are individually noted in the item descriptions and can be exchanged with more detailed information upon availability. Data Descriptions The following raster files are included in the data download: state_siting_region.tif — State areas identified by state FIPS code composite_siting_suitability.tif — Value of 1 indicates suitable siting location, 0 otherwise. The following areas are excluded from siting: Areas within 300 m of a federal airport runway or within an airport area boundary Waterbodies Areas with slope >16% Areas susceptible to sinkholes High coastal or inland flood risk areas Local, state, and federal parks, leisure areas, and cemeteries Areas >2 km away from electric substations Areas >5 km away from a municipal water supplier service area Areas >2 km away from high-speed fiber provider service territory USGS Protected Areas Database of the United States (PAD-US) GAP status 1, 2, or 3 areas US National Parks Wetlands USFWS critical habitats BIA land areas Railroads, major roadways, and minor roadways Military areas and training grounds NLCD developed lands Areas >0.8 km (0.5 miles) from NLCD developed lands land_value_dollar_per_sqft.tif — USD per square foot (sqft) derived from USDA $/acre land cost personal_property_tax_rate.tif — Personal property tax rate by state. Uses an assumed 0.0125 personal property tax rate for states with personal property tax, 0 for states without personal property tax. real_property_tax_rate.tif — Real property tax rate. Based on county level residential real estate property tax rates. sales_tax_rate.tif — Sales tax rate by state. mechanical_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through mechanical processes based on local water stress and humidity levels. water_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through evaporative (water cooled) processes based on local water stress and humidity levels. distance_to_substation.tif — Distance to nearest substation in hundreds of meters (i.e., value of 1 equals a distance of 100m). Offshore areas have a value of 0. industrial_electricity_rates_dollar_per_kwh.tif — USD/kWh industrial electricity rates. Represents the average industrial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. commercial_electricity_rates_dollar_per_kwh.tif — USD/kWh commercial electricity rates. Represents the average commercial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. data_center_market_locations.tif — Grid cells with positive values represent the centroid of existing data center market clusters. The value of non-zero grid cells represents the number of data centers in the market cluster. All other grid cells have a value of 0. Geospatial Metadata CRS: Albers Equal Area Conic (ESRI:102003) Extent: -2415585.0000000023283064,-1441981.2605773280374706 : 2384414.9999999976716936,1708018.7394226719625294 Dimensions: X: 48000 Y: 31500 Bands: 1 Origin: -2415585.0000000023283064,1708018.7394226719625294 Pixel Size: 100,-100 Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4.0 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall↗

Baltimore Social-Environmental Collaborative (BSEC) Doppler Lidar & Derived Products

This repository contains all processed Doppler‐lidar outputs from the PSU lidar deployed for the Baltimore Social‐Environmental Collaborative (BSEC) project. Vertical Stare Scans (fixed‐beam, vertical profiling): 1 Hz backscatter intensity (m⁻¹ sr⁻¹), signal‐to‐noise ratio (unitless), and Doppler vertical‐velocity (m s⁻¹) on 30 m range gates (and 3 m range gates), stored as CF-compliant NetCDF. Wind Profiles (horizontal‐wind retrieval): daily NetCDF outputs of retrieved horizontal wind speed (m s⁻¹) and direction (degrees), computed from the angled‐scan returns. Profile Statistics (summary statistics on the vertical velocity): 15 min windows (default) of mean, variance, skewness, kurtosis, high-frequency variance, etc., as a function of height; saved as CF-compliant NetCDF files. Boundary Layer Height (BLH) (fuzzy-logic output): 15 min BLH estimates (m), with lower/upper fuzzy bounds (m) and a quality flag (0–4) indicating data status (e.g., no data, good, ran out of signal, below range, cloud-topped). Cloud Base Height (Haar-gradient detection): 10 min estimates of cloud-base height (m). All five product streams are organized by year and date under their own top-level folders (01_Vertical_Stare_Scans/ through 05_Cloud_Height/). Each folder contains a data_ /YYYY/ subdirectory with daily CF-compliant NetCDF outputs (96 windows per day at 15 min intervals). Global attributes in each file include creation history, version (3.0.0), institution, and source. Instrument & MeasurementsThe PSU Doppler Lidar samples aerosol backscatter (m⁻¹ sr⁻¹), signal-to-noise ratio, and radial velocity at ~1 Hz. Vertical stare scans point the beam straight up; after collecting angled scans through multiple elevation angles, the "Wind Profiles" product contains the fully retrieved horizontal wind speed and direction. Data were collected continuously at ~30 m range resolution (and 3 m for the year of 2025), with a typical height ceiling of ~12 km. How to Use Open any NetCDF with Python's xarray, MATLAB, or similar CF-compliant tools. Stare scans and angled-scan retrievals (Wind Profiles) are CF-compliant daily NetCDF files. Profile-Statistics, BLH, and Cloud Height files are daily 15 min (10 min for Cloud Heights) summaries (96 time steps per file). Inspect the included variables (e.g., vertical_velocity_variance, wind_speed, BLH, cloud_base_height) for your analyses. Use the quality flags (BLH_flag, cloud_flag) to filter out poor-quality retrievals. For more information or questions about processing methods, please contact:Nicholas E. Prince ⟨nec5299@psu.edu⟩Penn State Department of Meteorology & Atmospheric Science

Air Quality↗