Search NASA⌕ Search

SEARCH · Search NASA

Results for “Python”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55

Package Data for CERF-Data Centers

This dataset contains sample input 100m resolution raster files for running the CERF-DC python package (see https://github.com/IMMM-SFA/cerf_data_centers) at the state level across the CONUS. Due to data availability constraints, some of the items included in this dataset are proxies or assumptions for siting factors used in the model. These are individually noted in the item descriptions and can be exchanged with more detailed information upon availability. Data Descriptions The following raster files are included in the data download: state_siting_region.tif — State areas identified by state FIPS code composite_siting_suitability.tif — Value of 1 indicates suitable siting location, 0 otherwise. The following areas are excluded from siting: Areas within 300 m of a federal airport runway or within an airport area boundary Waterbodies Areas with slope >16% Areas susceptible to sinkholes High coastal or inland flood risk areas Local, state, and federal parks, leisure areas, and cemeteries Areas >2 km away from electric substations Areas >5 km away from a municipal water supplier service area Areas >2 km away from high-speed fiber provider service territory USGS Protected Areas Database of the United States (PAD-US) GAP status 1, 2, or 3 areas US National Parks Wetlands USFWS critical habitats BIA land areas Railroads, major roadways, and minor roadways Military areas and training grounds NLCD developed lands Areas >0.8 km (0.5 miles) from NLCD developed lands land_value_dollar_per_sqft.tif — USD per square foot (sqft) derived from USDA $/acre land cost personal_property_tax_rate.tif — Personal property tax rate by state. Uses an assumed 0.0125 personal property tax rate for states with personal property tax, 0 for states without personal property tax. real_property_tax_rate.tif — Real property tax rate. Based on county level residential real estate property tax rates. sales_tax_rate.tif — Sales tax rate by state. mechanical_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through mechanical processes based on local water stress and humidity levels. water_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through evaporative (water cooled) processes based on local water stress and humidity levels. distance_to_substation.tif — Distance to nearest substation in hundreds of meters (i.e., value of 1 equals a distance of 100m). Offshore areas have a value of 0. industrial_electricity_rates_dollar_per_kwh.tif — USD/kWh industrial electricity rates. Represents the average industrial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. commercial_electricity_rates_dollar_per_kwh.tif — USD/kWh commercial electricity rates. Represents the average commercial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. data_center_market_locations.tif — Grid cells with positive values represent the centroid of existing data center market clusters. The value of non-zero grid cells represents the number of data centers in the market cluster. All other grid cells have a value of 0. Geospatial Metadata CRS: Albers Equal Area Conic (ESRI:102003) Extent: -2415585.0000000023283064,-1441981.2605773280374706 : 2384414.9999999976716936,1708018.7394226719625294 Dimensions: X: 48000 Y: 31500 Bands: 1 Origin: -2415585.0000000023283064,1708018.7394226719625294 Pixel Size: 100,-100 Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4.0 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall↗

Baltimore Social-Environmental Collaborative (BSEC) Doppler Lidar & Derived Products

This repository contains all processed Doppler‐lidar outputs from the PSU lidar deployed for the Baltimore Social‐Environmental Collaborative (BSEC) project. Vertical Stare Scans (fixed‐beam, vertical profiling): 1 Hz backscatter intensity (m⁻¹ sr⁻¹), signal‐to‐noise ratio (unitless), and Doppler vertical‐velocity (m s⁻¹) on 30 m range gates (and 3 m range gates), stored as CF-compliant NetCDF. Wind Profiles (horizontal‐wind retrieval): daily NetCDF outputs of retrieved horizontal wind speed (m s⁻¹) and direction (degrees), computed from the angled‐scan returns. Profile Statistics (summary statistics on the vertical velocity): 15 min windows (default) of mean, variance, skewness, kurtosis, high-frequency variance, etc., as a function of height; saved as CF-compliant NetCDF files. Boundary Layer Height (BLH) (fuzzy-logic output): 15 min BLH estimates (m), with lower/upper fuzzy bounds (m) and a quality flag (0–4) indicating data status (e.g., no data, good, ran out of signal, below range, cloud-topped). Cloud Base Height (Haar-gradient detection): 10 min estimates of cloud-base height (m). All five product streams are organized by year and date under their own top-level folders (01_Vertical_Stare_Scans/ through 05_Cloud_Height/). Each folder contains a data_ /YYYY/ subdirectory with daily CF-compliant NetCDF outputs (96 windows per day at 15 min intervals). Global attributes in each file include creation history, version (3.0.0), institution, and source. Instrument & MeasurementsThe PSU Doppler Lidar samples aerosol backscatter (m⁻¹ sr⁻¹), signal-to-noise ratio, and radial velocity at ~1 Hz. Vertical stare scans point the beam straight up; after collecting angled scans through multiple elevation angles, the "Wind Profiles" product contains the fully retrieved horizontal wind speed and direction. Data were collected continuously at ~30 m range resolution (and 3 m for the year of 2025), with a typical height ceiling of ~12 km. How to Use Open any NetCDF with Python's xarray, MATLAB, or similar CF-compliant tools. Stare scans and angled-scan retrievals (Wind Profiles) are CF-compliant daily NetCDF files. Profile-Statistics, BLH, and Cloud Height files are daily 15 min (10 min for Cloud Heights) summaries (96 time steps per file). Inspect the included variables (e.g., vertical_velocity_variance, wind_speed, BLH, cloud_base_height) for your analyses. Use the quality flags (BLH_flag, cloud_flag) to filter out poor-quality retrievals. For more information or questions about processing methods, please contact:Nicholas E. Prince ⟨nec5299@psu.edu⟩Penn State Department of Meteorology & Atmospheric Science

Air Quality↗

Sensitivity Analysis of Drivers Water Shortage in the Los Angeles Region During Drought

The code and detailed step-by-step instructions for generating the model output data, processing results, and analysis and plotting are provided at https://github.com/IMMM-SFA/Ferencz_et_al_2026_ER_Water. The PyArtes model is a python adaptation of the Artes model. PyArtes uses many of the same input data and optimization model architecture as Artes. Documentation for the PyArtes model is provided in the Supplement to the paper. The primary data product are simulated monthly water shortages for indoor and outdoor demand under a large ensemble of drought scenarios (>13,000). The droughts are hypothetical and are not based on historical time series data of supply sources - though historical data did help inform ranges explored for supply parameters. Demands are informed by recent 2017-2021 water supply data. Demands used for the model can be accessed at https://github.com/IMMM-SFA/Ferencz_et_al_2026_ER_Water. Simulations resolve demand for over 90 water providers in the study region. The results report 36 months of water shortage data for each indoor and outdoor demand node. The study also developed a multilayer perceptron (MLP) neural network trained on a subset of the simulated shortage ensemble to emulate worst annual water shortage for a given set of parameter multipliers -- provided the parameter values fall within the ranges sampled in the ensemble. Emulated water shortages for synthetic ensembles are in the MLP-generated shortages folder. The MLP model was used to generate larger ensembles to support Sobol analysis that would have been extremely computationally expensive to simulate. Datasets provided in this repository*: Simulated shortages. These results are used for the analysis for Figures 5, 8, and 9 in the paper, and also to train the MLP emulator. .zip file containing outputs for the 13,312 scenario ensemble. Separate .csv files for indoor and outdoor shortage for each scenario. Rows = demand ids (~100), Columns = months (36) Units = acre-feet/month of shortage (shortage = monthly demand - supply). 1 acft = 1233.48 m^3 .csv files of aggregated shortages derived from the 13,312 ensemble Rows = scenarios (13,312), Columns = demand ids (~100) Units = acre-feet/year (either worst annual shortage or total shortage over the 3-year drought) .csv file of the parameter multipliers scenarios for the ensemble .csv file of the parameter ranges and baseline values the multipliers were applied to MLP-generated shortages. These results are used for Figures 4, 6, and 7 in the paper. mwd higher folder: scenario ensembles, emulated worst year total shortages (acft), and Sobol results Emulated shortages. Rows = scenarios, columns = demand ids, units acft Sobol results. Rows = demand ids, columns Sobol (S1, ST, or 95% confidence interval) value for each parameter mwd lower folder: scenario ensembles, emulated worst year total shortages (acft), and Sobol results same organization as mwd higher MLP performance: performance metrics (R^2, RMSE, BIAS, MAPE) for the testing subset (20% or 2,662 scenarios) and simulated vs emulated worst year shortage (acre-feet/year) for every demand node, MWD wholesale regions, and the entire study region (LAC). Supporting data for figures. Figure plotting scripts in the associated GitHub repo. These files support analysis and visualization. Geospatial Data used for plotting simulated water shortages and Sobol results. Dictionary of full names for demand nodes in the model and estimates of water supply by source type informed by Artes input files and California Urban Water Management Planning data: https://water.ca.gov/Programs/Water-Use-And-Efficiency/Urban-Water-Use-Efficiency/Urban-Water-Management-Plans *Readme files provided for each folder.

drought↗

Baltimore Social-Environmental Collaborative (BSEC) Doppler Lidar & Derived Products

This repository contains all processed Doppler‐lidar outputs from the PSU lidar deployed for the Baltimore Social‐Environmental Collaborative (BSEC) project. Vertical Stare Scans (fixed‐beam, vertical profiling): 1 Hz backscatter intensity (m⁻¹ sr⁻¹), signal‐to‐noise ratio (unitless), and Doppler vertical‐velocity (m s⁻¹) on 30 m range gates (and 3 m range gates), stored as CF-compliant NetCDF. Wind Profiles (horizontal‐wind retrieval): daily NetCDF outputs of retrieved horizontal wind speed (m s⁻¹) and direction (degrees), computed from the angled‐scan returns. Profile Statistics (summary statistics on the vertical velocity): 15 min windows (default) of mean, variance, skewness, kurtosis, high-frequency variance, etc., as a function of height; saved as CF-compliant NetCDF files. Boundary Layer Height (BLH) (fuzzy-logic output): 15 min BLH estimates (m), with lower/upper fuzzy bounds (m) and a quality flag (0–4) indicating data status (e.g., no data, good, ran out of signal, below range, cloud-topped). Cloud Base Height (Haar-gradient detection): 10 min estimates of cloud-base height (m). All five product streams are organized by year and date under their own top-level folders (01_Vertical_Stare_Scans/ through 05_Cloud_Height/). Each folder contains a data_ /YYYY/ subdirectory with daily CF-compliant NetCDF outputs (96 windows per day at 15 min intervals). Global attributes in each file include creation history, version (3.0.1), institution, and source. Instrument & MeasurementsThe PSU Doppler Lidar samples aerosol backscatter (m⁻¹ sr⁻¹), signal-to-noise ratio, and radial velocity at ~1 Hz. Vertical stare scans point the beam straight up; after collecting angled scans through multiple elevation angles, the "Wind Profiles" product contains the fully retrieved horizontal wind speed and direction. Data were collected continuously at ~30 m range resolution (and 3 m for the year of 2025), with a typical height ceiling of ~12 km. How to Use Open any NetCDF with Python's xarray, MATLAB, or similar CF-compliant tools. Stare scans and angled-scan retrievals (Wind Profiles) are CF-compliant daily NetCDF files. Profile-Statistics, BLH, and Cloud Height files are daily 15 min (10 min for Cloud Heights) summaries (96 time steps per file). Inspect the included variables (e.g., vertical_velocity_variance, wind_speed, BLH, cloud_base_height) for your analyses. Use the quality flags (BLH_flag, cloud_flag) to filter out poor-quality retrievals. For more information or questions about processing methods, please contact:Nicholas E. Prince ⟨nec5299@psu.edu⟩Penn State Department of Meteorology & Atmospheric Science

Air Quality↗

Heat Pump Retrofits for Central Plant Hydronic Heating Systems: A Software Toolkit for Screening and Design

Retrofitting existing central plants with high-efficiency heat pump technologies can play a crucial role in achieving long-term planning goals. Modern heat pump technologies are able to use waste heat recovery to meet a building's heating demand, but there is a lack of accessible tools designed for non-HVAC experts, such as building owners, to quickly and easily conduct what-if analysis, e.g., estimating retrofit costs and payback period for their partial or full equipment replacement. This paper introduces an open-source software toolkit designed to facilitate the initial screening and decision-making of heat pump retrofits in existing central plants using a building's yearly load profile from metered or utility bill data. The toolkit evaluates the technical and economic viability of replacing traditional central plant equipment with various options including water-to-water or air-to-water heat pumps, which can provide efficient and lower-cost heating and cooling. It allows users to compare current central plant configurations with retrofit scenarios, assessing energy consumption, life-cycle costs, and environmental impact. The toolkit offers (1) a web-based tool designed for user-friendly access by a broad audience and (2) Python-based source code for researchers and engineers conducting parametric studies and design parameter optimization. The toolkit compares a typical central plant configuration to a configuration that uses a heat pump to supply hydronic heating and cooling. The output metrics include energy consumption and output of each equipment, life-cycle cost analyses and metrics, and environmental impact of the system.

Excell, L↗

Explainable Machine Learning for Functional Data

Black-box machine learning models are recognized as useful tools for prediction applications, but the algorithmic complexity of some models causes interpretation challenges. Explainability methods have been proposed to provide insight into these models, but there is little research focused on supervised modeling with functional data inputs. We argue that, especially in applications of high consequence, it is important to explicitly model the functional dependence in a black-box analysis to not obscure or misrepresent patterns in explanations. As such, we propose the V ariable importance E xplainable E lastic S hape A nalysis (VEESA) pipeline for training supervised machine learning models with functional inputs. The pipeline is an analysis process that includes the data preprocessing, modeling, and post-hoc explanations. The preprocessing is done using elastic functional principal components analysis, which accounts for vertical and horizontal variability in functional data and, ultimately, allows for explanations in the original data space that identify the important functional variability without bias due to correlated variables. Here, we demonstrate the pipeline on two high-consequence applications: explosives classification for national security and inkjet printer identification in forensic science. The applications exhibit the VEESA pipeline’s ability to provide an understanding of the characteristics of the functional data useful for prediction. Code for implementing the pipeline is available in the veesa R package (and supplemental python code).

Elastic Shape Analysis↗

Dielectric Resonator Design for Low Power and Low Temperature Microwave Plasma

Waveguide-based microwave plasmas generally operate at high temperatures (2000 - 6000K)[1], making it difficult to directly interface solid materials with the plasma without significant thermal damage. Dielectric microwave resonators (DMRs), long studied for wave-based manipulation of electromagnetic radiation for telecom and optics, can focus radiation to extremely small mode volumes, creating intense localized fields with low-power input.[2] This phenomenon can be used for applications ranging from efficient plasma electronics to near-ambient plasma-materials interactions. Such DMR-based plasmas have been demonstrated a handful of times in the literature, but the majority of research towards this utilize the lowest frequency resonance mode.[3], [4], [5] By carefully controlling the geometry of cylindrical resonators, a variety of electromagnetic modes can be excited. In this work, COMSOL Multiphysics simulations are used to study the electric field enhancement and absorption properties of CaTiO3 DMRs as a function of geometry and excitation frequency. Whereas previous studies have utilized the HEM111 resonance frequency to drive low power plasma excitation, we find that higher order resonance frequencies are more effective at field enhancement and result in less power loss within the dielectric material, hence less wasted heating. The effectiveness of these modes is also geometry dependent and can be computationally optimized for plasma generation. Complementing these computational efforts, we demonstrate a new closed-system reactor design built in a WR-650 waveguide and experimentally demonstrate the formation of atmospheric argon microwave plasma using < 30 W input power on DMR dimers. We observe a shifting resonance frequency as the DMRs heat in response to microwave excitation and develop a Python-based lock-in mechanism to effectively track the DMR resonance over time, leading to stable plasma operation. We use infrared thermal imaging to monitor the temperature of the DMR dimers and surrounding quartz chamber, demonstrating thermal temperatures < 60 degreesC. Finally, we utilize optical emission spectroscopy (OES) to probe the plasma properties as a function of the resonance mode.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

A PSCAD Library Component Featuring a Reduced-Order IBR Model for EMT-Based Fault Studies

This paper presents a fully implemented inverter reduce-order-model (ROM) in an EMT simulation (PSCAD) library component for direct user utilization in protection studies. The developed inverter ROM has the following features: Equivalent to a full IBR inverter model with positive- and negative-sequence current formulation and representation A python script is developed to fully automate this process, including training data generation, ROM parameter training, updating parameters, and model verification and validation. With this PSCAD ROM library component, protection engineers can utilize a trustworthy, accurate ROM for protection studies in an easy-to-use and streamlined manner.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Evaluation of a Reduced-Order Model for IBR Fault Response Representation via OEM Blackbox Models

This paper presents a fully implemented inverter reduced-order-model (ROM) in an EMT simulation (PSCAD) library component for direct user utilization in protection studies. The developed inverter ROM has the following features: Equivalent to a full inverter-based resource (IBR) inverter model with positive- and negative-sequence current formulation and representation. A Python script is developed to fully automate this process, including training data generation, ROM parameter training, updating parameters, and model verification and validation. The ROM is validated using both IEEE 2800-compliant and non-compliant OEM modes in a real-world system, building confidence of its usability by protection engineers.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Advancing Concentrating Solar Thermal Modeling Using System Advisor Model (SAM)

Concentrating solar thermal (CST) technologies play a critical role in enabling dispatchable power and high-temperature industrial heat applications. Accurate and flexible modeling tools are essential for evaluating system performance, guiding technology research and development, and informing investment decisions. The National Laboratory of the Rockies's System Advisor Model (SAM) is a widely used techno-economic simulation platform for CST systems, providing detailed performance and financial modeling capabilities for multiple CST system configurations. SAM integrates physics-based performance models with financial analysis to simulate the behavior of complex energy systems under realistic operating conditions. For CST technologies (including tower, parabolic trough, and linear Fresnel), SAM enables hourly simulations using site-specific weather data that ensure feasible operating conditions and convergence of mass and energy between core system components (i.e., solar field, receiver, thermal energy storage, and power cycle). These capabilities allow researchers and developers to evaluate annual energy production, capacity factors, levelized cost of energy (LCOE), and system dispatch strategies. A key advantage of SAM lies in its flexibility for parametric analysis and large-scale computational studies. Users can vary system design parameters such as heliostat field layout, receiver dimensions, thermal energy storage capacity, power block sizing, and installation cost assumptions to investigate their impact on system performance and financial metrics. When combined with automated scripting through LK, SDKTool, or Python interfaces, SAM enables high-throughput simulation workflows that support sensitivity analysis, technology benchmarking, and optimization studies. These approaches are particularly valuable for next-generation CST concepts, where design spaces are large and system interactions are complex. Another important capability of SAM is its support for dispatch optimization and thermal energy storage modeling, which are central to the value proposition of CST technologies. The ability to simulate integrated storage and flexible power generation allows researchers to explore strategies that maximize grid value, improve capacity utilization, and enhance integration with variable resources such as photovoltaic and wind generation. This poster will present an overview of SAM's thermal system modeling capabilities including concentrating solar. Additionally, we will highlight new feature developments including: 1) implementing Google's OR-Tools optimization platform for faster and more robust dispatch optimization, 2) developing a new power load following controller for modeling behind-the-meter applications, 3) enabling direct modeling of CSP-PV hybrid systems with the inclusion of battery storage, and 4) developing a multi-receiver falling particle Gen3 system model.

14 SOLAR ENERGY↗

3D Play Fairway Analysis for Examining of Superhot Reservoir Production Scenarios

The DEEPEN (DE-risking Exploration for geothermal Plays in magmatic ENvironments) project was a multi-laboratory, international effort to reduce uncertainty and improve resource characterization in superhot geothermal systems. Building on this foundation, this work advances open-source tools designed to lower the exploration risk and cost of superhot geothermal projects while promoting transparency, reproducibility, and efficiency in exploration workflows. These tools are being tested at two key sites: (1) the Nesjavellir Geothermal Area in Iceland, where the Icelandic Deep Drilling Project (IDDP) will drill its third well, and (2) Newberry Volcano in Oregon, USA, where Mazama Energy will pilot the first superhot enhanced geothermal system (EGS). A major outcome is the creation of a modular, open-source Python framework for play fairway analysis (PFA) in 2D and 3D, called geoPFA. The PFA workflow has been expanded to produce pseudo conceptual models, and will soon be refined to assess reservoir components through integration with the thermo-hydraulic-mechanical-chemical (THMC) simulator TReactMech, to enable iterative coupling between PFA and THMC models, improving characterization of superhot systems. All three of the Icelandic Deep Drilling Project's production scenarios were analyzed via this framework: (1) a superhot deep injection well paired with conventional production wells at Nesjavellir, (2) a superhot deep production well at Nesjavellir, and (3) superhot enhanced geothermal system at Newberry Volcano. This analysis provides useful insights around conceptual modeling of these production scenarios, helping to inform decisions around which scenario is best suited for which types of environments.

15 GEOTHERMAL ENERGY↗

Graph-Based Representations and Applications to Process Simulation

Rapid and robust convergence of a process flowsheet is critical to enable large-scale simulations that address core scientific questions related to process design, optimization, and sustainability. However, due to the highly coupled and nonlinear nature of chemical processes, efficiently solving a flowsheet remains a challenge. In this work, we show that graph representations of the underlying physical phenomena in unit operations may help identify potential avenues to systematically reformulate the network of equations and enable more robust topology-based convergence of flowsheets. To this end, we developed graph abstractions of the governing equations of vapor-liquid and liquid-liquid equilibrium separation equipment. These graph abstractions consist of a mesh of interconnected variable nodes and equation nodes that are systematically generated through PhenomeNode, a new open-source library in Python developed in this study. We show that partitioning the graph into separate mass, energy, and equilibrium subgraphs can help decouple nonlinearities and guide decomposition algorithms. By employing the graph abstraction on an industrial separation process for separating glacial acetic acid from water, we implemented a new block decomposition scheme in BioSTEAM and demonstrated that this can accelerate convergence over a traditional sequential modular approach.

Distillation↗

Predicting metabolic modules in incomplete bacterial genomes with MetaPathPredict

The reconstruction of complete microbial metabolic pathways using ‘omics data from environmental samples remains challenging. Computational pipelines for pathway reconstruction that utilize machine learning methods to predict the presence or absence of KEGG modules in incomplete genomes are lacking. Here, we present MetaPathPredict, a software tool that incorporates machine learning models to predict the presence of complete KEGG modules within bacterial genomic datasets. Using gene annotation data and information from the KEGG module database, MetaPathPredict employs deep learning models to predict the presence of KEGG modules in a genome. MetaPathPredict can be used as a command line tool or as a Python module, and both options are designed to be run locally or on a compute cluster. Benchmarks show that MetaPathPredict makes robust predictions of KEGG module presence within highly incomplete genomes.

59 BASIC BIOLOGICAL SCIENCES↗

NLR HPC Kestrel Jobs Data

Overview: Anonymized job-level records from the Kestrel HPC system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, utilization, energy estimates, and efficiency metrics. Sensitive fields (user, account, job name, submit line, working directory, submit script, and job type) are replaced with 7-character cryptographic hashes. System & Timeframe: Kestrel is located at the NLR campus. Standard compute nodes have 104 cores and 256 GB RAM; bigmem nodes have 2,000 GB. GPU nodes (gpu-h100 partition) use NVIDIA H100 GPUs. Data covers jobs submitted August 2023 through December 2025. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.kestrel.job-anon.zip — Anonymized job records (Hive-partitioned Parquet) datacard.md — Full dataset documentation ~11 million rows, 50 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct with timezone-aware export (SLURM_TIME_FORMAT="%Y-%m-%dT%H:%M:%S%z"), loaded into PostgreSQL. Calculated columns updated via database triggers and batch functions. All timestamps use timestamptz and correctly handle DST transitions. Preprocessing: Anonymization of name, user, account, submit_line, work_dir, submit_script, and job_type via 7-char hex hashes Derived columns: queue_wait, cpu_eff, max/min/avg_mem_eff, energy estimates Simplified job state mapping (e.g., "CANCELLED by 132357" → "CANCELLED") Boolean flags: python_job, reframe_job Temporal decomposition: year, month, day, day_of_week, hour, minute from submit_time Shared node tracking: shared_job_count, nodes_shared, jobs_shared Key Variables: Scheduling: job_id, partition, state_simple, submit_time, start_time, end_time, queue_wait Resources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max/min/avg_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, consumed_energy_raw_joules, consumed_energy_raw_watt_hours Sharing: shared_job_count, nodes_shared, jobs_shared Partitions: short, standard, debug, gpu-h100 Job States: CANCELLED, COMPLETED, FAILED, PENDING, RUNNING QoS Levels: normal, high Important Notes: Timestamps include timezone offsets; DST transitions are handled correctly, though adding intervals across DST boundaries requires offset adjustment shared_job_count reflects physical node co-residency, not use of the shared partition Job step records and raw Slurm JSONB fields are excluded Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING↗

Replication Data for: Measurement of the mean number of muons with energies above 500 GeV in air showers detected with the IceCube Neutrino Observatory

<b>Measurement of the mean number of muons with energies above 500 GeV in air showers detected with the IceCube Neutrino Observatory</b> <br><br> This data release accompanies results submitted to Physical Review D describing the measurement of the average multiplicity of TeV muons with IceCube. It contains the data necessary to reproduce the main plots from the paper (Figs. 7 and 9), i.e. the numerical results for the average number of muons with energies above 500 GeV as a function of primary cosmic ray energy. <br><br> For any questions about this data release, please write to analysis@icecube.wisc.edu. <br><br> Files included in this release: <ul> <li>A README file <li>Files including data to reproduce the results plots from the paper (see below for details) <li>An example python script showing how to read and plot the data </ul> <br> <u>What is in the files icecube_Nmu500_X_Y.txt:</u> <br> Y indicates wether the file contains values obtained from experimental data (Y="data") or air-shower simulations (Y="MC"). <br> X indicates the hadronic interaction model for which the plot is made. If Y="data", this means that the experimental data was interpreted using this model. If Y="MC", it means that the simulations were performed with this model. The three models included are Sibyll 2.1, QGSJet-II.04, and EPOS-LHC (see paper for references). The file with X="modelaverage" gives the average over the three individual results with the deviations from the average included in the systematic uncertainties. <br><br> Please see the README file for details on how the data is structured in the files.

Astroparticle Physics↗

Data and code from: Multivariate bayesian regression model for predicting disposed ash composition at U.S. coal fired power stations

This dataset contains the code and data files needed for implementation of a Multivariate Bayesian Regression model, described in Jin et al. (2025), for the historical prediction of the chemical composition of disposed coal ash at U.S. coal fired power plants as a function of annualized coal purchase data. The integrated coal supply data file (CoalSupplyDataset.csv) represents a compilation of monthly fuel purchase records for the period 1973-2022 at major U.S. power stations. These records were obtained from the U.S. Energy Information Administration. The CSV file also contains, for each coal purchase record, the coal region of the mine as defined by the U.S. Geological Survey. Data entry errors and data gaps in the EIA records were corrected as described in Jin et al. This CSV file represents the integrated coal supply data after corrections were made. The model structure and fitting parameters are encoded in pickle file format (Bayesian.pkl). The model was developed with the coal supply data and coal ash composition data, apportioned according to the Stratified Shuffle Split for training and testing subsets. The model was built using Python and the PyMC library. Reference Publication: Jin, Z.; Huang, J.; Hower, J.C.; Hsu-Kim, H.(2025). Predictive Assessment of the Chemical Composition of Coal Ash in Reserve at U.S. Disposal Sites. Environmental Science & Technology.

Coal ash composition↗

Recording and wear characteristics of 4 and 8 mm helical scan tapes

Performance data of media on helical scan tape systems (4 and 8 mm) is presented and various types of media are compared. All measurements were performed on a standard MediaLogic model ML4500 Tape Evaluator System with a Flash Converter option for time based measurements. The 8 mm tapes are tested on an Exabyte 8200 drive and 4 mm tapes on an Archive Python drive; in both cases, the head transformer is directly connected to a Media Logic Read/Write circuit and test electronics. The drive functions only as a tape transport and its data recover circuits are not used. Signal to Noise, PW 50, Peak Shift and Wear Test data is used to compare the performance of MP (metal particle), BaFe, and metal evaporate (ME). ME tape is the clear winner in magnetic performance but its susceptibility to wear and corrosion, make it less than ideal for data storage.

Peter, Klaus J.↗