Search NASA⌕ Search

SEARCH · Search NASA

Results for “CAN data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Data and script associated with “Shifts in Rain-Snow Partitioning Drive Faster Water Transit Times in the US Pacific Northwest”

This data package contains the data and code to use and run the Water Tracer enabled version of the Weather Research and Forecasting Hydrologic model (WT-WRF-Hydro) with the Sequential Precipitation Input Tagging (SPIT) framework. It is associated with the publication “Shifts in Rain-Snow Partitioning Drive Faster Water Transit Times in the US Pacific Northwest” published in Scientific Reports (Butler et al., 2026; https://doi.org/10.1038/s41598-026-46539-1). We use the Continental U.S. (CONUSII; Rasmussen et al., 2021) dataset to force the model with an historical climate (2006–2013) and a future climate (2086–2093) with a representative carbon pathway (RCP) 8.5 scenario. We use the model to calculate water transit times in five headwater catchments within the U.S. Pacific Northwest. We also show key hydrologic and environmental variables that affect water transit times and changes in the future. Finally, we use observed data to validate the model such as stream water isotopes, snowpack characteristics, and stream discharge. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. The data package consists of 11 folders: (1) "Figures" contains the exported figures used in the manuscript; (2) "Model_Isotope_Date" contains the WT-WRF-Hydro isotope date used in model validation; (3) “Model_Outputs_Future” contains the WT-WRF-Hydro future climate outputs; (4) “Model_Outputs_Historical” contains the WT-WRF-Hydro historical climate outputs; (5) “Model_Outputs_Weights_Areas” contains the WT-WRF-Hydro weights per catchment used to calculate water transit times and isotopes in stream water; (6) “MODIS_data_scripts” contains data used to validate snow conditions in the study area; (7) “Observed_Flow_Data” contains the observed streamflow data used in model validation; (8) “Observed_Isotope_Data” contains the observed stream water isotope data used in model validation; (9) “Scripts” contains the Python scripts used to general results and the figures; (10) “Statistic_Outputs” contains the water transit time statistical outputs reported in this manuscript; (11) “Validation_SNOTEL” contains the SNOTEL data used in model validation. The files in this data package have the following file extensions: .tif, .txt, .csv, .pdf, .py, .jpg, and .png.

American River↗

IM3 + EPRI Data Center Load Projections

This dataset contains scenarios of hourly total electricity demand with and without projected loads from data centers over the period 2022-2040. The root projections without data center demands are identical to those documented in Burleyson et al. 2024. In short, those projections encompass hourly electricity demands for 54 Balancing Authorities (BAs) in the United States across a range of eight of weather and socioeconomic scenarios. Refer to the root dataset and accompanying publication, Burleyson et al. 2025, for information about how those projections were generated. For this derivative dataset we used the base loads from the following scenarios: rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 The root load projections did not reflect the drastic expansion of data centers that has occurred in the last several years to support artificial intelligence and cloud computing. To reflect growth in data center demand, a second set of load projections were created in which we layered in additional data center load projections based on the data center load growth scenarios described in a 2024 report by the Electric Power Research Institute (EPRI): "Powering Intelligence: Analyzing Artificial Intelligence and Data Center Energy Consumption". The EPRI projections from the report are included in this dataset (EPRI_2024_Projections.xlsx). That report contained annual state-level data center load projections for four year-over-year growth rates for data center demands: Low (3.71% annual growth) Moderate (5% annual growth) High (10% annual growth) Higher (15% annual growth) To homogenize the load projections with and without data centers we had to get them to a common scale. The first step was to take the EPRI annual state-level data center energy consumption values and convert them to 8760-hr loads for each year. We did that by assuming a flat (e.g., not weather- or time-sensitive) load profile and distributing the data center loads in each state evenly across all hours in a year. From there the loads were downscaled from the state-level to the county-level using 2019 county-level populations as weights. Finally, the county-level hourly data center loads were summed to the BA-level using the county-to-BA mapping underpinning the root load projections. The net result is 16 (4 weather and socioeconomic scenarios crossed with 4 data center load growth scenarios) unique load projections for the period 2022-2040. The file format follows that of the root dataset with a single additional column "Scaled_TELL_BA_Load_with_DC_MWh" that contains the hourly loads with the added data center loads for a given BA-year-scenario combination. Please refer to the readme file in the root dataset for more information on the file format.

Burleyson, Casey [Pacific Northwest National Labor↗

Electricity Baseline 2021 Background Data and Log File

The ElectricityLCI v2 Python package (https://github.com/USEPA/ElectricityLCI/tree/v2.0) was used to generate the 2021 electricity baseline: a regionalized life cycle inventory model of U.S. electricity generation, consumption, and distribution using standardized facility and generation data. ElectricityLCI implements a local data store for downloading and accessing public data on an individual's computer. The data store follows the folder definition provided by USEPA's esupy Python package (https://github.com/USEPA/esupy), which utilizes the appdirs Python dependency (https://pypi.org/project/appdirs/). An overview of the ElectricityLCI data stores may be found on the README (https://github.com/USEPA/ElectricityLCI/blob/v2.0/README.md#data-store). This submission includes the background data used to generate the 2021 electricity baseline inventory. Each zip archive stores the source files as found in their data stores. Sub-folders in each of the data stores are archived separately. For example, stewi.zip contains the JSON files, while stewi.facility.zip is the 'facility' sub-folder of stewi data store that stores the parquet files. To reproduce the data store, extract each zip file and drag-and-drop sub-folders in to their appropriate root folders to recreate the data stores, then copy the root folders to your data store folder (as returned by running the following on the command line: python -c "import appdirs; print(appdirs.user_data_dir())"). The main five data stores include: 'electricitylci', 'facilitymatcher', 'fedelemflowlist', 'stewi', and 'stewicombo'. The log file generated by the 2021 model run is also included, which contains the statements at the DEBUG level and above.

Electricity; LCA; LCI; Life Cycle; data inventory↗

Electricity Baseline 2020 Background Data and Log File

The ElectricityLCI v2 Python package (https://github.com/USEPA/ElectricityLCI/tree/v2.0) was used to generate the 2020 electricity baseline: a regionalized life cycle inventory model of U.S. electricity generation, consumption, and distribution using standardized facility and generation data. ElectricityLCI implements a local data store for downloading and accessing public data on an individual's computer. The data store follows the folder definition provided by USEPA's esupy Python package (https://github.com/USEPA/esupy), which utilizes the appdirs Python dependency (https://pypi.org/project/appdirs/). An overview of the ElectricityLCI data stores may be found on the README (https://github.com/USEPA/ElectricityLCI/blob/v2.0/README.md#data-store). This submission includes the background data used to generate the 2020 electricity baseline inventory. Each zip archive stores the source files as found in their data stores. Sub-folders in each of the data stores are archived separately. For example, stewi.zip contains the JSON files, while stewi.facility.zip is the 'facility' sub-folder of stewi data store that stores the parquet files. To reproduce the data store, extract each zip file and drag-and-drop sub-folders in to their appropriate root folders to recreate the data stores, then copy the root folders to your data store folder (as returned by running the following on the command line: python -c "import appdirs; print(appdirs.user_data_dir())"). The main five data stores include: 'electricitylci', 'facilitymatcher', 'fedelemflowlist', 'stewi', and 'stewicombo'. The log file generated by the 2020 model run is also included, which contains the statements at the DEBUG level and above.

Electricity; LCA; LCI; Life Cycle; data inventory↗

3D Deep Learning Joint Inversion of Active Seismic Full Waveform and Passive Seismic Traveltime Data for Reservoir Imaging and Uncertainty Quantification

Here, we present deep learning (DL) networks for three-dimensional (3D) joint inversion of active seismic full waveform and passive seismic traveltime data to image reservoirs and their properties and quantify imaging uncertainties. Active seismic full-waveform data can provide high-resolution monitoring images but are collected only intermittently because of their high acquisition cost. In contrast, passive seismic data can be gathered at relatively low cost between regular active surveys, although their imaging quality can be compromised by factors such as low signal-to-noise ratios and limited ray coverage of the target. Although these datasets are routinely acquired together at CO 2 storage sites, their combined inversion within a 3D DL framework has not been previously demonstrated. To our knowledge, this is the first study to address this gap, combining the strength of both data types. For efficient data storage and DL training with large 3D seismic datasets, we use a 3D data matrix in which a random number of passive seismic traveltime data are stored as parabolic envelopes using one-hot encoding and a 3D full-waveform data matrix in which multiple shot gathers are summed. Two network architectures are evaluated: a single-encoder U-Net for single-data type inversion and a dual-encoder U-Net for joint inversion of active and passive seismic data. We also evaluate the single-encoder U-Net for joint inversion by concatenating full-waveform data and traveltime data. We propose a systematic approach for selecting an optimal dropout rate that balances regularization during training and Monte Carlo dropout-based uncertainty quantification during prediction by examining the correlation coefficient between standard deviation and prediction error, along with the training misfit, across a range of dropout rates. 3D DL inversion experiments include five different network configurations, with evaluations under ideal, noisy and dropout-enabled conditions. Both model and data uncertainties are assessed, as well as their combined effects. Across all conditions, the networks consistently predict accurate CO 2 saturation models with low prediction errors, such as a structural similarity index of 0.993 and CO 2 difference of 1.1%. Uncertainty estimates show strong spatial correlation with prediction errors, confirming the effectiveness of the proposed dropout selection approach. The results demonstrate that our DL approach, utilizing compact data representations and appropriate uncertainty quantification, yields accurate subsurface images under various inversion conditions and provides valuable insights into the reliability of predictions.

Um, Evan Schankee [Lawrence Berkeley National Labo↗

Data From: "Warming and snow loss increase reliance on old groundwater in a Colorado River headwater"

This repository contains the data and code associated with the paper titled "Warming and snow loss increase reliance on old groundwater in a Colorado River headwater," published in Nature Geoscience, 2026. This study seeks to answer how various ages of groundwater interact with mountainous streamflow in mountainous headwaters such as the East River. It includes various model-data processing scripts, primarily for ParFlow-CLM analysis of simulated water years 2015-2021, and two numerical warming experiments (+2.5 and +4.0 degrees C), including run scripts, forcing scripts, and post-processing, as well as comparison to observation datasets, detailed below. This data requires the use of R (.r, .rmd), Python (.py), Jupyter Notebook or Jupyter Lab (.ipynb), ParFLOW-CLM, EcoSLIM. Further information on the use of all file formats mentioned below (e.g. .tff. .nc) are provided within the associated scripts and directory where the files are located. Contents & Usage ASO/: ​​Contains the bash and python scripts used to convert airborne snow observatory (ASO) data (ASO, 2023) in various data formats (georeferenced tiff file, NetCDF, UTM, and to latitude/longitude) then regrided to the ParFlow equivalent grid. Output data are in regrid_regll_data.zip and subsequently visualized and analyzed in plot_and_compare.py for Supplementary Figures A14 and A15. The wksht_ASO_comparison.xlsx spreadsheet is used to calculate the data for Supplementary Figure A16. EcoSLIM/: Contains the scripts and input files to run the EcoSLIM particle tracking simulations (/run_scripts) and the post-processing python script (/plot_scripts/eco_agedist_plots.ipynb). Jasechko et al./: Contains the jupyter notebook (Extract_Elevation.ipynb) to determine the outlet elevations of the 260 watersheds used in Jasechko et al. (2016), and the corresponding table, Table_S1_Watersheds_alt.csv. Used to create Supplementary Information Figure A2. PLM_Wells/: Contains the QA/QC-ed groundwater level time series of the PLM-1 and PLM-6 Monitoring Wells from Faybishenko et al. (2023), reformatted to water years used for Supplementary Figures A19 and and A20. ParFlow/: Contains the input files and run scripts to run ParFlow-CLM (/run_scripts), the python and tool command language (Tcl) scripts to create and distribute the ParFlow forcing simulation files (/forcing), and various scripts and intermediary files to analyze the model outputs (/post_process). SQUIRE/: Contains the processing scripts and intermediary files for the Surface QUantitatIve pRecipitation Estimation (SQUIRE) data (Grover, 2023) used to generate Supplementary Figure A18. USGS_Streamflow/: Contains the raw and gap-filled United States Geological Survey streamflow data (U.S. Geological Survey, 2026) used at the Almont station (site number 09112500). Gap-filling is performed in the R script with data from the Taylor station (site number 09110000). (/USGS_09112500_EAST_RIVER_AT_ALMONT_GAP_FILLED/code_almont_streamflow_gap_fill.Rmd). discharge/: Contains the gap-filled discharge data at the Watershed Function SFA East River pumphouse site (Newcomer et al., 2022) used to generate Supplementary Figure A13 and to compute hourly Nash-Sutcliffe model efficiency coefficients (NSE) in Table A4. snotel_and_flux_tower/: Contains the snow telemetry data (U.S. Department of Agriculture, 2024) from the Butte (site ID 380) and Schofield (site ID 737) stations, reformatted by water year, accessed with the snotelr R package. Used to create Supplementary Figure A17. Also contains the flux tower observational data (FluxTower_Pumphouse_ESS-DIVE.ET_only.h.txt) from Ryken et al. (2022) and sap flux transpiration data (MaxB_Transpiration_5Sites.daily_sums.h.txt) from Ryken (2021), used to create Supplementary Figures A22 and A23, respectively. Raw EcoSLIM model outputs are in excess of 24TB, and are stored on National Energy Research Scientific Computing Center (NERSC) and publicly available via the external link provided in the paper.

atmospheric warming↗

COMPASS-FME Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) Experiment Level 2 Sensor Data v2-1

This is the version v2-1 Level 2 (L2) data release for COMPASS-FME environmental sensors located at our Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) experimental site. This manipulative, ecosystem-scale TEMPEST experiment addresses the potential for freshwater and estuarine-water disturbance events to alter tree function, species composition, and ecosystem processes in a deciduous coastal forest in MD, USA. The experiment uses a large-unit (2000 m2), un-replicated experimental design, with three 50 m × 40 m plots serving as control, freshwater, and estuarine-water treatments. Level 2 (L2) data consist of sensor observations from the COMPASS-FME synoptic sites, TEMPEST, and DELUGE. Compared to the L1 data, these are more consistent (always 15-minute timestamps for the entire year); better QA/QC’d (out of bounds, out of service, and extreme outlier values are removed); and more complete, with a gap-filled time series available alongside the main observations, and additional derived (calculated) variables. L2 data are intended to be rapidly and easily usable in analyses and simulations. However, algorithmic outlier identification always carries the risk of removing valid data, and Level 1 data may be more suitable for analyses that focus on variability or extreme events. This dataset includes: - An overall dataset README file that describes the current version, gives citation and contact information, etc. - Site- and year-specific folders, each holding variable-specific Parquet (a high performance, space efficient format; see https://parquet.apache.org) data files for each site and plot in that year. - Metadata files within each site-year folder provide full information on data units, expected ranges, contact information, detailed flood times, as well as a general description of the site. - Environmental sensor types that appear in the data files include weather (ClimaVUE50, CS, RM Young, and LI instruments in the graphs below); soil conditions (TEROS12); soil redox state (Redox); groundwater variables (AquaTROLL200 and AquaTROLL600); open water sondes (Exo); tree sap velocity (Sapflow); and system voltage and state (Datalogger). Data are reported every 15 minutes. Please see v2-1 TEMPEST L2 Sensor Package Quick Start.pdf for detailed information on data package structure, temporal coverage, and versioning. Data files are in Apache Parquet, a high performance, space efficient format for tabular data. These files can be read using R's `arrow` package (https://arrow.apache.org/docs/r/), with similar tools available in other languages. The TEMPEST flood events occurred on the following dates. They lasted for ~10 hours each day and delivered ~80,000 gallons to each plot; many data streams are available at 1 or 5 minute frequency during these periods. * Tests: Aug 25 (fresh plot) and Sep 9 (salt plot), 2021 * TEMPEST 1: June 22, 2022 * TEMPEST 2: June 6-7, 2023 * TEMPEST 3: June 11-13, 2024

EARTH SCIENCE > ATMOSPHERE > ATMOSPHERIC TEMPERATU↗

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis - NLR Historical Solar PV

The U.S. Department of Energy and National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis from variable sources, hydrogen compression and storage, and hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) research platform. This dataset represents part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure nationwide. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with various energy technologies. Future datasets will demonstrate how existing hydrogen fuel cell technologies can provide controllable, dispatchable, and variable power output for artificial intelligence data centers and other variable loads. This dataset entry describes the behavior of a 1.25-MW proton exchange membrane MC250 electrolyzer system, manufactured by Nel Hydrogen , [1] when fed historical data generated by the 430-kW, fixed-axis solar photovoltaic (PV) array located at NLR’s Flatirons Campus. (While the electrolyzer balance of plant supports up to 2.5 MW of electrolysis, NLR only has a single 1.25-MW electrolysis stack.) Solar PV power output data for the 2020 calendar year were categorized on a daily basis by total energy generation and standard deviation. Each day was then ranked by these metrics, and the 25th, 50th, and 100th percentiles were selected. The 75th percentile day did not exhibit sufficient variability to make for a valuable experiment. A similar process was used for the related historical wind dataset . [2] The historical days in 2020 that represented these percentiles are Dec. 19, March 29, and May 4, respectively. The entire solar day’s power profile was then fed through the MC250 electrolyzer. Due to its length, the 100th percentile day experiment was split into two parts, and the final 3 hours of the solar day were not captured. These final 3 hours contained no spikes or dips of interest and simply represented a slow decay of input solar power. Also, a single timestamp (13:13:47 on Jan. 14, 2026) was lost in the hydrogen system supervisory control and data acquisition. Finally, during the 25th percentile experiment (solar day Dec. 19, 2020) data recording was lost from 11:00:13 to 11:14:45. The roughly 15 minutes of the solar profile were rerun at the end of the experiment and spliced into this time slot during post-processing. The electrolysis system controls hydrogen production by varying direct current applied to the stack, from a maximum of 3,000 A to a minimum safe operation of 300 A, or 10%. Because the current–voltage characteristic changes as the stack ages and efficiency degrades, the actual minimum safe operating power changes over time. The historical solar profiles were translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1-Hz frequency. For more details on the statistical analysis process, see the slide deck “Public Reference Data for Megawatt-Scale Hydrogen Electrolysis: NLR Historical Solar PV Analysis and Profile Generation” accessible with this data entry. These datasets report relevant hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. Each .zip file represents a single solar PV electrolysis experiment and is formatted as: {technology}_{percentile}_{scaling factor} For instance, “solarPV-430kW_25_2x.zip” reports the experiment using the 25th percentile solar data from the historical 2020 solar PV dataset, scaled to 200%. Scaling factors were applied to the generated solar PV power output files to more closely match the 1.25-MW capacity of the electrolyzer. Each .zip folder contains the following files: A .csv file containing raw data. An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production, electrolysis power consumption, and solar power input. A PDF file detailing the historical solar data statistical analysis used to generate the solar profile. An experiment labeled “characterization_200.zip” demonstrates the MC250 electrolyzer steady-state response with 30-minute load steps for a total duration of 5 hours. Finally, a .csv file is provided with all experiments combined into one dataset labeled "combined_solarPV_experiments.csv". [1] nelhydrogen.com/product/mc-series-electrolyser . [2] data.nlr.gov/submissions/316 .

08 HYDROGEN↗

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

Impact of recent ENDF nuclear data update, high initial enrichment and high burnup fuel on critical experiments applicability determination via the integral index c k for burnup credit validation

In 2012, NUREG/CR-7109 reported on the validation of burnup credit calculations involving major and minor actinides and major fission products which was investigated for pressurized and boiling water reactor (PWR and BWR) fuel enrichments up to 5 wt% 235 U and assembly-average burnups up to 60 GWd/MTU. Recently, there has been interest in increasing the maximum enrichment used in PWR fuel as high as 8 wt% 235 U and correspondingly increasing the maximum assembly-average burnups to approximately 75 GWd/MTU. These proposed increases in enrichment and burnup necessitate reinvestigation of the validation basis for k eff calculations for this expanded application space. Additionally, the 2012 study was performed by using the Evaluated Nuclear Data File (ENDF)/B-VII.0 nuclear data with the SCALE 6 covariance library, and the effects of using the newly released ENDF/B-VII.1 and ENDF/B-VIII.0 nuclear data and covariance libraries should be evaluated. In this work, published in NUREG/CR-7309 in 2025, the validation assessment was performed consistently with NUREG/CR-7109: modeling irradiated fuel assemblies in the Generic Burnup Credit (GBC)-32 cask defined in NUREG/CR-6747. The TSUNAMI-3D sequence was used to generate sensitivity data for the application model, and the data were compared with sensitivity data from select benchmark models. The integral parameter c k is the metric of similarity used in this study and is consistent with NUREG/CR-7109, where a c k value in excess of 0.8 indicates sufficient similarity for use in validation. A new set of benchmark experiments with sensitivity data has been assembled for this effort. The number of experiments with available sensitivity data is now 2,104, compared to 474 in NUREG/CR-7109. This increase was facilitated by the efforts of the Nuclear Energy Agency to generate sensitivity data for a majority of the experiments in the International Criticality Safety Benchmark Evaluation Project (ICSBEP) Handbook to supplement the data available in the Oak Ridge National Laboratory (ORNL) Verified, Archived Library of Inputs and Data (VALID). The complete set of benchmarks considered here includes experiments for low-enriched uranium (LEU), intermediate enriched uranium (IEU), and a mixture of uranium and plutonium (MIX) from the ICSBEP Handbook and VALID, as well as ORNL models of the Haut Taux de Combustion (HTC) experiments and other potentially relevant models not included in VALID. The updated similarity study shows that none of the extended burnup and higher enrichment combinations considered show a significant decrease in the number of potentially applicable experiments, meaning sufficient critical experiments exist for the validation of BUC criticality safety calculations, with initial enrichments up to 8 wt% 235 U and burnups up to 80 GWd/MTU. Additionally, both the ENDF/B-VII.1 and ENDF/B-VIII.0 nuclear data libraries can be used for validation since the number of critical experiments applicable for validation increases for most cases with the most recent nuclear data compared to the previous one. As in previous BUC validation studies, the French HTC experiments are the most similar in a majority of the application cases studied, especially from representative discharge burnups ranging from 40 to 80 GWd/MTU. In conclusion, these results match the conclusions presented in NUREG/CR-7109 regarding validation of the primary actinides in BUC analyses.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Exploring Ion Mobility Mass Spectrometry Data File Conversions to Leverage Existing Tools and Enable New Workflows

Ion mobility (IM) is often combined with LC-MS experiments to provide an additional dimension of separation for complex sample analysis. While highly complex samples are better characterized by the full dimensionality of LC-IM-MS experiments to uncover new information, downstream data analysis workflows are often not equipped to properly mine the additional IM dimension. For many samples the data acquisition benefits of including IM separations are all that is necessary to uncover sample information and the full dimensionality of the data is not required for data analysis. Post-acquisition reduction and adaptation of the dimensions of LC-IM-MS and IM-MS experiments into an LC-MS format opens the possibility to use a plethora of existing software tools. In this work, we developed data file conversion tools to reduce the complexity of IM data analysis. Three data file transformations are introduced in the PNNL PreProcessor software: 1) mapping the IM axis to the LC axis for IM-MS data, 2) converting the drift time vs. m/z space to CCS/z vs m/z space, and 3) transforming All Ions IM/MS mobility aligned fragmentation data to a standard LC-MS DDA data file format. Finally, these new data file conversions are demonstrated with corresponding lipidomics and proteomics workflows that leverage existing LC-MS data analysis software to highlight the benefits of the data transformations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Optimizing Management of Persistent Data Structures in High-Performance Analytics

Large-scale data analytics workflows ingest massive input data into various data structures, including graphs and key-value datastores. These data structures undergo multiple transformations and computations and are typically reused in incremental and iterative analytics workflows. Persisting in-memory views of these data structures enables reusing them beyond the scope of a single program run while avoiding repetitive raw data ingestion overheads. Memory-mapped I/O enables persisting in-memory data structures without data serialization and deserialization overheads. However, memory-mapped I/O lacks the key feature of persisting consistent snapshots of these data structures for incremental ingestion and processing. The obstacles to efficient virtual memory snapshots using memory-mapped I/O include background writebacks outside the application’s control, and the significantly high storage footprint of such snapshots. To address these limitations, we present Privateer, a memory and storage management tool that enables storage-efficient virtual memory snapshotting while also optimizing snapshot I/O performance. Here, we integrated Privateer into Metall, a state-of-the-art persistent memory allocator for C++, and the Lightning Memory-Mapped Database (LMDB), a widely-used key-value datastore in data analytics and machine learning. Privateer optimized application performance by 1.22× when storing data structure snapshots to node-local storage, and up to 16.7× when storing snapshots to a parallel file system. Privateer also optimizes storage efficiency of incremental data structure snapshots by up to 11× using data deduplication and compression.

Computer science↗

Data and Scripts associated with a manuscript on ecosystem responses to wildfires in the Columbia River Basin

This data package is associated with the publication “Ecosystem leaf area, gross primary production, and evapotranspiration responses to wildfire in the Columbia River Basin” submitted to Biogeosciences (Shi et al., 2024; doi: 10.22541/au.171053013.30286044/v1). In this research, data products, leaf area index (LAI), gross primary production (GPP), and evapotranspiration (ET), from the Moderate Resolution Imaging Spectroradiometer (MODIS) are used to quantify the resistance and resilience of different ecosystem types in the Columbia River Basin (CRB). A machine learning algorithm, random forest (RF), was used to examine the impacts of precipitation, vapor pressure deficit (VPD), and burn severity from Monitoring Trends in Burn Severity (MTBS) on ecosystem resilience. The data package includes the processed MODIS data products, precipitation, VPD, and burn severity in 138 fire regions in CRB and the input files for RF model training. This data package includes six folders. The MODIS products are included in three MODIS_* folders with shell scripts for data clipping and *ncl files for data processing: (1) “/MODIS_LAI_CRB”; (2) “/MODIS_GPP_CRB”; and (3) “/MODIS_ET_CRB”. All the processed data for each fire event are NetCDF formatted. The MTBS burn severity data and the shell and *ncl scripts used for data processing are in the folder named (4) “MTBS_fire”. The ERA meteorological fields and the data processing scritps are in (5) “ERA_Var_CR”. All the scripts for figure development are in the format of *ncl and in the folder (6) “paper_scripts”. See the file ending in “flmd.csv” for a list of all files contained in this data package and descriptions for each. Tabular column headers and units are described in the data dictionary file ending in “dd.csv”.

54 ENVIRONMENTAL SCIENCES↗

Electricity Baseline 2022 Background Data and Log File

The ElectricityLCI v2 Python package (https://github.com/USEPA/ElectricityLCI/tree/v2.0) was used to generate the 2022 electricity baseline: a regionalized life cycle inventory model of U.S. electricity generation, consumption, and distribution using standardized facility and generation data. ElectricityLCI implements a local data store for downloading and accessing public data on an individual's computer. The data store follows the folder definition provided by USEPA's esupy Python package (https://github.com/USEPA/esupy), which utilized the appdirs Python dependency (https://pypi.org/project/appdirs/). This submission includes the background data used to generate the 2022 electricity baseline inventory. Each zip archive stores the source files as found in their data stores. Sub-folders in each of the data stores are archived separately. For example, stewi.zip contains the JSON files, while stewi.facility.zip is the 'facility' sub-folder of stewi data store that stores the parquet files. To reproduce the data store, extract each zip file and drag-and-drop sub-folders in to their appropriate root folders to recreate the data stores, then copy the root folders to your data store folder (as returned by running the following on the command line: `python -c "import appdirs; print(appdirs.user_data_dir())"`). The main five data stores include: 'electricitylci', 'facilitymatcher', 'fedelemflowlist', 'stewi', and 'stewicombo'. The log file generated by the 2022 model run is also included, which contains the statements at the DEBUG level and above.

Electricity; LCA; data inventory↗

Potential of Data Center Controls in Grid Services

The rapid proliferation of large data centers brings both challenges and opportunities for grid reliability. The data center resources and their potential flexibility have the potential to contribute resources to grid operations. Through capabilities like energy shifting and resource coordination, data centers can help reduce their net demand on the transmission network, as well as provide additional grid services to support reliable operation on the grid. While transient and long-term grid planning and operations are the scenarios that draw most attention, the quasi-steady state timeseries (QSTS) operation of data centers and grid bring interesting scenarios that can help evaluate the data center controls to aid grid services. This work is focused on modeling data centers for QSTS applications – incorporating the AI data center load profiles and building on the PNNL digital twin model for the thermal management loads to enable simulation studies to reveal the impact of data center controls on grid performance. This includes the integration of a QSTS battery and natural gas generator model to incorporate local resource impacts to the system. The simulation study is performed with a modified IEEE 24-Bus transmission system. Scenarios are focused on evaluating the data center load impacts on the transmission system and leveraging both data center and local generation controls to mitigate those impacts and provide additional grid services. The data center controls revealed the ability to contribute to two main kinds of grid services: preventing congestion on a weak grid by coordinating the data center resources with the collocated BESS and onsite generation; and the ability to help the grid operations during stressed times of operation like during a contingency. Leveraging these and other capabilities has the potential to help data centers become grid responsive assets, aiding in both their integration into the power system and grid reliability.

power grid simulation↗

Aligning NASA Earth Science Data Stewardship with FAIR Principles: Outcomes, Recommendations, and Future Directions

The FAIR Principles—Findable, Accessible, Interoperable, and Reusable—offer a widely accepted framework for improving the sharing and reuse of digital scientific data by both human and machine users. Following these principles is critical for effective scientific data stewardship, broader scientific collaboration, and compliance with federal and agency data policies. This paper, based on the work of NASA’s Open, Free, and FAIR Working Group (O’FAIR WG) under the Earth Science Data Systems Program, presents an overview of how FAIR is being applied within NASA’s Earth science data landscape. It highlights ongoing progress and challenges, identifies FAIR-enabling resources, and offers recommendations and strategic actions to enhance the FAIRness of NASA-funded open and free Earth science data products. The FAIR-enabling resources identified underscore the vital role of NASA's existing enterprise processes, standards, tools, and infrastructures in supporting FAIR implementation. Our findings show strong performance in making NASA Earth science data more findable and accessible. However, further work is needed—especially in enhancing interoperability, so that different systems and tools can better understand and exchange data. This is especially important for enabling machine-driven discovery and analysis. We emphasize the importance of a balanced strategy that combines a centralized, top-down approach—focused on building enterprise-level capabilities and processes—with a decentralized, bottom-up approach driven by discipline-specific needs and community practices. We advocate for coordinated efforts to enhance (meta)data interoperability to facilitate seamless data and information sharing and exchange of Earth science data both within NASA and across other agencies managing Earth science data.

Data Product↗

LinkML: an open data modeling framework

Background Scientific research relies on well-structured, standardized data; however, much of it is stored in formats such as free-text lab notebooks, nonstandardized spreadsheets, or data repositories. This lack of structure challenges interoperability, making data integration, validation, and reuse difficult. Findings LinkML (Linked Data Modeling Language) is an open framework that simplifies the process of authoring, validating, and sharing data. LinkML can describe a range of data structures, from flat, list-based models to complex, interrelated, and normalized models that utilize polymorphism and compound inheritance. It offers an approachable syntax that is not tied to any one technical architecture and can be integrated seamlessly with many existing frameworks. The LinkML syntax provides a standard way to describe schemas, classes, and relationships, allowing modelers to build well-defined, stable, and optionally ontology-aligned data structures. Once defined, LinkML schemas may be imported into other LinkML schemas. These key features make LinkML an accessible platform for interdisciplinary collaboration and a reliable way to define and share data semantics. Conclusions LinkML helps reduce heterogeneity, complexity, and the proliferation of single-use data models while simultaneously enabling compliance with FAIR (Findable, Accessible, Interoperable, and Reusable) data standards. LinkML has seen increasing adoption in various fields, including biology, chemistry, biomedicine, microbiome research, finance, electrical engineering, transportation, and commercial software development. In short, LinkML makes implicit models explicitly computable and allows data to be standardized at their origin. LinkML documentation and code are available at https://linkml.io/.

AI-ready data↗

Data from TropiRoot 1.0 database: tropical root characteristics across environments

TropiRoot 1.0 is a new tropical root database with root characteristics across environment gradients. It has data extracted from 104 new sources, resulting in more than 8000 rows of data (either species or community data). Most of the data in TropiRoot 1.0 includes root characteristics such as root biomass, morphology, root dynamics, mass fraction, architecture, anatomy, physiology and root chemistry. This initiative represents an approximately 30% increase in the currently available data for tropical roots in the Fine Root Ecology Database (FRED). TropiRoot 1.0, contains root characteristics from 25 different countries where seven are located in Asia, six in South America, five in Central America and the Caribbean, four in Africa, two in North America, and 1 in Oceania. Due to the volume of data, when ancillary data was available, including soil data, these data was either extracted and included in the database or their availability was recorded in an additional column. Multiple contributors checked the entries for outliers during the collation process to ensure data quality. For text-based observations, we examined all cells to ensure that their content relates to their specific categories. For numerical observations, we ordered each numerical value from least to greatest and plotted the values, checking apparent outliers against the data in their respective sources and correcting or removing incorrect or impossible values. Some data (soil and aboveground) have different columns for the same variable presented in different units, including originally published units, but root characteristics data had units converted to match the ones reported in FRED. By filling a gap from global databases, TropiRoot 1.0 expands our knowledge of otherwise so far underrepresented regions, and our ability to assess global trends. This advancement can be used to improve tropical forest representation in vegetation models.

54 ENVIRONMENTAL SCIENCES↗