Search NASA⌕ Search

SEARCH · Search NASA

Results for “gridded data products”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

ROOT’s RNTuple I/O Subsystem: The Path to Production

The RNTuple I/O subsystem is ROOT’s future event data file format and access API. It is driven by the expected data volume increase at upcoming HEP experiments, e.g. at the HL-LHC, and recent opportunities in the storage hardware and software landscape such as NVMe drives and distributed object stores. RNTuple is a redesign of the TTree binary format and API and has shown to deliver substantially faster data throughput and better data compression both compared to TTree and to industry standard formats. In order to let HENP computing workflows benefit from RNTuple’s superior performance, however, the I/O stack needs to connect efficiently to the rest of the ecosystem, from grid storage to (distributed) analysis frameworks to (multithreaded) experiment frameworks for reconstruction and ntuple derivation. With the RNTuple binary format soon arriving at its first production release, we present RNTuple’s feature set, integration efforts, and its performance impact on the time-to-solution. We show the latest performance figures of RDataFrame analysis code of realistic complexity, comparing RNTuple and TTree as data sources. We discuss RNTuple’s approach to functionality critical to the HENP I/O (such as multithreaded writes, fast data merging, schema evolution) and we provide an outlook on the road to its use in production.

Blomer, Jakob↗

The Role of the U.S. Electric Distribution System in Serving Data Center and Other Large Loads

The rapid expansion of data centers in the United States is reshaping how the electric distribution system must plan for and accommodate large load interconnections. This report evaluates the role of the distribution grid in serving these loads, from small edge facilities to hyperscale campuses. Using national datasets, utility filings, and industry studies, we assess demand growth, reliability requirements, interconnection thresholds, and infrastructure needs at substations and feeders. The analysis highlights the mismatch between fast data center development timelines and slower utility planning and construction cycles, as well as strategies such as phased energization, on-site generation, hosting capacity maps, and structured interconnection frameworks. While focused on data centers, the insights also apply to other high-density loads such as advanced manufacturing, hydrogen production, and electrified transportation. The report concludes with approaches to align planning processes, transparency tools, and regulatory frameworks so utilities can manage new large loads in ways that support a reliable and resilient grid.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Optimal Pathways from Alternative Carbon Feedstocks to Organic Commodity Chemicals

The use of biogenic and waste feedstocks is a promising strategy to improve the chemical sector's supply chain resiliency and carbon intensity. To help inform research efforts that transform these feedstocks into industrial chemicals, we used a systematic analysis framework to consistently evaluate the economics and environmental impacts of >200 alternative production pathways for 51 organic commodity chemicals in the United States under an optimistic future scenario that reflects the potential upper bounds of process scalability, energy availability, and carbon uptake. Lower-impact and lower-cost alternative pathways were identified for all but three chemicals, with 75% using thermochemical routes and half leveraging existing manufacturing infrastructure. Scenario analysis shows that the ranking of these pathways for half of the assessed chemicals is particularly sensitive to carbon uptake assumptions and criteria prioritization (i.e., cost only, environmental impact only, or both), with changes in electricity grid mix, hydrogen source, and underlying mass and energy flow data proving less influential. Implementing alternative pathways for just 11 chemicals could support a transition to net-zero greenhouse gas emissions from chemical production by 2050, with 11% lower cost than business as usual, similar water requirements, quadrupled electricity demand, and the use of most available woody biomass. These findings provide an exploratory guide toward a future chemical industry that harnesses alternative feedstocks.

09 BIOMASS FUELS↗

WE-Validate: An Open-Source Framework For Wind Power Validation

Grid operators rely on historical weather time series at existing and planned wind power plants to make informed decisions when planning for a future power grid with very high penetration of renewable power. While synthetic wind power time series have been developed based on historical weather models, their validation with actual power production data remains complex due to variations in modeling practices and methodologies. This paper introduces the WE-Validate framework, originally designed for wind speed validation and now enhanced for wind power validation with a graphical user interface to support users with minimal programming experience. Validation of wind power with WE-Validate is based on robust metrics consisting of RMSE, centered RMSE, average bias, average percent bias, mean absolute error, mean absolute percent error, cross correlation, and calculation of ramping magnitude, rate, and duration. This paper showcases WE-Validate with validation of synthetically derived power for a wind plant in Washington state for one month in 2018. Validation of the synthetic power from two comparison data sets compared with observations shows both comparison series have strong correlation with observed across weekly and monthly aggregations while suffering from persistent negative bias. The suite of metrics within WE-Validate facilitates immediate insight into the utility of the comparison data sets through compression across multiple axes. This user-friendly, open-source tool can be extended beyond wind power, making it a valuable resource for system planners and operators in different domains.

Moncheur de Rieudotte, Malcolm P.↗

Assessment of the polygeneration approach in wastewater treatment plants for enhanced energy efficiency and green hydrogen/ammonia production

Wastewater treatment plants (WWTPs) offer opportunities to optimize resource utilization and enhance energy efficiency. Here, this study provides a comprehensive analysis of using the polygeneration approach in WWTPs to reduce grid energy dependence, optimize energy distribution, and utilize surplus energy for hydrogen (H 2 ) and ammonia (NH 3 ) production. Several models were employed, including photovoltaic (PV) cells, parabolic trough collectors (PTCs), steam methane reforming, and polymer electrolyte membranes, to assess the feasibility of this approach. Three scenarios were evaluated and compared: Scenario 1 (Baseline) represents the current situation, Scenario 2 maximizes the Net Present Value (NPV), and Scenario 3 minimizes NH 3 production costs. Real data from As-Samra WWTP in Jordan was used to accurately assess the feasibility of each scenario. The results show that Scenario 2 offers the highest profitability and efficiency, with a NPV of 87.48 million USD and an annual NH 3 production of 15,417 tons, reducing both grid dependency and biogas fuel consumption. Both Scenarios 2 and 3 demonstrate the ability to meet thermal demands efficiently while generating significant revenue from NH 3 production. Scenario 3, in particular, achieves competitive H 2 and NH 3 production costs. Environmentally, Scenario 2 significantly reduces annual greenhouse gas emissions by 12.66 kilotons of CO 2eq , with near-zero carbon intensity for thermal energy due to solar reliance. In conclusion, the polygeneration approach offers a promising pathway for WWTPs to achieve greater sustainability, economic gains, and reduced environmental impact, providing valuable insights for decision-makers.

42 ENGINEERING↗

Object-Based Evaluation of Dynamical and Statistical Downscaled Precipitation Products over CONUS

High-resolution precipitation data, generated through dynamical downscaling (DD) or statistical downscaling (SD) of global climate model output, provide critical information for regional climate assessment and adaptation planning. Most downscaling development and validation have focused on accurate gridscale precipitation construction and ignored the spatial structure of precipitation across model grids and at the event scale. However, many applications, e.g., hydrologic modeling and the analysis using the downscaled precipitation, require a reasonable representation of the spatial structure of precipitation within watersheds. Therefore, a set of standard metrics to evaluate the representation of the spatial structure of individual storms across diverse downscaled precipitation products is desired. To address this need, we conducted an object-based evaluation of precipitation in decades-long DD and SD products over the contiguous United States (CONUS). Specifically, we evaluate their ability to reproduce various features of precipitation objects in the observations: total volume, precipitation area, peak intensity, and spatial structure. Multiple metrics (bias, Perkins score, and nonparametric statistical tests) are used to quantify model performance. Our evaluation reveals notable variations in performance among individual products across different climate zones and seasons, as well as between extreme and nonextreme events. In general, most DD products exhibit balanced performance across the four precipitation object features, while SD products vary more significantly in their performance across products. Based on this comprehensive evaluation, we provide guidance on choosing downscaled products for specific regions, seasons, and precipitation object features. These findings and recommendations can inform precipitation-relevant modeling and analysis over CONUS, guide future downscaling technique developments, and provide actionable information for climate impact assessment and adaptation.

Downscaling↗

Object-Based Evaluation of Dynamical and Statistical Downscaled Precipitation Products over CONUS

High-resolution precipitation data, generated through dynamical downscaling (DD) or statistical downscaling (SD) of global climate model output, provide critical information for regional climate assessment and adaptation planning. Most downscaling development and validation have focused on accurate gridscale precipitation construction and ignored the spatial structure of precipitation across model grids and at the event scale. However, many applications, e.g., hydrologic modeling and the analysis using the downscaled precipitation, require a reasonable representation of the spatial structure of precipitation within watersheds. Therefore, a set of standard metrics to evaluate the representation of the spatial structure of individual storms across diverse downscaled precipitation products is desired. To address this need, we conducted an object-based evaluation of precipitation in decades-long DD and SD products over the contiguous United States (CONUS). Specifically, we evaluate their ability to reproduce various features of precipitation objects in the observations: total volume, precipitation area, peak intensity, and spatial structure. Multiple metrics (bias, Perkins score, and nonparametric statistical tests) are used to quantify model performance. Our evaluation reveals notable variations in performance among individual products across different climate zones and seasons, as well as between extreme and nonextreme events. In general, most DD products exhibit balanced performance across the four precipitation object features, while SD products vary more significantly in their performance across products. Based on this comprehensive evaluation, we provide guidance on choosing downscaled products for specific regions, seasons, and precipitation object features. These findings and recommendations can inform precipitation-relevant modeling and analysis over CONUS, guide future downscaling technique developments, and provide actionable information for climate impact assessment and adaptation.

Environmental sciences↗

An ML-based terrestrial data fusion and augmentation framework to enable advanced understanding of the terrestrial carbon and water interactions

Soil moisture is essential to the terrestrial carbon and water cycles and land–atmosphere interactions. There are various types of soil moisture data, and each type has the distinct spatiotemporal strengths and limitations, depending on the diverse applications and retrieval methodologies of different data types (Li et al., in review; The PNNL-82151 FY23 Report). However, the limitations of different soil moisture data in terms of accuracy and spatiotemporal coverage hinder our ability to further understand the soil moisture dynamics across scales. To have a gap free soil moisture data product with a fine spatiotemporal coverage and vertical profiles, we train extreme gradient boosting (XGBoost) models by using (1) in-situ soil moisture measurements from the International Soil Moisture Network (ISMN), (2) soil moisture from the ECMWF reanalysis (ERA) at the 9 km and sub-daily spatiotemporal resolution, (3) the Daymet meteorological fields, and (4) data products that characterize surface conditions, including soil texture, organic content, topography, vegetation type, and rooting depth. We use the trained XGBoost models that have consistent performance across seven soil layers, i.e., 0–5 cm, 5–10 cm, 10–20 cm, 20–40 cm, 40–60 cm, 60–100 cm, and 100–200 cm, and the gridded model predictors to generate a soil moisture data at the 1 km and daily spatiotemporal resolution for the Continental United States (CONUS) from 2001–2020. This dataset can be broadly used for Earth system model benchmark, monitoring extreme weathers, making informed decisions regarding agriculture, water resource management, climate change mitigation, and ecosystem preservation.

58 GEOSCIENCES↗

Model Data Archive Associated with Manuscript "Fire-altered Carbon Pools Create Disturbance Memory in Stream Dissolved Organic Carbon"

This data package supports the publication “Fire-altered Carbon Pools Create Disturbance Memory in Stream Dissolved Organic Carbon” by Li et al. (2026). The package contains processed model inputs, configuration files, restart files, simulation outputs, scripts, and visualization products used to evaluate post-fire dissolved organic carbon (DOC) dynamics in the Naches River Watershed, Washington, USA, following the 2021 Schneider Springs Fire. The modeling workflow couples ELM-BGC, the biogeochemistry-enabled Energy Exascale Earth System Model Land Model; ATS, the Advanced Terrestrial Simulator for integrated surface-subsurface hydrology; and PFLOTRAN, a reactive transport model for multicomponent aqueous geochemistry. Together, these models simulate how wildfire-induced changes in vegetation, litter, coarse woody debris, and soil organic matter influence DOC production, transport, and reaction from burned hillslopes to stream networks. The archive includes preprocessed meteorological, geospatial, hydrologic, and biogeochemical forcing data; ELM-BGC-derived DOC source terms; ATS mesh files; PFLOTRAN reactive-transport inputs; model configuration files; spin-up and transient restart files; watershed-scale diagnostic outputs; stream concentration time series; and figures or visualization files used to inspect and reproduce key results. File types include Hierarchical Data Format 5 (HDF5) files for gridded forcing and model-coupling data, model input and configuration files for ELM-BGC, ATS, and PFLOTRAN, restart and simulation-output files generated by the modeling workflow, tabular or time-series diagnostic outputs, scripts for post-processing and figure generation, and image or visualization products associated with the manuscript. Use of the package depends on the intended task. Re-running the simulations requires the relevant modeling software, including ELM-BGC, ATS, and PFLOTRAN as ATS's geochemical engine. Inspecting outputs and reproducing figures requires Python with scientific plotting libraries such as Matplotlib, and three-dimensional model outputs may be viewed with ParaView. Geographic information system files or maps may be inspected with ArcGIS Pro or comparable GIS software. The data package is intended to enable traceability, reuse, and partial reproduction of the coupled land-to-watershed hydro-biogeochemical modeling workflow used to test how wildfire disturbance affects terrestrial carbon pools and downstream DOC dynamics.

ATS↗

Assessing Spatial Representativeness of Global Flux Tower Eddy-Covariance Measurements Using Data from FLUXNET2015

Large datasets of carbon dioxide, energy, and water fluxes were measured with the eddy-covariance (EC) technique, such as FLUXNET2015. These datasets are widely used to validate remote-sensing products and benchmark models. One of the major challenges in utilizing EC-flux data is determining the spatial extent to which measurements taken at individual EC towers reflect model-grid or remote sensing pixels. To minimize the potential biases caused by the footprint-to-target area mismatch, it is important to use flux datasets with awareness of the footprint. This study analyze the spatial representativeness of global EC measurements based on the open-source FLUXNET2015 data, using the published flux footprint model (SAFE-f). The calculated annual cumulative footprint climatology (ACFC) was overlaid on land cover and vegetation index maps to create a spatial representativeness dataset of global flux towers. The dataset includes the following components: (1) the ACFC contour (ACFCC) data and areas representing 50%, 60%, 70%, and 80% ACFCC of each site, (2) the proportion of each land cover type weighted by the 80% ACFC (ACFCW), (3) the semivariogram calculated using Normalized Difference Vegetation Index (NDVI) considering the 80% ACFCW, and (4) the sensor location bias (SLB) between the 80% ACFCW and designated areas (e.g. 80% ACFCC and window sizes) proxied by NDVI. Finally, we conducted a comprehensive evaluation of the representativeness of each site from three aspects: (1) the underlying surface cover, (2) the semivariogram, and (3) the SLB between 80% ACFCW and 80% ACFCC, and categorized them into 3 levels. The goal of creating this dataset is to provide data quality guidance for international researchers to effectively utilize the FLUXNET2015 dataset in the future.

54 ENVIRONMENTAL SCIENCES↗

CLM5 Simulations of Soil Moisture and Gross Primary Productivity for CONUS at 0.125 degrees

This dataset provides 0.125-degree gridded simulations of soil moisture and gross primary productivity (GPP) for the Contiguous United States (CONUS), generated using the Community Land Model version 5 (CLM5) with the biogeochemistry module enabled. The data covers a historical baseline (1980-2015) and mid-century future projections (2020-2055). Future projections are organized into two sets of scenarios to distinguish the impacts of different drivers: (1) Atmospheric Only (ATM): These scenarios apply future atmospheric forcings while holding land use and land cover (LULC) at historical baseline levels. The atmospheric forcings represent moderately versus severely hotter/drier atmospheric conditions (dynamically downscaled perturbed thermodynamics simulations based on CMIP6 SSP245 and SSP585 warming signals), each with cooler versus hotter Earth System Model temperature sensitivity instantiations. These scenarios are identified in the folder names as rcp45_cooler_near, rcp45_hotter_near, rcp85_cooler_near, and rcp85_hotter_near. (2) Coupled Atmospheric and Land-Use (LAND+ATM): These scenarios apply future atmospheric forcing together with future LULC by pairing atmospheric pathways with lower versus higher population/economic growth scenarios representing Shared Socioeconomic Pathways 3 and 5 (SSP3 and SSP5). These scenarios are identified in the file names as ssp3_rcp45_cooler_near, ssp3_rcp45_hotter_near, ssp5_rcp85_cooler_near, and ssp5_rcp85_hotter_near. Please refer to the "README_first.md" file for detailed information on file structure, variables, units, and data formats.

drought↗

Powering Data Centers with Clean Energy: A Techno-Economic Case Study of Nuclear and Renewable Energy Dependability

Rising data demands from artificial intelligence (AI) and large language models (LLMs) generating images, videos, and text have prompted increased need for larger and more robust data centers in the United States. Major companies interested in these larger data centers face the choice of linking them to existing regional grids, building stand-alone power supplies onsite, or a combination of both. The request, review, and approval process for new transmission lines to grids in the United States, however, has grown in recent years to times spans rivaling those of new construction for nuclear power plants. Building an islanded power supply for each data center is therefore becoming a prominent option. In this case study, several technologies are modeled in techno-economic simulations for long-term system costs subject to fixed electricity demand from a singular data center. A 250 MWe data center is assumed with additional 50 MWe for resiliency. Techno-economic simulations are conducted using the Holistic Energy Resource Optimization Network (HERON) software, which is a part of the Framework for Optimization of Resources and Economics (FORCE) tool suite. Technologies considered include solar, wind, lithium-ion batteries, and several types of nuclear reactors: large-scale reactors, small modular reactors, and microreactors. A low- and high-cost estimate for each technology is assumed to develop a range of expected economic performance. Low-cost estimates included several clean energy production tax credits. Different combinations of renewable energy generators with nuclear reactors are considered, ranging from a fully renewable-powered data center to a fully nuclear-powered data center. Historic time series of wind and solar availability from the Texas grid are used to train a reduced order model; this model then generates unique time series with similar characteristics of the training dataset. Multiple scenarios of weather and subsequent operations are simulated for each renewable-nuclear combination to determine total costs throughout the project lifetime. Fully renewable-powered configurations required large amounts of installed capacity (GW scale) in the simulations to meet the fixed demand of the data center. This is due to some scenarios in the historical dataset which captured low-wind and low-solar days, requiring over-building of these technologies as well as batteries to compensate for the low amounts of electricity generation. Fully nuclear-powered configurations outperformed the fully renewable and mixed renewable-nuclear configurations in terms of cost, with ranges between $1B and $10B in 2023 USDs compared to $40B+ for fully renewable configurations. Of the nuclear technologies, small modular reactors performed better economically than large-scale nuclear models due to lower projected capital costs, and both performed better than the microreactor models. These results demonstrate the applicability of firm, dispatchable electricity resources from baseload generators like nuclear power plants for operating facilities that run at constant power without daily variability.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Electrification Options for Multi-Family Water Heating in Cold Climates - Final Report

In multi-family buildings, large water storage tanks in centralized domestic hot water (DHW) systems can serve as thermal energy storage (TES) batteries to mitigate grid impact. These systems offer demand shift and efficiency benefits, significantly reducing peak power consumption, particularly in cold climates. This study evaluates the load-shifting benefits of a centralized heat pump water heater (HPWH) system equipped with a CO2 heat pump in multi-family buildings through simulation. The heat pump system and water storage tank are sized using design-day sizing. A finite-element-based stratified tank model and CO2 heat pump performance map from a commercial DHW product are used. Annual simulations are conducted to assess the benefits of the centralized DHW system for energy efficiency improvements, load shifting, and emission reductions. These simulations incorporate utility tariffs and marginal grid emission data from Los Angeles and Chicago. In Los Angeles, using a water tank as a thermal battery achieves 7.4% utility cost savings and 10.2% emission reduction. In Chicago, compared to HPWH conventional operation without preheating, TES-enabled central HPWH provides 15% utility cost savings and 13% emission reduction. The case study demonstrates that the demand reduction potential of central CO2 HPWHs is significant in cold climate regions.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Assessment and Coordination of EVSE Cybersecurity Standards

Cybersecurity certification programs for Electric Vehicle Supply Equipment (EVSE) are fragmented due to no single certification covering all aspects of the device and additionally the existence of multiple programs and under different levels of regulation. These devices are also confronted by the intricate assembly of product software, firmware, and hardware. Devices contain both logical and physical interfaces. These multifaceted devices have vulnerabilities at many levels and interconnect with other potentially vulnerable systems including the electric vehicle, the cloud where data and payment information are stored, and the electric grid and electric grid equipment including utilities. Of the EVSE certification programs that are found, none are directly for the cybersecurity of EVSE. Many standards are for safety, specifically battery safety, some are cybersecurity standards for other types of equipment and can be modeled for EVSE. In specific, ISA/IEC 62443 is found to be significantly in line with EVSE security needs and will be used in future testing to certify EVSE and help guide the project to demonstrate where gaps exist, where strengths lie in the standard and how this can be used to lead the certification efforts in harmonizing EVSE cybersecurity standards. In addition, there are multiple efforts that are currently seeking to build EVSE standards or revise existing standards to address gaps. This effort is seeking to establish a cybersecurity program for EVSE that will inform customers and help increase the level of security across products and state EVSE procurements to achieve consistency across different jurisdictions.

33 ADVANCED PROPULSION SYSTEMS↗

Second-generation downscaled earth system model data using generative machine learning

The second-generation Sup3rCC dataset provides high-resolution meteorological data generated through the downscaling of multiple earth system models (ESMs) from the Coupled Model Intercomparison Project Phase 6 (CMIP6). This downscaling is performed through application of a generative machine learning approach called Super-Resolution for Renewable Resource Data (sup3r). This dataset builds on the first-generation Sup3rCC data by applying improved bias correction methods and adding downscaled precipitation to the output variables. As with the first Sup3rCC version, the data still include temperature, wind speed and direction at multiple heights, pressure, three components of downwelling solar radiation, and relative humidity—all at 4-kilometer (km) hourly resolution over the contiguous United States. This is a 25x spatial enhancement and 24x temporal enhancement of the source 100-km daily-average ESM data. This extension of the Sup3rCC dataset includes data from six ESMs from two shared socioeconomic pathways (SSPs) totaling 400 years of data with multiple future projections of changing meteorological conditions. The scenario selection was based on a structured evaluation of historical ESM skill and comprehensive representation of possible trajectories of future climate change in temperature, humidity, precipitation, solar irradiance, and near-surface wind speeds. The inclusion of multiple future projections is intended to enable users to assess key drivers of un 36 certainty and variability. All data are double-bias corrected, resulting in a product that can be used out-of-the-box for energy system analysis with minimal historical bias. The potential applications of Sup3rCC data extend to various topics in renewable energy resource assessment, energy systems modeling, and grid resilience studies. High-resolution future meteorological projections are critical for evaluating the effects of changing meteorological conditions on renewable energy generation, energy demand, and for optimizing energy storage and grid infrastructure. The 4-km hourly resolution of the downscaled data enables understanding of spatial and temporal variability at the scales necessary for energy system operational planning. In addition, the dataset can support risk assessments by providing detailed information on possible future extreme weather events and long-term meteorological variability at scales relevant to energy infrastructure. By offering an enhanced representation of possible future meteorological conditions, the second-generation Sup3rCC dataset enables more precise modeling of energy resilience and adaptation strategies in response to changing meteorological conditions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Remote Sensing of Live Fuel Moisture for Wildfires Using SMAP Satellite Observations

Live Fuel Moisture (LFM) is a critical parameter for wildfire risk assessment, traditionally measured by labor-intensive field sampling. However, sampled LFM data are influenced by site-specific factors, such as local vegetation types and plant traits, and are often collected retrospectively after wildfire events, making it difficult to obtain pre-fire data for predictive applications. Here, we evaluate the relationship between LFM and Vegetation Water Content (VWC) and Soil Moisture (SM) retrieved from SMAP L-band brightness temperature using the Maximum Entropy Production (MEP) approach. The MEP-retrieved VWC exhibited strong correlation with in situ measurements of LFM ( r > 0.6) in the Western U.S. The integration of high-resolution vegetation coverage data enhances the detection of sub-grid vegetation heterogeneity. This study demonstrates the operational potential of remote sensing derived VWC as a scalable proxy of LFM, supporting its application in regional assessment of wildfire risk.

Cho, Kyeungwoo [Georgia Institute of Technology, A↗

I/O performance studies of analysis workloads on production and dedicated resources at CERN

The recent evolutions of the analysis frameworks and physics data formats of the LHC experiments provide the opportunity of using central analysis facilities with a strong focus on interactivity and short turnaround times, to complement the more common distributed analysis on the Grid. In order to plan for such facilities, it is essential to know in detail the performance of the combination of a given analysis framework, of a specific analysis and of the installed computing and storage resources. This contribution describes performance studies performed at CERN, using the EOS disk-based storage, either directly or through an XCache instance, from both batch resources and highperformance compute nodes which could be used to build an analysis facility. A variety of benchmarks, both synthetic and based on real-world physics analyses and their corresponding input datasets, are utilized. In particular, the RNTuple format from the ROOT project is put to the test and compared to the latest version of the TTree format, and the impact of caches is assessed. In addition, we assessed the difference in performance between the use of storage system specific protocols, like XRootd, and FUSE. The results of this study are intended to be a valuable input in the design of analysis facilities, at CERN and elsewhere.

Sciabà, Andrea↗

Non-linear relationships between daily temperature extremes and US agricultural yields uncovered by global gridded meteorological datasets

Global agricultural commodity markets are highly integrated among major producers. Prices are driven by aggregate supply rather than what happens in individual countries in isolation. Furthermore, estimating the effects of weather-induced shocks on production, trade patterns and prices hence requires a globally representative weather data set. Recently, two data sets that provide daily or hourly records, GMFD and ERA5-Land, became available. Starting with the US, a data rich region, we formally test whether these global data sets are as good as more fine-scaled country-specific data in explaining yields and whether they estimate similar response functions. While GMFD and ERA5-Land have lower predictive skill for US corn and soybeans yields than the fine-scaled PRISM data, they still correctly uncover the underlying non-linear temperature relationship. All specifications using daily temperature extremes under any of the weather data sets outperform models that use a quadratic in average temperature. Correctly capturing the effect of daily extremes has a larger effect than the choice of weather data. In a second step, focusing on Sub Saharan Africa, a data sparse region, we confirm that GMFD and ERA5-Land have superior predictive power to CRU, a global weather data set previously employed for modeling climate effects in the region.

54 ENVIRONMENTAL SCIENCES↗