Search NASASearch

SEARCH · Search NASA

Results for “proxy data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

North‐East Peri‐Tethyan Water Column Deoxygenation and Euxinia at the Paleocene Eocene Thermal Maximum

Abstract The Paleocene–Eocene Thermal Maximum (PETM) is associated with climatic change and biological turnover. It shares features with the Oceanic Anoxic Events (OAEs) of the Mesozoic, such as transient global warming and biogeochemical perturbations. However, the PETM experienced a more muted expansion of marine anoxia compared to the Mesozoic OAEs (especially OAE2), with geographically limited evidence for photic zone euxinia (PZE). We explore the extent and drivers of marine deoxygenation during the PETM using biomarkers for water column euxinia and anoxia as well as an intermediate complexity Earth system model (cGEnIE). These reveal that the water column in the North‐East Peri‐Tethys became anoxic, with euxinic conditions reaching the photic zone (PZE) during the PETM. Our model shows that euxinia developed due to a global increase in the ocean nutrient inventory with concomitant oxygen consumption, similar to findings for OAE2. The particularly strong regional response in the NE Peri‐Tethys appears to arise from a combination of global CO 2 ‐weathering forcing, regionally restricted circulation and upwelling of sulphidic thermocline waters. Unlike OAE2, anoxia and PZE do not become widespread in our PETM simulations, consistent with new and existing geochemical and biological proxy data. This globally muted response could result from reduced biogeochemical feedbacks to climate forcing relative to the mid‐Cretaceous climate. Our observations suggest that similar mechanisms operated in response to disparate Cenozoic (PETM) and Mesozoic (OAEs) transient global warming events, while also highlighting that background conditions are crucial in modulating the sensitivity of Earth's system to them.

Behrooz, L.

Machine Learning Classification Strategy to Improve Streamflow Estimates in Diverse River Basins in the Colorado River Basin

Streamflow in the Colorado River Basin (CRB) is significantly altered by human activities including land use/cover alterations, reservoir operation, irrigation, and water exports. Climate is also highly varied across the CRB which contains snowpack-dominated watersheds and arid, precipitation-dominated basins. Recently, machine learning methods have improved the generalizability and accuracy of streamflow models. Previous successes with LSTM modeling have primarily focused on unimpacted basins, and few studies have included human impacted systems in either regional or single-basin modeling. We demonstrate that the diverse hydrological behavior of river basins in the CRB are too difficult to model with a single, regional model. We propose a method to delineate catchments into categories based on the level of predictability, hydrological characteristics, and the level of human influence. Lastly, we model streamflow in each category with climate and anthropogenic proxy data sets and use feature importance methods to assess whether model performance improves with additional relevant data. Overall, land use cover data at a low temporal resolution was not sufficient to capture the irregular patterns of reservoir releases, demonstrating the importance of having high-resolution reservoir release data sets at a global scale. On the other hand, the classification approach reduced the complexity of the data and has the potential to improve streamflow forecasts in human-altered regions.

54 ENVIRONMENTAL SCIENCES

A Proxy Method to Bridge LCA Data Gaps Using Automated Material Classification and Probabilistic Under-Specification

Life cycle assessments (LCAs) are essential for understanding the environmental impacts of material production. However, gaps in life cycle inventory (LCI) data for material and chemical inputs present a key challenge for LCA practitioners, especially in the early design stages. Strategies for filling in these gaps require additional time and expertise, which can hinder the LCA’s completion. This study combined automatic material classification and probabilistic under-specification to create a time-efficient method to fill material LCI data gaps. To illustrate the proposed method, proxy environmental impact distributions were generated using publicly available material LCI data classified into the ChemOnt chemical taxonomy using the open-source chemical classification software ClassyFire. Input materials with data gaps were then classified into the same taxonomy, where proxy environmental impact values could be selected from the available distributions to quickly fill in any data gaps. Although these methods were applied to classify material production processes available in the Federal LCA Commons and Ecoinvent databases, they can be applied to any LCA database. This study shows that classifying materials by their chemical structure produces taxonomies with increased granularity relative to industrial classification, improving the ability of under-specified proxy data to be used for differentiating the environmental impacts of competing designs.

biological databases

Towards self-consistent integrated modeling of the tokamak pedestal, scrape-off layer, and divertor using SOLPS-ITER and EPED

A predictive, self-consistent integrated model of the tokamak edge plasma and neutral system has been developed and validated against experimental data from DIII-D, spanning from the divertor target and scrape-off layer to the top of the pedestal. The model uses the SOLPS-ITER fluid boundary plasma and neutral code to constrain the EPED prediction of the pedestal height and width by taking the separatrix temperature and density calculated from SOLPS as inputs to EPED. An empirical proxy function based on data from a subset of these discharges is then used to infer the ratio of density at the pedestal top to that at the separatrix based on the electron temperature at the divertor target. This allows the density at the top of the pedestal to be provided to EPED as well, so that knowledge of this quantity is no longer required a priori, while SOLPS radial transport coefficients are needed instead. The gas puff rate provided to SOLPS is then a primary model input for this forward model coupling (along with radial transport coefficients, which are kept fixed through each scan in this study). Comparisons with measurements from DIII-D density scans through a range of target conditions find significant pedestal pressure degradation as the density is increased towards detachment in three different divertor configurations, consistent with the experimental data in these ballooning-limited pedestal regimes.

core-edge integration

Object Proxy Patterns for Accelerating Distributed Applications

Workflow and serverless frameworks have empowered new approaches to distributed application design by abstracting compute resources. However, their typically limited or one-size-fits-all support for advanced data flow patterns leaves optimization to the application programmer—optimization that becomes more difficult as data become larger. The transparent object proxy, which provides wide-area references that can resolve to data regardless of location, has been demonstrated as an effective low-level building block in such situations. Here we propose three high-level proxy-based programming patterns—distributed futures, streaming, and ownership—that make the power of the proxy pattern usable for more complex and dynamic distributed program structures. We motivate these patterns via careful review of application requirements and describe implementations of each pattern. As a result, we evaluate our implementations through a suite of benchmarks and by applying them in three meaningful scientific applications, in which we demonstrate substantial improvements in runtime, throughput, and memory usage.

Distributed Computing

From Edge to HPC: Investigating Cross-Facility Data Streaming Architectures

In this paper, we investigate three cross-facility data streaming architectures, Direct Streaming (DTS), Proxied Streaming (PRS), and Managed Service Streaming (MSS). We examine their architectural variations in data flow paths and deployment feasibility, and detail their implementation using the Data Streaming to HPC (DS2HPC) architectural framework and the SciStream memory-to-memory streaming toolkit on the production-grade Advanced Computing Ecosystem (ACE) infrastructure at Oak Ridge Leadership Computing Facility (OLCF). We present a workflow-specific evaluation of these architectures using three synthetic workloads derived from the streaming characteristics of scientific workflows. Through simulated experiments, we measure streaming throughput, round-trip time, and overhead under work sharing, work sharing with feedback, and broadcast and gather messaging patterns commonly found in AI-HPC communication motifs. Our study shows that DTS offers a minimal-hop path, resulting in higher throughput and lower latency, whereas MSS provides greater deployment feasibility and scalability across multiple users but incurs significant overhead. PRS lies in between, offering a scalable architecture whose performance matches DTS in most cases.

George, Anjus [ORNL] (ORCID:0000000179737061)

Empirical Indicators of Transmission Value in the Southeast United States

Concurrent differences in energy price between different parts of the electric grid are a key indicator of the value of additional transmission. In areas without a wholesale electricity market, such as the Southeast, an alternative indicator to price is the Federal Energy Regulatory Commission’s (FERC) system lambda data. This economic metric represents the minimized marginal production costs of thermal generators, including fuel and other variable operation and maintenance expenses. Balancing Authorities report a single system lambda for their entire balancing area. Most Southeastern lambdas exhibit sufficient price variation to support a transmission valuation analysis, although incomplete accounting of congestion costs or scarcity rents during peak load hours may underestimate the true value of transmission capacity. With transmission value defined as the annual average hourly absolute price difference between two regions and FERC’s system lambda data used as a price proxy, we find the following results in the Southeast region during 2012-2023 (reported in $\$2024$/MWh): Intra‐regional findings: Annual averages historically span $\$2$–$\$28$/MWh and average $\$12$/MWh in SERTP and span $\$4$–$\$19$/MWh and average $\$9$/MWh in FRCC, disregarding transmission value driven by anomalous data. The ranges of transmission value reported here are large, spanning an order of magnitude in some cases. Much of this variation is driven by year-to-year changes, with 2022 having a particularly high intra-regional transmission value due to elevated natural gas prices. Inter‐regional corridors: Annual average transmission values across three broader regions range from $\$6$ to $\$28$/MWh with a long-term average of $\$11$/MWh. Much of the transmission value is concentrated in a small portion of hours. Across all regions, severe weather—particularly polar vortex events in January 2018, February 2021, and December 2022—drives the largest price spreads. Seasonal patterns also emerge, with summer afternoons and fall mornings contributing consistently to transmission value, as for example between MISO and SOCO in 2023.

24 POWER TRANSMISSION AND DISTRIBUTION

Continuous snow depth and temperature measurements from dense network of above-ground distributed temperature profiling systems from 2021-09-23 to 2024-08-23, Seward Peninsula, Alaska

The dataset contains temperature measurements from distributed temperature profiling (DTP) systems (Dafflon et al., 2022; Wielandt et al., 2022; Wang et al., 2024a; Fiolleau et al., 2024) deployed vertically above the ground surface at a large number of locations from 2021 to 2024. The research is designed to improve understanding of the local heterogeneity in snow depth and snow thermal insulation dynamics, as well as their interactions in a discontinuous permafrost region (Wang et al., 2025). The DTP systems were deployed at 96 locations in a watershed along the Nome-Teller road at mile marker 27 (T27) and at 54 locations on a hillslope along the Kougarok road at mile marker 64 (K64) in the Seward Peninsula, Alaska. The probe location information is stored in Probe_locations_T27.csv and Probe_locations_K64.csv. Temperature measurements were recorded at 15-minute intervals using high-precision digital sensors (accuracy: ±0.1°C, resolution: 0.0078°C). The temperature probes, either 1.4 m or 1.6 m long, contain sensors spaced every 5 cm or 10 cm along their length. The temperature data are stored in compressed files following the format: DTP_snow_air_temperature_(site)_(start)_(end).zip, where site is either T27 or K64, and start and end represent the time series period. Within each ZIP file, individual CSV files are named by probe ID and contain temperature records at different heights above the ground surface.This dataset also includes derived snow depth time series over three snow seasons, estimated from temperature measurements. Snow depth was estimated by identifying the consecutive sensor pair that exhibited the largest drop in high-frequency temperature fluctuations (detailed in the methods). These data are stored in: Snow_depths_flags_(site)_(start)_(end).csv, which includes snow depth time series and corresponding quality flags (defined in the methods) from different probes. Additionally, the dataset includes derived metrics and supporting measurements at selected locations over two snow seasons, contributing to the manuscript of Wang et al., 2025. These locations were chosen based on the availability of high-quality snow depth time series during both seasons. The additional data include: (1) Air temperature proxies measured from the top sensors on the pole when they were not buried by snow, stored in Air_temperature_proxies_(site)_(start)_(end).csv (2) Ground interface temperature, recorded at 3 cm above the ground, stored in Ground_interface_temperature_(site)_(start)_(end).csv (3) Site characteristics, including vegetation height, elevation, and the topographic position index (TPI) within a 50 m radius, stored in Selected_probe_locations_gps_vegheight_tpi_elevation_(site).csv. These metrics were derived from 1 m resolution summer LiDAR-based digital elevation models and digital surface models from Singhania et al., 2023, DOI:10.5440/1832016. Metadata files include data descriptions (_dd.csv) for tabular data. All included files are listed and described in xxxx_flmd.csv.This dataset is an updated version of a previous archive (Wang et al., 2024b, DOI: 10.15485/2475020), incorporating multiple seasons and improved snow depth estimation. Please note that due to large amount of information present in this dataset, many specificities associated with the acquisition of snow temperature, air temperature proxy and estimation of snow depth, and the future archiving of additional datasets on the soil temperature, thaw depth and soil characteristics at these locations, the author would welcome being contacted by people planning to use this dataset.The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES

Estimating Return on Investment for Energy Technical Assistance Programs

The U.S. Department of Energy's Office of State and Community Energy Programs engaged the National Laboratory of the Rockies to assess the return on investment (ROI) of technical assistance (TA) programs that support state, local, and Tribal energy planning. Although TA delivers value through capacity building, stakeholder engagement, and knowledge transfer, these benefits are often intangible and challenging to monetize. This study reviews existing ROI frameworks and synthesizes the most relevant elements into a hybrid approach tailored to energy TA programs. The proposed framework integrates monetary and non-monetary outcomes through early logic model development, baseline data collection, and the use of proxies for intangible benefits. As a case study, this paper applies this approach to the Communities Local Energy Action Program (Communities LEAP), demonstrating how ROI can inform program design, data strategy, and performance assessment. Findings underscore that ROI should be applied selectively and planned from the outset to ensure data alignment and attribution accuracy. The framework offers TA practitioners a structured approach that can be leveraged for future programs to evaluate and communicate the multifaceted value of TA investments.

29 ENERGY PLANNING, POLICY, AND ECONOMY

Annual and sub-seasonal dynamics of a rapidly eroding permafrost coastline along the Beaufort Sea in northern Alaska

Drew Point, an unlithified ice-rich permafrost coastline along the Alaskan Beaufort Sea, is among the most rapidly eroding Arctic coastlines, with an average erosion rate of 19 m/yr from 2007 to 2019. We use 16 high-resolution remote sensing datasets (satellite, airborne, and UAV imagery) to analyze erosion mechanisms (thermal abrasion and denudation) in relation to environmental forcings along a 1.5 km stretch of coastline during the 2018 and 2019 open water seasons. In a striking contrast, 2019 exhibited the highest mean erosion rate (34.5 m) within the 2007–2019 record, while 2018 had the second lowest (11.2 m). Block failure contributed to sub-seasonal erosion rates 6 to 21 times higher than thermal denudation, with staggered block fall timing, lag responses post-storm, and non-storm block collapse influencing overall erosion magnitude and timing. To quantify wind effects, we developed wind sums, a metric combining cumulative wind speed and directional data that can be used as a proxy for integrated storm intensity capable of incorporating lagged responses that correlated strongly with erosion at sub-seasonal and annual scales. Our findings emphasize the dominant role of wind during periods of open water and air temperature during the thaw season in driving permafrost coastline erosion dynamics, while highlighting the importance of spatiotemporally high-resolution datasets for understanding Arctic coastal change dynamics.

Alaska Beaufort Sea Coast

Unraveling the Threads of Environmental Justice in Critical Mineral Extraction: A Framework for Regionalized Life Cycle Data

As we transition to a more sustainable energy system, the extraction and processing of critical minerals becomes increasingly important. However, these industries often raise concerns about environmental and social impacts, particularly in disadvantaged communities. To address these concerns, our research focuses on regionalizing environmental life cycle data to connect it with communities affected by mineral extraction. A framework was developed for collecting life cycle background data that supports the Justice40 Toolset, a market-based approach to evaluating net benefits and costs of critical mineral material recovery pathways. A goal was to alleviate public skepticism around mineral extraction processes, including secondary and unconventional feedstocks, by highlighting both environmental and social impacts. To achieve this, computational analysis and geospatial data science techniques were employed, such as within-scale and across-scale methods and proxy dataset usage. By doing so, we were able to develop a framework for identifying and disaggregating data down to regions small enough to support Justice40 goals. Our approach not only provides valuable tools fo insight into the environmental implications of mineral extraction but also helps policymakers evaluate the social impacts on local communities. This research contributes to a more just and equitable transition to a sustainable energy system, ensuring that marginalized voices are heard in decision-making process.

Davis, Tyler [NETL Site Support Contractor, Nation

Package Data for CERF-Data Centers

This dataset contains sample input 100m resolution raster files for running the CERF-DC python package (see https://github.com/IMMM-SFA/cerf_data_centers) at the state level across the CONUS. Due to data availability constraints, some of the items included in this dataset are proxies or assumptions for siting factors used in the model. These are individually noted in the item descriptions and can be exchanged with more detailed information upon availability. Data Descriptions The following raster files are included in the data download: state_siting_region.tif — State areas identified by state FIPS code composite_siting_suitability.tif — Value of 1 indicates suitable siting location, 0 otherwise. The following areas are excluded from siting: Areas within 300m of a federal airport runway Waterbodies Areas with slope >16% Areas susceptible to sinkholes High coastal or inland flood risk areas Local, state, and federal parks, leisure areas, and cemeteries Areas >2 km away from electric substations Areas >5 km away from a municipal water supplier service area Areas >2 km away from high-speed fiber provider service territory Protected Areas Database of the United States (PAD-US) areas Railroads, major roadways, and minor roadways Military areas and training grounds Developed lands Areas >0.8 km (0.5 miles) from developed lands land_value_dollar_per_sqft.tif — USD per square foot (sqft) derived from USDA $/acre land cost personal_property_tax_rate.tif — Personal property tax rate by state. Uses an assumed 0.0125 personal property tax rate for states with personal property tax, 0 for states without personal property tax. real_property_tax_rate.tif — Real property tax rate. Based on county level residential real estate property tax rates. sales_tax_rate.tif — Sales tax rate by state. mechanical_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through mechanical processes based on local water stress and humidity levels. water_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through evaporative (water cooled) processes based on local water stress and humidity levels. distance_to_substation.tif — Distance to nearest substation in hundreds of meters (i.e., value of 1 equals a distance of 100m). Offshore areas have a value of 0. industrial_electricity_rates_dollar_per_kwh.tif — USD/kWh industrial electricity rates. Represents the average industrial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. commercial_electricity_rates_dollar_per_kwh.tif — USD/kWh commercial electricity rates. Represents the average commercial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. data_center_market_locations.tif — Grid cells with positive values represent the centroid of existing data center market clusters. The value of non-zero grid cells represents the number of data centers in the market cluster. All other grid cells have a value of 0. Geospatial Metadata CRS: Albers Equal Area Conic (ESRI:102003) Extent: -2415585.0000000023283064,-1441981.2605773280374706 : 2384414.9999999976716936,1708018.7394226719625294 Dimensions: X: 48000 Y: 31500 Bands: 1 Origin: -2415585.0000000023283064,1708018.7394226719625294 Pixel Size: 100,-100 Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4.0 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall

Package Data for CERF-Data Centers

This dataset contains sample input 100m resolution raster files for running the CERF-DC python package (see https://github.com/IMMM-SFA/cerf_data_centers) at the state level across the CONUS. Due to data availability constraints, some of the items included in this dataset are proxies or assumptions for siting factors used in the model. These are individually noted in the item descriptions and can be exchanged with more detailed information upon availability. Data Descriptions The following raster files are included in the data download: state_siting_region.tif — State areas identified by state FIPS code composite_siting_suitability.tif — Value of 1 indicates suitable siting location, 0 otherwise. The following areas are excluded from siting: Areas within 300 m of a federal airport runway or within an airport area boundary Waterbodies Areas with slope >16% Areas susceptible to sinkholes High coastal or inland flood risk areas Local, state, and federal parks, leisure areas, and cemeteries Areas >2 km away from electric substations Areas >5 km away from a municipal water supplier service area Areas >2 km away from high-speed fiber provider service territory USGS Protected Areas Database of the United States (PAD-US) GAP status 1, 2, or 3 areas US National Parks Wetlands USFWS critical habitats BIA land areas Railroads, major roadways, and minor roadways Military areas and training grounds NLCD developed lands Areas >0.8 km (0.5 miles) from NLCD developed lands land_value_dollar_per_sqft.tif — USD per square foot (sqft) derived from USDA $/acre land cost personal_property_tax_rate.tif — Personal property tax rate by state. Uses an assumed 0.0125 personal property tax rate for states with personal property tax, 0 for states without personal property tax. real_property_tax_rate.tif — Real property tax rate. Based on county level residential real estate property tax rates. sales_tax_rate.tif — Sales tax rate by state. mechanical_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through mechanical processes based on local water stress and humidity levels. water_cooling_fraction.tif — Fraction of year (values between 0 and 1, inclusive) that the data center would be cooled through evaporative (water cooled) processes based on local water stress and humidity levels. distance_to_substation.tif — Distance to nearest substation in hundreds of meters (i.e., value of 1 equals a distance of 100m). Offshore areas have a value of 0. industrial_electricity_rates_dollar_per_kwh.tif — USD/kWh industrial electricity rates. Represents the average industrial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. commercial_electricity_rates_dollar_per_kwh.tif — USD/kWh commercial electricity rates. Represents the average commercial rate across all utilities that operate within a given county. Values are derived from the US Utility Rate Database. data_center_market_locations.tif — Grid cells with positive values represent the centroid of existing data center market clusters. The value of non-zero grid cells represents the number of data centers in the market cluster. All other grid cells have a value of 0. Geospatial Metadata CRS: Albers Equal Area Conic (ESRI:102003) Extent: -2415585.0000000023283064,-1441981.2605773280374706 : 2384414.9999999976716936,1708018.7394226719625294 Dimensions: X: 48000 Y: 31500 Bands: 1 Origin: -2415585.0000000023283064,1708018.7394226719625294 Pixel Size: 100,-100 Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4.0 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall

Measuring Thread Timing to Assess the Feasibility of Early-Bird Message Delivery Across Systems and Scales

Early-bird communication is a communication/computation overlap technique that leverages fine-grained communication to improve application run-time. Communication is divided such that each individual thread can initiate transmission of its portion of the data upon completion rather than waiting for a dedicated communication phase. The benefit of early-bird communication depends on the completion timing of the individual threads: On the one hand, if all threads are complete at nearly the same time, the overheads of sending multiple messages will accumulate, leading to performance that is worse than if a single message had been sent. On the other hand, if thread completions are spread out in time, those that complete earlier can send data while others continue working, leading to performance that is better than if a single message had been sent. The challenge is that the completion times are currently unknown and can vary based on application, problem size, system software, and underlying hardware. In this paper, we address this lacuna by measuring and evaluating the potential overlap afforded by early-bird communication for a selection of proxy applications. These measurements help us understand whether a given application could benefit from early-bird communication. Here, we present our technique for gathering this data and evaluate data collected from three proxy applications: MiniFE, MiniMD, and MiniQMC. Each application is run on three systems with distinct CPU architectures and strong scales across three run sizes. To characterize the behavior of these workloads, we study the trends of thread timings at both a macro level, across all threads across all runs of an application, and a micro level, that is, within a single process of a single run. We observe that our tested applications exhibit significantly different thread arrival distributions. The machine used had a significant impact, with the window of potential overlap varying by as much as an order of magnitude.

97 MATHEMATICS AND COMPUTING

Effects of 9.5 Years of Whole-Soil Warming on the Fatty Acid and n-Alkanes Composition in Bulk Soil and Density Fractions at Blodgett Experimental Forest, California, USA

Original data of molecular data (fatty acids and n-alkanes) including concentrations and calculated molecular proxies in a whole-soil warming experiment at the Blodgett Forest Research Station after 9.5 years of warming. The study site has a Mediterranean climate with annual average temperature of 12.5 ℃ and annual average precipitation of 1774 mm. The study site is characterized by a mesic Ultic Alfisol formed from granitic parent material, corresponding to a Dystric Cambisol under the World Reference Base for Soil Resources (WRB) classification system. Experimental warming is applied throughout the soil profile to a depth of 1 m using vertically embedded heating cables that raise soil temperature by 4 °C relative to ambient conditions. Soil samples were collected on 1 May 2023, after the experiment had been operating continuously for about 9.5 years since its initiation in January 2014.The data has been processed from raw data and cross-validated by other peers. The dataset includes: - Bulk_Fattyacid_9.5-year_Soil_Warming_Blodgett, California, USA: fatty acid concentrations and proxies including Carbon Preference Index (CPI) and Average Chain Length (ACL) of bulk soil organic carbon; - Fractions_Fattyacid_9.5-year_Soil_Warming_Blodgett, California, USA: fatty acid concentrations and proxies including CPI and ACL of free particulate organic matter (fPOM) and mineral-associated organic matter (MAOM); - Bulk_Alkanes_9.5-year_Soil_Warming_Blodgett, California, USA: n-alkanes concentrations and proxies including CPI and ACL of bulk soil organic carbon; - Fractions_Alkanes_9.5-year_Soil_Warming_Blodgett, California, USA: n-alkanes concentrations and proxies including CPI and ACL of fPOM and MAOM; - n-Alkanes_All_Monomer_Concentration_9.5-year_Soil_Warming_Blodgett, California, USA: concentration of all the n-alkane monomers identified and integrated for bulk soil, fPOM and MAOM; - Fattyacid_All_Monomer_Concentration_9.5-year_Soil_Warming_Blodgett, California, USA: concentration of all the fatty acid monomers including diacids identified and integrated for bulk soil, fPOM, and MAOM. All data are provided in CSV format and can be viewed using Microsoft Excel. We specifically look at fatty acids (FA) and n-alkanes in bulk soil, fPOM and MAOM and calculated molecular proxies such as CPI and ACL to understand the source of oragnic carbon (with ACL) and degree of decomposition (CPI) of each soil fraction. Due to lack of long-chain fatty acids (carbon number ⩾ 20), microorganism-derived organic carbon is characterized by shorter ACL in comparison to plant-derived organic carbon. Fresh SOC is characterized by even-over-odd dominance for fatty acids and odd-over-even dominance for n-alkanes. Therefore, CPI indicates whether soil organic carbon (SOC) represents fresh input (CPI > 10) or is strongly decomposed (close to 1). The research questions should be then, after 9.5-year warming: 1. whether the relative contribution between microorganism-derived and plant-derived SOC in each soil fraction? 2. whether fPOM became more decomposed whereas MAOM remained relatively persistent in each soil fraction across the soil depth?

Carbon

"Hidden" hydrothermal technical potential & technoeconomics: Revealing permeability & fluids with more data

Historical hydrothermal estimates have largely relied on temperature or heat flow estimates ignoring the need for natural flowing fluids. More accurate hydrothermal estimates require some indication of permeability and fluids that naturally exist in the subsurface. This paper describes a novel approach that includes proxies of permeability and fluids in hydrothermal estimates by leveraging the relatively data-rich Great Basin. Specifically, nameplate capacities (megawatts) of operating geothermal plants, negative (0 megawatt) locations and 48 geophysical and geologic features are used to used in eXtreme Gradient Boosting (XGBoost) regression to make hydrothermal capacity predictions. Additionally, this work inputs the XGBoost-based hydrothermal predictions into the Renewable Energy Potential (reV) model to quantify technical capacity, its uncertainty and techno-economics. Compared to historical hydrothermal estimates, these predictions adhere to the 37 operating geothermal plants and negative locations. We present a method for subsampling the negative sites to bring the labels into balance that uses the geologic domain knowledge to proportionally represent negatives. Overall, the distributions of the hydrothermal technical capacity and the site levelized cost of energy are respectively much tighter, lower and more accurate than the previous estimates for the Great Basin, as they include geological and geophysical surrogates for permeability and fluids. Percentile (50th and 90th, median and high estimate, respectively) models provide bookends for these metrics.

13 HYDRO ENERGY

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES

Studying CPU and memory utilization of applications on Fujitsu A64FX and Nvidia Grace Superchip

ARM-based manycore CPU architectures are well-positioned to provide the rising memory throughput requirements of modern data intensive scientific applications in High Performance Computing (HPC). The Fujitsu A64FX CPU platform is based on the ARM v8.2A architecture, and is the processor of the flagship Japanese supercomputer - "Fugaku", which was previously ranked as the #1 supercomputer in the world according to the Top500 list. The Nvidia Grace superchip features 144 Neoverse V2 cores based on the ARMv9 architecture with 4x128b SVE2, providing exceptional computational power. The chip supports up to 480GB of memory, making it ideal for AI, machine learning, and scientific computing workloads. In this paper, we conduct a thorough performance exploration of a variety of parallel bandwidth-sensitive benchmarks and applications compiled with the native Fujitsu compiler on a Fugaku A64FX compute node and ARM (LLVM) Compiler on an NVIDIA Grace superchip compute node, engaging all the computational cores per cluster using OpenMP multithreading (assuming the cores can drive the available bandwidth). Our ultimate goals are to study the resource utilization of scientific applications and benchmarks on A64FX and Grace superchip, considering graph application scenarios ( GAP Benchmark suite) and eleven appli- cation proxies from the Rodinia heterogeneous benchmark suite (considering domains such as Data Mining, Bioinformatics, Fluid Dynamics, Pattern Recognition, etc.). Through exhaustive performance monitoring, we quantify the resource utilization of diverse OpenMP-based HPC applications on both the Fujitsu A64FX and the Nvidia Grace Superchip platforms.

benchmarking, Performance Analysis, High performan