Search NASA⌕ Search

SEARCH · Search NASA

Results for “data sharing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

FAIR Data Meets FAIR Software

Modern scientific research is increasingly defined by the interplay between data, software, and the workflows that connect them. Yet while the FAIR (Findable, Accessible, Interoperable, Reusable) principles have become foundational for scientific data stewardship, the same level of structure and expectation has only recently begun to extend to research software. This talk covers why and how FAIR principles are being applied to data and software to support data reuse. It outlines the gaps in current sharing norms, the growing federal emphasis on persistent identifiers and public access, and the opportunities created when datasets, computational workflows, code, and models are linked through rich, standardized metadata. Practical implementation pathways for the EIC and JLab communities are described, including datacards for structured dataset documentation and provenance-aware workflows. By aligning data lifecycle management with FAIR-aligned software practices, the scientific community can advance toward autonomous knowledge graphs, generative workflows, and high-quality, AI-ready scientific datasets.

McSpadden, Diana [Thomas Jefferson National Accele↗

SetGo: Metadata Readiness for Scientific AI Datasets

Scientific datasets intended for AI use require both computational readiness for model training and metadata readiness for discovery, sharing, and reuse. The Readiness Engine for Data Integration (REDI) addresses computational readiness, but no corresponding tool evaluates whether a dataset’s metadata are sufficiently complete, governed, and standards-compliant for publication and agent-based consumption. Existing FAIR assessors operate only on published repository records, and no single system covers FAIR compliance, licensing, provenance, governance, reproducibility, and catalog readiness together. We present SetGo, an open-source Python toolkit that assesses and repairs metadata readiness across these six dimensions before a dataset is published or archived. Applied to four scientific corpora, SetGo surfaces deficiencies that general-purpose tools do not detect: ERA5 climate metadata scores 4% on ACDD 1.3 compliance; materials datasets fail OPTIMADE species-definition requirements; and PDB-derived proteomics data carries licensing terms incompatible with standard SPDX identifiers. Guided enrichment raises overall FAIR scores from 52–57% to 81–91%, and a single setgo publish command pushes to Hugging Face Hub, CKAN, or OpenMetadata with ML Commons Croissant 1.0 metadata sidecars. To support interactive and automated workflows, SetGo integrates with coding agents powered by large language models (LLMs) through a /setgo skill that enables natural-language execution of the full assess–enrich–publish loop, with user involvement limited to supplying missing metadata values.

Wilkinson, Sean [ORNL] (ORCID:0000000214437479)↗

Radsource Mr: Mixed Reality Planning Tool For Radioactive Recovery

The RadSource MR system leverages Meta Quest 3's advanced mixed reality capabilities to create a comprehensive spatial planning platform for end-of-life sealed radioactive source recovery operations. The application utilizes the Quest 3's high-resolution passthrough cameras and spatial mapping algorithms to generate accurate 3D environmental models. Core technical components include: (1) Real-time spatial measurement algorithms calculating distances, angles, slopes, and surface areas with sub-centimeter accuracy; (2) Virtual object placement system allowing users to position digital representations of recovery equipment (trailers, containment vessels, protective barriers) within the real environment; (3) Voice recording and annotation system for hands-free documentation in protective equipment; (4) 3D mesh capture and storage capabilities for post-operation analysis and regulatory documentation. (5) Procedure documentation is available for viewing in Mixed Reality, providing an innovative and convenient way to access the information during pre-visit and inspection activities. (6) Support for screen capture for the view for real world and virtual objects together to use it later for planning. The system integrates computer vision techniques for environmental understanding, spatial mathematics for precise measurements, and human-computer interaction principles optimized for hazardous environment operations. Data persistence allows teams to save and share planning sessions across multiple stakeholders while maintaining operational security requirements.

Khadka, Rajiv [Idaho National Laboratory (INL), Id↗

Progress in end-to-end optimization of fundamental physics experimental apparata with differentiable programming

In this article we examine recent developments in the research area concerning the creation of end-to-end models for the complete optimization of measuring instruments. The models we consider rely on differentiable programming methods and on the specification of a software pipeline including all factors impacting performance — from the data-generating processes to their reconstruction and the inference on the parameters of interest — along with the careful specification of a utility function well aligned with the end goals of the experiment. Building on previous studies originated within the MODE Collaboration, we focus specifically on applications involving instruments for particle physics experimentation, as well as industrial and medical applications that share the detection of radiation as their data-generating mechanism. This report illustrates the most recent advancements in the area, and outlines, for each of the discussed applications as well as for automatic differentiation itself, ongoing and future work.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Applying a Multisector Scenario Framework to Evaluate Past and Future Public Surface Water Supply Infrastructure Strategies in Texas

Datasets supporting the index model and scenario analysis used in evaluating surface water supply strategies across different water system types in Texas. These data underpin the scenario development and application of five key indicators: Water Availability Index (WAI), Water Quality Index (WQI), Energy Requirement Index (ERI), Water Treatment Cost (WTC), and Water Infrastructure Cost (WIC). The datasets are organized by system type—stream reaches (flowlines), waterbodies, and reservoirs—and include both raw and standardized index values. The integrated datasets also provide scenario classifications (original and adjusted) based on infrastructure and planning priorities, enabling comparison across Shared Socioeconomic Pathways (SSPs). Additional strategy-level data are included to support evaluation of state-level new reservoir projects in relation to cost and availability tradeoffs. Please refer to the README file provided in Files for more details. Descriptions of the datasets are provided below. Dataset(s) Descriptions Folder: Index_model_database.zip Subfolder: Stream_reach.zip Fl_wf.csv, Fl_wq.csv, Fl_er.csv, Fl_wf_wtcUV.csv, Fl_wf_wtcnoUV.csv, Fl_allfac_wic1.csv, Fl_allfac_wic2.csvDatasets for computing WAI, WQI, ERI, WTC, and WIC for surface water systems classified as stream reaches (flowlines). Subfolder: Waterbody.zip Wb_wf.csv, Wb_wq.csv, Wb_er.csv, Wb_wf_wtcUV.csv, Wb_wf_wtcnoUV.csv, Wb_allfac_wic1.csv, Wb_allfac_wic2.csvEquivalent index model datasets for waterbodies, reflecting hydrologic and infrastructure attributes specific to impounded natural systems. Subfolder: Reservoir.zip Rs_wf.csv, Rs_wq.csv, Rs_er.csv, Rs_wf_wtcUV.csv, Rs_wf_wtcnoUV.csv, Rs_allfac_wic1.csv, Rs_allfac_wic2.csvIndex model datasets specific to regulated reservoir systems, incorporating both resource indicators and cost parameters. Folder: Integrated data.zip combined_merged_data.csv, combined_merged_data_scenario.csvDatasets integrating index model indicators (both raw and scaled) with scenario classifications, including adjustments reflecting SSP-aligned transitions and planning shifts. Folder: Additional data.zip wai_supplystrat_wic_merged.csvCurated dataset capturing proposed major reservoir-based municipal water supply strategies in Texas. Integrates site-level planning data with estimated capital infrastructure costs and water availability scores for comparative assessment.

geospatial↗

UBW (USLCI-Brightway2) [SWR-25-169]

Life cycle inventory (LCI) data are critical for robust life cycle assessment (LCA), yet many widely used datasets such as the U.S. Life Cycle Inventory (USLCI) are not natively compatible with advanced modeling frameworks like Brightway2. This work presents an automated pipeline to transform USLCI data into a fully functional Brightway2 project. The workflow performs systematic data cleaning, resolves duplicate process and exchange identifiers, and applies allocation to multi-output processes. Technosphere and biosphere flows are harmonized through unit conversions and a bridge mapping to the biosphere3 database, with comprehensive logging of missing flows and cutoff issues. The resulting Brightway2 database is validated using matrix diagnostics to ensure consistency of the technosphere, and is benchmarked via life cycle impact assessment (LCIA) methods such as ReCiPe and IPCC GWP. Outputs include reproducible CSV exports of corrected processes, elementary flows, characterization factors, and LCIA results, alongside backup utilities for project sharing. This pipeline lowers barriers for integrating USLCI data into open-source LCA workflows, enabling reproducible, validated LCA inventories within the Brightway 2 framework.

Ghosh, Tapajyoti [National Laboratory of the Rocki↗

AI Model Benchmarking for Nonproliferation Applications: Steel Thread Benchmarking Task Force Technical Report (Rev. 2)

Steel Thread is a NA-22 venture that seeks to build trustworthy, reliable AI models that can be used in a wide variety of nonproliferation tasks. A key aspect of building these models is developing appropriate benchmarks and evaluation methods, which will enable the venture to identify and adapt models to provide the most value in the nonproliferation domain. Benchmarks must be relevant to key tasks in this domain, such as question answering, information retrieval, document summarization and classification, consensus analysis, and image and data analysis. This report 1) provides an overview of benchmark design, evaluation, and challenges; 2) reviews a variety of open benchmarks, with a focus on language models and tasks; and 3) identifies benchmarks that are most relevant to Steel Thread. This report is intended to serve as a basis for further efforts to classify and evaluate benchmarks and their correlation with success on nonproliferation-specific tasks. The Steel Thread venture has defined benchmarks to be a particular combination of a dataset (or datasets) and a metric (or metrics) conceptualized as representing one or more specific tasks or sets of abilities for a specific modality. It is adopted by a research community as a shared framework for comparing methods.1 It includes 1) Data: Labeled (a designated subset not used for training, which could be all the data), 2) Metric: A way to quantify performance, 3) Task/Ability: What the benchmark is testing, 4) Protocol: A structured and repeatable evaluation process, 5) Baseline/Reference Model: For comparison; could be statistical, rule-based, SME-derived, or another model, and 6) Maintenance Plan: to update with new information over time; important for long-term utility. For further clarity, the definition includes what a benchmark, in this context, is not. It is not a corpus of training data, specific to a model (it is intended to apply to a range of models), a universal evaluation of performance, a guarantee that the ‘top’ model on the leaderboard will be the best fit for every specific use case, an all-encompassing proof of a model’s universal quality, nor is it a one-size-fits-all measure of success. It does not cover every real-world constraint (like operational, ethical, or cost considerations), a systems integration test, or a unit test. This definition was inspired by and resulted from discussions within the Steel Thread Benchmarking Task Force. This group was formed to define what we would mean as a benchmark within Steel Thread but persisted as the need to develop a thorough understanding of the large and expanding existing benchmarking space. This technical report is a result of the group’s divide and conquer approach to exploring this space. The release of benchmarks might not be progressing as quickly as model development, but it is moving very fast, as many benchmarks quickly become saturated, when state-of-the-art models score so close to the benchmark’s ceiling that their results are virtually indistinguishable. At that point, the test no longer differentiates between new systems, so researchers usually stop reporting scores as the benchmark no longer informs about improvements from the next generation of models. In the OpenAI announcement of GPT-5, they reported results on six flagship public benchmarks (AIME 2025, SWE-bench Verified, Aider Polyglot, MMMU, HealthBench Hard, GPQA) but the full system-card covers roughly thirty-five separate evaluations, comprising hundreds of test task items in total. There have been some efforts to summarize benchmarks in specific fields, like for text-to-image generation, but these surveys have had a narrow methodology scope. Therefore, a comprehensive survey of all benchmarks or even all benchmarks that could be relevant to Steel Thread is outside of the scope of this report. We chose some specific benchmarks to investigate in detail.

97 MATHEMATICS AND COMPUTING↗

Mitigating Data Center Impact on Grid Stability: A Coordinated Control Strategy Using Verrus StabiliGrid Architecture

Large data centers, which now represent a significant and growing share of the total U.S. grid load, can inadvertently destabilize the electrical grid when they disconnect simultaneously during brief voltage disturbances. The July 10, 2024, Eastern Interconnection incident, in which a sub-100-millisecond transmission fault triggered the cascading loss of approximately 1,500 MW of data center load, illustrates this vulnerability. While commercial battery energy storage systems (BESS) deployed in data centers provide device-level fault ride-through per IEEE 1547, they lack coordination with facility protection logic and uninterruptible power supplies (UPS), limiting their effectiveness as grid-stabilizing assets. This report presents the Verrus StabiliGrid architecture, a coordinated control framework that integrates BESS, UPS, and point-of-interconnection (POI) protection settings to enable data centers to ride through both undervoltage and overvoltage grid contingencies without disconnecting. The four-step strategy encompasses: (1) high-resolution power quality monitoring to detect the grid state during events such as undervoltage, overvoltage, underfrequency, and overfrequency; (2) POI protection settings that allow for extended ride-through and grid-connected operation during grid contingencies; (3) grid state-driven autonomous dispatch of assets to improve grid resilience by reducing power draw during undervoltage or absorbing more power during overvoltage events; and (4) coordinated post-recovery dispatch of data center assets to restore firm load to pre-contingency levels. Validated through controller-hardware-in-the-loop (C-HIL) simulations at the National Laboratory of the Rockies, results show grid import restoration to pre-fault levels within 100 milliseconds of voltage recovery. This work advances the ability of data centers to transition from passive, disturbance-sensitive loads to active participants in grid stability, a capability increasingly required by emerging NERC and ERCOT regulatory frameworks.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Extending Component Lifetime And Improving Inverter Reliability (ECLAIIR)

Inverter reliability remains one of the most persistent challenges limiting the performance, availability, and economic viability of utility‑scale photovoltaic (PV) plants. Industry data consistently show that inverters account for the highest share of corrective maintenance events and unplanned outages across PV fleets. These failures result in energy losses, increased O&M costs, and reduced confidence in long‑term solar asset performance. Motivated by these challenges, this project—Extending Component Lifetime and Improving Inverter Reliability (ECLAIIR)—was undertaken to systematically investigate inverter degradation and failure mechanisms, develop predictive maintenance capabilities, and establish data‑driven pathways to improve service life and reduce the Levelized Cost of Energy (LCOE) for large‑scale PV systems. The primary goal of the project was to identify pre‑failure signatures in string inverters using both lab‑based accelerated lifetime testing and field‑based data and to develop predictive maintenance algorithms that can anticipate inverter faults before they occur. Through collaboration with inverter testing laboratory, solar PV plant owner, and failure‑analysis experts, the project advanced the technical understanding of inverter reliability. By instrumenting inverters with thermistors, humidity sensors, power‑quality meters, and acoustic sensors, the research established how multiple sensing modalities can reliably detect deviations from normal behavior hours to days before failure. These findings substantially enhance scientific understanding of inverter failure kinetics and provide the PV industry with the most comprehensive cross‑OEM characterization of early‑stage failure indicators reported to date. Technically, the project demonstrated the effectiveness of predictive maintenance by developing and validating the PreDICT (Predictive Diagnostics of PV Inverters Using Condition Monitoring and Trend Analysis) framework—a multi‑layer diagnostic architecture combining peer‑to‑peer analytics, historical trend modeling, and advanced machine‑learning techniques such as the Sequential Conditional Variational Autoencoder (SCVAE). This predictive model achieved more than 90% accuracy in detecting pre‑failure conditions and provided up to four days of lead time before inverter failure in field scenarios. Economically, the project’s LCOE analysis showed that predictive maintenance can reduce lifetime energy losses and minimize corrective maintenance interventions. Modeling indicated that, depending on inverter failure rates and replacement timelines, predictive maintenance can significantly reduce LCOE impacts associated with inverter downtime: from as high as 19.4% under conventional maintenance strategies to 0.1%–10.17% when predictive analytics are adopted. These results confirm that predictive maintenance is both technically feasible and economically advantageous for utilities and plant operators. The project’s findings also have broad public benefit. By improving inverter reliability and reducing downtime, predictive maintenance directly increases electricity generation from existing PV assets. Enhanced reliability lowers operational costs for utilities, which can translate over time into lower energy costs for consumers. Furthermore, the project’s technical publications, conference presentations, and industry workshops ensure that knowledge gained is shared broadly across the solar industry, supporting workforce development and enabling utilities of all sizes to adopt modern asset‑health monitoring practices. The retrofitting case study and service‑life prediction framework further support informed decision‑making for aging PV fleets, helping operators extend system life and reduce electronic waste. In summary, the ECLAIIR project significantly advanced the state of knowledge on inverter degradation, demonstrated the technical and economic value of predictive maintenance, and delivered actionable tools and insights that support more reliable, cost‑effective, and sustainable PV plant operation. The outcomes of this project will continue to inform utility practices, guide inverter design improvements, and strengthen the long‑term performance of solar assets nationwide.

14 SOLAR ENERGY↗

Instantiation of the Damara Tern Platform for Advanced Materials and Manufacturing Technologies (AMMT) Program Collaborative Data Management

This work package focused on deploying an instance of the Damara Tern platform to support AMMT collaborative research activities. The objectives were to provide selected AMMT collaborators with access to a shared environment for capturing operations, trackables, and associated metadata, and to implement data entry functionalities that reflect site-specific procedures. Key activities included creating configurable, schema-driven entry forms and validating the data collection process. The report details the deployment process, the platform infrastructure, and the implemented data entry workflows, providing a reference for end users and establishing a foundation for future production-scale deployments.

36 MATERIALS SCIENCE↗

N 2 Onet: a global collaborative network facilitating advances in measurement, modeling, and mitigation of agricultural soil nitrous oxide emissions

Nitrogen (N) fertilizer supports global food production, but its use and overuse drive emissions of nitrous oxide (N 2 O), a potent and long-lived greenhouse gas. Understanding the drivers of N 2 O fluxes remains elusive, making it difficult to predict emissions in time and space and to develop and evaluate ways to lower emissions through management. Major scientific uncertainties underlying the understanding of the drivers of N 2 O fluxes identified in a workshop of N 2 O emissions experts include poor process-based understanding of controls on soil N 2 O emissions in the field; insufficient data to reduce uncertainty in N 2 O budgets from the field to regional scales, including N 2 O emission measurements and importantly, field-scale N balances; and high uncertainty in model predictions of soil N 2 O emissions across environmental and management conditions. To reduce these uncertainties, we present the concept of N 2 Onet, a global collaborative initiative to accelerate advances in N 2 O measurement, analyses, and mitigation. N 2 Onet will serve as an observational network of supersites with multi-scale measurements; a database hub for N 2 O flux and ancillary data; and a catalyst for community building, information sharing, and training. By coalescing and coordinating the global community of researchers, N 2 Onet will provide a roadmap for reducing N 2 O emissions from agriculture worldwide.

54 ENVIRONMENTAL SCIENCES↗

Network-Aware and Welfare-Maximizing Dynamic Pricing for Energy Sharing

The proliferation of behind-the-meter (BTM) distributed energy resources (DER) within the electrical distribution network presents significant supply and demand flexibilities, but also introduces operational challenges such as voltage spikes and reverse power flows. In response, this paper proposes a network-aware dynamic pricing framework tailored for energy-sharing coalitions that aggregate small, but ubiquitous, BTM DER downstream of a distribution system operator's (DSO) revenue meter that adopts a generic net energy metering (NEM) tariff. By formulating a Stackelberg game between the energy-sharing market leader and its prosumers, we show that the dynamic pricing policy induces the prosumers toward a network-safe operation and decentrally maximizes the energysharing social welfare. The dynamic pricing mechanism involves a combination of a locational ex-ante dynamic price and an ex-post allocation, both of which are functions of the energy sharing's BTM DER. The ex-post allocation is proportionate to the price differential between the DSO NEM price and the energy-sharing locational price. Simulation results using real DER data and the IEEE 13-bus test systems illustrate the dynamic nature of network-aware pricing at each bus, and its impact on voltage.

aggregates↗

Network-Aware and Welfare-Maximizing Dynamic Pricing for Energy Sharing: Preprint

The proliferation of behind-the-meter (BTM) distributed energy resources (DER) within the electrical distribution network presents significant supply and demand flexibilities, but also introduces operational challenges such as voltage spikes and reverse power flows. In response, this paper proposes a network-aware dynamic pricing framework tailored for energy-sharing coalitions that aggregate small, but ubiquitous, BTM DER downstream of a distribution system operator's (DSO) revenue meter that adopts a generic net energy metering (NEM) tariff. By formulating a Stackelberg game between the energy-sharing market leader and its prosumers, we show that the dynamic pricing policy induces the prosumers toward a network-safe operation and decentrally maximizes the energysharing social welfare. The dynamic pricing mechanism involves a combination of a locational ex-ante dynamic price and an ex-post allocation, both of which are functions of the energy sharing's BTM DER. The ex-post allocation is proportionate to the price differential between the DSO NEM price and the energy sharing locational price. Simulation results using real DER data and the IEEE 13-bus test systems illustrate the dynamic nature of network-aware pricing at each bus, and its impact on voltage.

energy communities↗

The CanBikeCO Full Pilot: Long-Term Results and Analysis From an E-Bike Program in Colorado, USA

Personal micromobility devices like bicycles, e-bikes, and scooters are low- or zero-energy alternatives to single-occupancy vehicles. However, a lack of data has led to a dearth of data-driven research on personally owned e-bike usage. We present longitudinal findings from the CanBikeCO program, focused on e-bike adoption and use across demographics, trip characteristics, and geographies in the state of Colorado. CanBikeCO recorded travel survey data from low-income individuals provided with personal e-bikes by the Colorado Energy Office in six communities across Colorado from July 2021 to December 2022. The data were collected using a custom instance of the National Renewable Energy Laboratory OpenPATH platform, which combines passive data collection with semantic information such as trip mode and purpose labels. To our knowledge, there are no prior travel survey data on personally owned e-bikes with this range and scope. Insights from this unique dataset include: (i) work trips were 17% more likely than average trips to be taken on an e-bike, (ii) e-bikes were most often reported to replace cars (34% of e-bike trips) and other personal micromobility devices (22%), and (iii) participants favored walking for trips less than 1 mile, e-bikes for trips of 1-3 miles, and e-bikes, cars, or shared rides for trips of 3-20 miles. The data used to generate these results have been made available in the Transportation Secure Data Center. We find e-bike use is appealing across age groups and may be related to characteristics of land use, urban form, occupation, income, and car ownership. We conclude for this population that the energy demand added by e-bike use (induced demand and replacing non-motorized modes) is outweighed by the reduction in energy demand from replacement of single-occupancy vehicle trips with e-bike trips. Our findings suggest considerable potential for energy savings from personal e-bike ownership.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Produced water sharing: Improved economics and reduced community impact – A Pennsylvania case study

Here, hydraulic fracturing for oil and gas extraction from unconventional reservoirs is water intensive. Between water sourced for fracturing purposes and water present in rock formations, operators often produce a greater volume of water than oil or gas. Historically, this surplus of produced water has mainly been disposed of via deep injection wells. Rising disposal costs, seasonally limited water availability, and concerns over induced seismicity have incentivized produced water recycling practices, where an operator uses produced water for hydraulic fracturing operations. The logistical challenges associated with produced water recycling have also encouraged operators to adopt ad-hoc water exchange practices, in which competing operators will exchange produced water for mutual cost savings. In this paper, we investigate the potential benefit from systematic produced water exchange among operators in Northeastern Pennsylvania. We leverage PARETO, a free and open-source modeling framework for produced water management optimization, to quantify the benefits of water exchange practices. In an example drawn from FracFocus data, we find that the adoption of systematic water sharing could improve produced water recycling rates from 49.2% to 99%, decreasing operating and trucking costs.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Historical and Future Global Irrigation Energy Consumption by Fuel and Region

Irrigation energy use is a significant component of agricultural production costs, contributing directly to the energy and emissions intensity of crop production and ultimately to food prices. Understanding the existing structure of irrigation energy consumption help achieve food-energy-water security and environmental goals. We present a comprehensive global data set detailing country-level irrigation energy consumption, emphasizing the comparative use of electric, diesel, and emerging solar pumps. To our knowledge, no such data set exists. We draw from a literature review to develop a logistic transformed regression model to estimate the shares of fuel sources for irrigation across countries over historical years to construct a global data set of country-level irrigation energy consumption by multiple fuel sources. Additionally, we compare our estimates of irrigation energy use with agricultural energy use as reported by the International Energy Agency and other external sources. We then use this data to project future irrigation energy use with the Global Change Analysis Model, which is a multisector dynamics model, to showcase the usage of this data set. Projections under the reference scenario show a global shift in fuel types for irrigation pumping, while patterns vary across regions, with India and Pakistan leading in solar-powered irrigation growth and countries like the USA and China continuing to rely primarily on grid electricity. This data set provides a resource to understand the role of irrigation fuel choices within the broader energy sector, as well as the connected agricultural, land use, and water sectors under alternative future scenarios, enabling informed decision making toward efficient agricultural practices.

Global Change Analysis Model (GCAM)↗

Multi-task Parallelism for Robust Pre-training of Graph Foundation Models on Multi-source, Multi-fidelity Atomistic Modeling Data

Graph foundation models using graph neural networks promise sustainable, efficient atomistic modeling. To tackle challenges of processing multi-source, multi-fidelity data during pre-training, recent studies employ multi-task learning, in which shared message passing layers initially process input atomistic structures regardless of source, then route them to multiple decoding heads that predict data-specific outputs. This approach stabilizes pre-training and enhances a model’s transferability to unexplored chemical regions. Preliminary results on approximately four million structures are encouraging, yet questions remain about generalizability to larger, more diverse datasets and scalability on supercomputers. We propose a multi-task parallelism method that distributes each head across computing resources with GPU acceleration. Implemented in the open-source HydraGNN architecture, our method was trained on over 24 million structures from five datasets and tested on the Perlmutter, Aurora, and Frontier supercomputers, demonstrating efficient scaling on all three highly heterogeneous super-computing architectures.

Lupo Pasini, Massimiliano [ORNL] (ORCID:0000000249↗

Daily water stable isotopes, transpiration, and matrix potential data for an aspen and engelmann stand in the East River Watershed (version 2)

We provide daily stable isotope (2H & 18O) ratios in soil water and xylem (plant stem) water, as well as the sap flow (transpiration) and the soil's matric potential at a forested site near Gothic, Colorado, in the East River catchment. We measured the stable isotopic composition of the transpiration and the daily transpiration flux sum of three aspen and three engelmann spruce. In both forest stands, we installed a soil profile and measured the soil matric potential at 15, 30, and 60 cm depth as well as the stable isotopes of soil pore water at 5, 10, 30, 60, and 90 cm depths. All isotope measurements were done in situ via vapor probes connected to a cavity ring down spectrometer (Picarro L1240i).We further report the daily meteorological data observed at billy barr near our study site. We also provide for each tree the relative share of root water uptake derived from the isotope measurements via a Bayesian mixing model (MixSIAR).The daily data is provided as a time series in "Iso_MP_Sap_DataDaily_ESSDiveUpload.csv" and the units are provided in "dd.csv"; the location of the instrumented trees and soil profiles are given as latitude and longitude coordinates saved as CSV and KMZ files; and a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata.The data was gathered to investigate the short-term changes of the water sources (i.e., variation of root water uptake from different soil depths) of the studied subalpine trees.Update 07/16/2025: The relative and absolute plant water uptake depths were grouped to ensure that the MixSIAR model was applied with endmembers that differed in their d2H value by at least 3 permill and at least 1 permill for d18O. Whenever the difference between observed d2H values for two or more probes at neighboring depths was less than 3 permill, we used the average value for the source water endmember. For days at which probes that were not next to each other measurements did not differ at least 3 permill, the average of all probes between these two depths was used as the water source endmember representing the depth range between these two probes.

54 ENVIRONMENTAL SCIENCES↗