Search NASA⌕ Search

SEARCH · Search NASA

Results for “Large-Scale, Realistic Dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

High-Fidelity, Large-Scale, Realistic Dataset Development

The final report summarizes the work performed for supporting the ARPA-E Grid Optimization Competition (Challenge 2 and Challenge 3) within the stated period. Challenge 2 For the challenge period, the main responsibility of the team is to investigate, gen- erate, and deliver parts of the data sets for the competition, based on the competition model for Challenge 2, existing data sets from Challenge 1, and data source supplied by other data set teams. Challenge 3 For the challenge period, the main responsibility of the team is to propose, create, deliver, and maintain the data format during the competition period. The data format will specify how the benchmark data will be represented and communicated to competitors. It will also specify how competitors should report back the solutions. The data format will be closely aligned with the problem formulation (maintained by the formulation team) and the solution validation process (maintained by the validation team). Our team is also responsible in investigating, generating, and delivering parts of the data sets for the competition. The data sets will be created based on the competition model for Challenge 3, existing data sets from Challenge 1 and Challenge 2, and data source supplied by other data set teams.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

RADAI: A Large-Scale Realistic Dataset for Radiation Detection Algorithm Development

Open, realistic datasets are essential for developing and benchmarking radiation detection algorithms, yet they remain scarce. The Radiological Anomaly Detection and Identification (RADAI) project was develop to create datasets that meet the training and testing needs for sophisticated radiation detection algorithms. The RADAI dataset is a large-scale synthetic resource that integrates high-fidelity Monte Carlo simulations with realistic urban scenarios to capture both background variability and source signatures. RADAI models construction-material NORM, people and vehicles, urban clutter, and dynamic environmental effects such as cosmic-ray and rain-induced transients, and they provide list-mode detector data with motion and response modeling suitable for algorithm training and evaluation. The RADAI project resulted in three publicly-released complementary datasets together with an online scoring portal for standardized performance assessment and an open software toolkit that supports data access, augmentation, model development, and evaluation. These resources enable reproducible comparisons across methods and promote rigorous studies at the scale required by contemporary machine learning. By grounding algorithm development in realistic, well-documented conditions, RADAI supports progress toward more robust detection, identification, and localization in complex urban environments.

Ghawaly, James M. [Division of Computer Science an↗

Conditional distribution estimation of building characteristics with diffusion models for urban energy modeling

Understanding current energy consumption behavior in communities is critical for informing future energy use decisions and enabling efficient energy management. Urban energy models, which are used to simulate these energy use patterns, require large datasets with detailed building characteristics for accurate outcomes. However, such detailed characteristics at the individual building level are often unknown and costly to acquire, or unavailable. Through this work, we propose using a generative modeling approach to generate realistic building attributes to fill in the data gaps and finally provide complete characteristics as inputs to energy models. Our model learns complex, building-level patterns from training on a large-scale residential building stock model containing 2.2 million buildings. We employ a tabular diffusion-based framework that is designed to handle heterogeneous (discrete and continuous) features in tabular building data, such as occupancy, floor area, heating, cooling, and other equipment details. We develop a capability for conditional diffusion, enabling the imputation of missing building characteristics conditioned on known attributes. We conduct a comprehensive validation of our conditional diffusion model, firstly by comparing the generated conditional distributions against the underlying data distribution, and secondly, by performing a case study for a Baltimore residential region, showing the practical utility of our approach. Our work is one of the first to demonstrate the potential of generative modeling to accelerate building energy modeling workflows.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Dark Energy Survey: Modeling strategy for multiprobe cluster cosmology and validation for the Full Six-year Dataset

We introduce an updated To&Krause2021 model for joint analyses of cluster abundances and large-scale two-point correlations of weak lensing and galaxy and cluster clustering (termed CL+3x2pt analysis) and validate that this model meets the systematic accuracy requirements of analyses with the statistical precision of the final Dark Energy Survey (DES) Year 6 (Y6) dataset. The validation program consists of two distinct approaches, (1) identification of modeling and parameterization choices and impact studies using simulated analyses with each possible model misspecification (2) end-to-end validation using mock catalogs from customized Cardinal simulations that incorporate realistic galaxy populations and DES-Y6-specific galaxy and cluster selection and photometric redshift modeling, which are the key observational systematics. In combination, these validation tests indicate that the model presented here meets the accuracy requirements of DES-Y6 for CL+3x2pt based on a large list of tests for known systematics. In addition, we also validate that the model is sufficient for several other data combinations: the CL+GC subset of this data vector (excluding galaxy--galaxy lensing and cosmic shear two-point statistics) and the CL+3x2pt+BAO+SN (combination of CL+3x2pt with the previously published Y6 DES baryonic acoustic oscillation and Y5 supernovae data).

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Dark energy survey: Modeling strategy for multiprobe cluster cosmology and validation for the full six-year dataset

Here, we introduce an updated To&Krause2021 model for joint analyses of cluster abundances and large-scale two-point correlations of weak lensing and galaxy and cluster clustering (termed CL+3×2 pt analysis) and validate that this model meets the systematic accuracy requirements of analyses with the statistical precision of the final Dark Energy Survey (DES) Year 6 (Y6) dataset. The validation program consists of two distinct approaches, (i) identification of modeling and parametrization choices and impact studies using simulated analyses with each possible model misspecification and (ii) end-to-end validation using mock catalogs from customized Cardinal simulations that incorporate realistic galaxy populations and DES-Y6-specific galaxy and cluster selection and photometric redshift modeling, which are the key observational systematics. In combination, these validation tests indicate that the model presented here meets the accuracy requirements of DES-Y6 for CL+3×2 pt based on a large list of tests for known systematics. In addition, we also validate that the model is sufficient for several other data combinations: the CL+GC subset of this data vector (excluding galaxy–galaxy lensing and cosmic shear two-point statistics) and the CL+3×2 pt+BAO+SN (combination of CL+3×2 pt with the previously published Y6 DES baryonic acoustic oscillation and Y5 supernovae data).

79 ASTRONOMY AND ASTROPHYSICS↗

Simulating Snow Over Sea Ice In Climate Models

We have evaluated two methods of simulating the seasonal cycle of snow over sea ice in and around the Arctic: The NCAR global climate model CCM3, with its standard snow hydrology, and the snow pack model SNTHERM, forced with hourly atmospheric output from CCM3. A new dataset providing dates for the onset of snow melt over Arctic sea ice provides a means for assessing basin-wide how well the models simulate melt onset, but contains no information on how long it then takes for all the snow to melt. Use of data from the SHEBA site provides very detailed information on the behavior of the snow before and during the melt season, but only for a very limited area. Russian drift data provide climatological data on the seasonal cycle of snow water equivalent and snow density, over multi-year sea ice in the central Arctic basin. These datasets are used to compare the two modeling methods, and to see if use of the more physically-realistic SNTHERM provides any significant improvements. Conclusions obtained so far include: 1. Both CCM3 and CCM3/SNTHERM do a good job overall of matching the onset of snow melt dataset; although CCM3/SNTHERM consistently trends to underestimate the date and CCM3 to overestimate it. 2. SHEBA and ice drift data for the Arctic show that CCM3/ SNTHERM does a better job than CCM3 at simulating the total melt period. 3. Ice drift snow density and accumulation data suggest that while providing superior results, CCM3/SNTHERM may still suffer from overly vigorous melting. 4. Both the large-scale atmospheric forcing and snow pack physical processes are important in proper simulation of the snow seasonal cycle. Ongoing work includes further diagnosis of CCM3/SNTHERM, use of more observational datasets, especially from marginal seas in the pan-Arctic, and full coupling of SNTHERM into CCM3 (work to date has all been off-line simulations).

Arnold, James E.↗

Comparison of DeePMD, MTP, GAP, ACE and MACE Machine‐Learned Potentials for Radiation‐Damage Simulations: A User Perspective

Accurate and efficient interatomic potentials are essential for molecular dynamics (MD) simulations of radiation damage, gas diffusion, and phase stability in complex ceramics such as LiAlO 2 , especially under extreme conditions relevant to tritium production. Here, we evaluate the performance of six machine-learned interatomic potentials (MLIPs), moment tensor potential (MTP), Gaussian approximation potential, deep potential (DeePMD), atomic cluster expansion (ACE), message-passing ACE (multilayer atomic cluster expansion (MACE) pretrained) and MACE (trained from-scratch), all trained on the same density functional theory dataset with inclusion of tritium. The MLIPs are benchmarked against traditional Buckingham and ReaxFF potentials in terms of energy accuracy, density predictions, thermal equilibration behavior, threshold displacement energy (E d ), tritium diffusivity, and computational cost. Among the models, MTP shows the best overall balance between efficiency and accuracy, with low force and energy errors and realistic E d values for Li and Al. The ACE and MACE (pretrained and trained from scratch) models exhibit high E d (>200 eV) and unphysical pair interactions. DeePMD underestimates Ed due to overly repulsive behavior even at equilibrium distances. All models over-estimate tritium diffusion but the pretrained MACE model behaves well during tritium-diffusion simulations up to 500 K, maintaining diffusivities in the physically consistent 10 −11 m 2 /s range. Finally, we quantify the computational cost of each potential in large-scale atomic/molecular massively parallel simulator, finding that only MTP is more efficient than traditional empirical potentials, while others are significantly more expensive. These findings explain the trade-offs between accuracy and computational cost in MLIP development and provide essential guidance for use in high-throughput radiation damage and gas diffusion simulations in nuclear ceramics.

74 ATOMIC AND MOLECULAR PHYSICS↗

Machine-learning interatomic potentials for interfaces in all-solid-state batteries: Perspectives on training data, model selection, and validation

Interfaces play a pivotal role in dictating the performance and reliability of all-solid-state batteries (ASSBs), where complex electro-chemo-mechanical phenomena at grain boundaries (GBs) and interfaces can lead to degradation and failure. Traditional atomistic simulation methods, such as first-principles calculations and classical molecular dynamics, face limitations in modeling these interfaces due to either high computational cost or insufficient transferability to the diverse atomic environments evolving at interfaces. Machine-learning interatomic potentials (MLIPs) have emerged as a transformative approach, enabling large-scale, high-accuracy simulations of disordered and chemically complex systems by leveraging the predictability of machine learning models trained on first-principles data. Recent applications of MLIPs have demonstrated their ability to capture intricate behaviors at ASSB interfaces, including ion transport, interfacial evolution, and degradation mechanisms, with accuracy and efficiency unattainable by conventional methods. This prospective paper presents comprehensive analysis and practical guidance for MLIP development for GBs and interfaces in ASSBs, with a focus on three key pillars: data generation, model selection, and validation. Here, we review the current state of MLIP applications for GBs and interfaces in both general and ASSB-specific materials, highlighting best practices and challenges in constructing diverse and representative datasets, choosing appropriate machine learning architectures, and rigorously validating model performance. We also discuss emerging strategies and opportunities for improved reliability and efficiency of MLIPs to simulate realistic interfaces in ASSBs.

Energy - Storage↗

Galaxy cluster matter profiles - I. Self-similarity, mass calibration, and observable-mass relation validation employing cluster mass posteriors

We present a study of the weak lensing inferred matter profiles ΔΣ(R) of 698 South Pole Telescope (SPT) thermal Sunyaev-Zel’dovich effect (tSZE) selected and MCMF optically confirmed galaxy clusters in the redshift range 0.25 < z < 0.94 that have associated weak gravitational lensing shear profiles from the Dark Energy Survey (DES). Rescaling these profiles to account for the mass dependent size and the redshift dependent density produces average rescaled matter profiles ΔΣ(R/R200c)/(ρcritR200c) with a lower dispersion than the unscaled ΔΣ(R) versions, indicating a significant degree of self-similarity. Galaxy clusters from hydrodynamical simulations also exhibit matter profiles that suggest a high degree of self-similarity, with RMS variation among the average rescaled matter profiles with redshift and mass falling by a factor of approximately six and 23, respectively, compared to the unscaled average matter profiles. We employed this regularity in a new Bayesian method for weak lensing mass calibration that employs the so-called cluster mass posterior P(M200|ζ̂, λ̂, z), which describes the individual cluster masses given their tSZE (ζ̂) and optical (λ̂, z) observables. This method enables simultaneous constraints on richness λ-mass and tSZE detection significance ζ-mass relations using average rescaled cluster matter profiles. We validated the method using realistic mock datasets and present observable-mass relation constraints for the SPT×DES sample, where we constrained the amplitude, mass trend, redshift trend, and intrinsic scatter. Our observable-mass relation results are in agreement with the mass calibration derived from the recent cosmological analysis of the SPT×DES data based on a cluster-by-cluster lensing calibration. Our new mass calibration technique offers a higher efficiency when compared to the single cluster calibration technique. We present new validation tests of the observable-mass relation that indicate the underlying power-law form and scatter are adequate to describe the real cluster sample but that also suggest a redshift variation in the intrinsic scatter of the λ-mass relation may offer a better description. In addition, the average rescaled matter profiles offer high signal-to-noise ratio (S/N) constraints on the shape of real cluster matter profiles, which are in good agreement with available hydrodynamical ΛCDM simulations. This high S/N profile contains information about baryon feedback, the collisional nature of dark matter, and potential deviations from general relativity.Key words: gravitational lensing: weak / galaxies: clusters: general / large-scale structure of Universe

79 ASTRONOMY AND ASTROPHYSICS↗

The DESI-Lensing Mock Challenge: large-scale cosmological analysis of 3x2-pt statistics

The current generation of large galaxy surveys will test the cosmological model by combining multiple types of observational probes. Realising the statistical promise of these new datasets requires rigorous attention to all aspects of analysis including cosmological measurements, modelling, covariance and parameter likelihood. In this paper we present the results of an end-to-end simulation study designed to test the analysis pipeline for the combination of the Dark Energy Spectroscopic Instrument (DESI) Year 1 galaxy redshift dataset and separate weak gravitational lensing information from the Kilo-Degree Survey, Dark Energy Survey and Hyper-Suprime-Cam Survey. Our analysis employs the 3x2-pt correlation functions including cosmic shear and galaxy-galaxy lensing, together with the projected correlation function of the spectroscopic DESI lenses. We build realistic simulations of these datasets including galaxy halo occupation distributions, photometric redshift errors, weights, multiplicative shear calibration biases and magnification. We calculate the analytical covariance of these correlation functions including the Gaussian, noise and super-sample contributions, and show that our covariance determination agrees with estimates based on the ensemble of simulations. We use a Bayesian inference platform to demonstrate that we can recover the fiducial cosmological parameters of the simulation within the statistical error margin of the experiment, investigating the sensitivity to scale cuts. This study is the first in a sequence of papers in which we present and validate the large-scale 3x2-pt cosmological analysis of DESI-Y1.

79 ASTRONOMY AND ASTROPHYSICS↗

Human Factors and Technologies Design to Improve User Acceptance of Pooled Rideshare for Increasing Transportation System Energy Efficiency

This multi-year project delivered a comprehensive, human-factors-driven framework to understand, model, and improve pooled rideshare (PR) adoption in the United States. Through three large-scale national survey studies involving more than 16,000 participants across multiple cities and demographic groups, the research established one of the most extensive datasets to date on user perceptions, behavioral barriers, and service expectations related to pooled rideshare. These data revealed key human factors barriers of user acceptance of PR and suggested potential actionable experience optimizations that could lead to increased PR usage. This foundational knowledge guided the development of novel human-factors models and behavioral choice models that quantify how psychological, demographic, and trip-level factors influence willingness to pool. Building on these empirical insights, the project developed advanced behavioral modeling tools, including mixed logit and integrated choice and latent variable models, to capture both observable and latent influences on PR adoption. These models significantly improved the ability to predict riders’ acceptance of pooled trips, explaining choice heterogeneity through latent constructs such as safety, service experience, privacy concerns, time sensitivity, and environmental attitudes. Together, these models provide a robust analytical foundation for designing PR systems that more effectively meet user needs. The project translated human-factors insights and behavioral models into actionable technology innovations by extending POLARIS—an agent-based, activity-based travel simulation platform—into a fully functional pooled rideshare simulation environment. New PR modules, acceptance models, and regional scenarios were implemented for Greenville, SC and Austin, TX, enabling high-fidelity validation of algorithmic strategies under realistic demand and traffic conditions. The simulation platform supported the development and evaluation of adaptive discount-based assignment algorithms, enhanced willingness-to-pay formulations, demographic-aware incentive mechanisms, and a proactive joint assignment and repositioning strategy. Simulation results demonstrated substantial gains in pooling uptake, average vehicle occupancy, energy efficiency, and fleet profitability. In Greenville, pooling adoption more than doubled, while reductions in vehicle-miles traveled and energy consumption were significant. In Austin, pooling improvements were achieved with minimal service-quality trade-offs, and profitability increased across all fleet sizes. Through this research, we developed a comprehensive understanding of the human factors barriers that limit user acceptance of pooled rideshare services. These insights enabled the design of human-factors-aware pooled rideshare technologies that more effectively address user concerns and improve adoption rates. By integrating these models into an advanced agent-based simulation framework, we demonstrated that higher adoption of pooled rideshare can lead to measurable improvements in energy efficiency and system performance. Together, these contributions establish a validated pathway from human-centered analysis to technology development and energy-saving outcomes, supporting national goals for more sustainable and efficient mobility systems.

Jia, Yunyi↗

PHASE: Personalized Head-based Automatic Simulation for Electromagnetic properties in 7T MRI

Accurate and individualized human head models are becoming increasingly important for electromagnetic (EM) simulations. These simulations depend on precise anatomical representations to realistically model electric and magnetic field distributions, particularly when evaluating Specific Absorption Rate (SAR) within safety guidelines. State of the art simulations use the Virtual Population due to limited public resources and the impracticality of manually annotating patient data at scale. Here, this paper introduces Personalized Head-based Automatic Simulation for EM properties (PHASE), an automated open-source toolbox that generates high-resolution, patient-specific head models for EM simulations using paired T1-weighted (T1w) magnetic resonance imaging (MRI) and computed tomography (CT) scans with 14 tissue labels. To evaluate the performance of PHASE models, we conduct semi-automated segmentation and EM simulations on 15 real human patients, serving as the gold standard reference. The PHASE model achieved comparable global SAR and localized SAR averaged over 10 grams of tissue (SAR-10g), demonstrating its potential as a promising tool for generating large-scale human model datasets in the future. The code and models of PHASE toolbox have been made publicly available: https://github.com/hrlblab/PHASE.

Deep learning↗

Foundation models for atomistic simulation of chemistry and materials

Conventional computational methods for modeling chemical and materials systems are limited by system size and timescale, forcing a trade-off between quantum-mechanical accuracy and the sampling needed for realistic observables. Large language and vision foundation models — pre-trained on massive datasets using transformer architectures — have revolutionized many fields. It is thus interesting to ask whether a foundation model — subject to suitable data, parameter scaling and training — could enable learned simulations of chemistry and materials. Here, in this study, we review the field of machine-learned interatomic potentials (MLIPs) and posit that scaling up large and diverse chemical and materials datasets and highly expressive architectures using advanced training strategies should result in models that are: more efficient, transferable, robust to out-of-distribution scenarios, and easier to fine-tune to a variety of downstream physical observables than models trained from scratch on small datasets corresponding to specific, targeted atomistic simulation tasks. We provide specific criteria for creating such large-scale MLIP foundation models, coordinated strategies for their development, evaluation and deployment, and highlight potential emergent capabilities that could transform predictive simulations in chemistry and materials science and accelerate discovery across multiple technological domains.

Yuan, Eric C.-Y. [University of California, Berkel↗

Spectroscopy-guided discovery of three-dimensional structures of disordered materials with diffusion models

Spectroscopy techniques such as x-ray absorption near edge structure (XANES) provide valuable insights into the atomic structures of materials, yet the inverse prediction of precise structures from spectroscopic data remains a formidable challenge. In this study, we introduce a framework that combines generative artificial intelligence models with XANES spectroscopy to predict three-dimensional atomic structures of disordered systems, using amorphous carbon (a-C) as a model system. In this work, we introduce a new framework based on the diffusion model, a recent generative machine learning method, to predict 3D structures of disordered materials from a target property. For demonstration, we apply the model to identify the atomic structures of a-C as a representative material system from the target XANES spectra. We show that conditional generation guided by XANES spectra reproduces key features of the target structures. Furthermore, we show that our model can steer the generative process to tailor atomic arrangements for a specific XANES spectrum. Finally, our generative model exhibits a remarkable scale-agnostic property, thereby enabling generation of realistic, large-scale structures through learning from a small-scale dataset (i.e. with small unit cells). Our work represents a significant stride in bridging the gap between materials characterization and atomic structure determination; in addition, it can be leveraged for materials discovery in exploring various material properties as targeted.

36 MATERIALS SCIENCE↗

Dark Energy Survey Year 3 results: Simulation-based 𝑤CDM inference from weak lensing and galaxy clustering maps with deep learning: Analysis design

Data-driven approaches using deep learning are emerging as powerful techniques to extract non-Gaussian information from cosmological large-scale structure. Here, this work presents the first simulation-based inference (SBI) pipeline that combines weak lensing and galaxy clustering maps in a realistic Dark Energy Survey Year 3 (DES Y3) configuration and serves as preparation for a forthcoming analysis of the survey data. We develop a scalable forward model based on the CosmoGridV1 suite of N-body simulations to generate over one million self-consistent mock realizations of DES Y3 at the map level. Leveraging this large dataset, we train deep graph convolutional neural networks on the full survey footprint in spherical geometry to learn low-dimensional features that approximately maximize mutual information with target parameters. These learned compressions enable neural density estimation of the implicit likelihood via normalizing flows in a ten-dimensional parameter space spanning cosmological 𝑤CDM, intrinsic alignment, and linear galaxy bias parameters, while marginalizing over baryonic, photometric redshift, and shear bias nuisances. To ensure robustness, we extensively validate our inference pipeline using synthetic observations derived from both systematic contaminations in our forward model and independent Buzzard galaxy catalogs. Our forecasts yield significant improvements in cosmological parameter constraints, achieving 2−3× higher figures of merit in the 𝛺 𝑚 − 𝑆 8 plane relative to our implementation of baseline two-point statistics and effectively breaking parameter degeneracies through probe combination. These results demonstrate the potential of SBI analyses powered by deep learning for upcoming Stage-IV wide-field imaging surveys.

Thomsen, A. [Zurich, ETH] (ORCID:0000000203099021)↗

2D reactive transport model of shale chemical weathering and biogeochemical fluxes along a mountainous hillslope, East River Watershed, Colorado: Input files and simulation results

This data package contains input files and simulation results for a two-dimensional (2D) reactive transport model used to quantitatively analyze the coupled hydrological and biogeochemical processes governing shale weathering and associated biogeochemical fluxes under realistic environmental conditions in the high-elevation East River Watershed. These data support the conclusions presented in Stolze et al. (Water Resources Research, under review), "Model-based interpretation of solute exports and carbon partitioning during shale weathering in a mountainous hillslope". The model simulates atmospheric-subsurface gas exchange, subsurface water flow, and shale weathering processes under dynamic, year-scale conditions along a shale-underlain hillslope located in the East River watershed. The simulations were performed using the PFLOTRAN flow and reactive transport code and executed on the Perlmutter supercomputer to leverage its large-scale parallel computing capabilities. The data package contains two zipped folders, "model_input_files" and "simulation_results", and one readme.txt file. "model_input_files" contains the necessary input files to run the calibrated base-base model presented in Stolze et al. (Water Resources Research, under review). "simulation_results" contains a single hdf5 file ("Output_2D_hillslope_model.h5") which includes the results of simulation performed using the base-case model. This file can be opened with HDFView 3.1.4, Python, or MATLAB. "readme.txt" contains relevant information about the base-case model and provides guidelines on how to run the associated input files provided in the folder "model_input_files". Furthermore, readme.txt provides information regarding the model results provided in "Output_2D_hillslope_model.h5" such as matrix dimensionality and output units. Field datasets used to evaluate model performance were collected at three monitoring wells located along a hillslope transect (PLM1, PLM2, and PLM3). Dissolved ion concentration data were collected from November 2016 to October 2021 for Ca, Mg, DIC, Na, K, SO4 (Dong et al., 2025 - dic_npoc_data_2014_2024.zip - DOI:10.15485/1660459; Williams et al., 2025 - anion_data_2014_2024.zip - DOI:10.15485/1668054; Dong et al., 2025 - cation_data_2014_2024.zip - DOI:10.15485/1668055). Note that we used the files named er_PLM1_xx_yy, er_PLM2_xx_yy, and er_PLM3_xx_yy where xx stands for the name of the aqueous species and yy stands for the depth where the measurements were performed. Soil water content ([0 - 1] m) and water table depth were collected from November 2016 to October 2021 (Wan et al., 2024 - Dynamic_water_table__depthsFig2b.csv and Soil_water_content_Fig4e.csv - DOI:10.15485/2322567). Gaseous CO2 concentration were collected from October 2020 to December 2021(Wan et al., 2024 - Soil_CO2_concentrations_Fig4h.csv - DOI:10.15485/2322567) Gaseous CO2 flux from the subsurface to the atmosphere were collected in the vicinity of PLM2 from October 2019 to May 2022 (Wu et al., 2025). Soil microbial biomass concentration was measured from August 2016 to June 2017 (Sorensen et al., 2019 - 2017_East_River_Pumphouse_Microbial_Biomass__1_.csv - DOI:10.15485/1577267) All field data are published as CSV files compatible with Microsoft Excel, MATLAB, and Python, or as text files. The coordinates of the monitoring wells and the CO2(g) flux sensor in the coordinate system WGS84 are: -PLM1: [38.9197710 ; -106.9492750] -PLM2: [38.9201580 ; -106.9487170] -PLM3: [38.9207843 ; -106.9483668] -PLM4: 38.9210060 ; -106.9479528] -CO2(g) flux sensor: [38.9199180 ; -106.9489906] ------------------------------------------------------------------------------------------- This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. This research used resources of the National Energy Research Scientific Computing Center (NERSC), a Department of Energy User Facility using NERSC award BER-ERCAP 23980, BER-ERCAP 28550, and BER-ERCAP 33789.

54 ENVIRONMENTAL SCIENCES↗

Long-term hydro-economic analysis tool for evaluating global groundwater cost and supply: Superwell v1.1

Abstract. Groundwater plays a key role in meeting water demands, supplying over 40 % of irrigation water globally, with this role likely to grow as water demands and surface water variability increase. A better understanding of the future role of groundwater in meeting sectoral demands requires an integrated hydro-economic evaluation of its cost and availability. Yet substantial gaps remain in our knowledge and modeling capabilities related to groundwater availability, recharge, feasible locations for extraction, extractable volumes, and associated extraction costs, which are essential for large-scale analyses of integrated human–water system scenarios, particularly at the global scale. To address these needs, we developed Superwell, a physics-based groundwater extraction and cost accounting model that operates at sub-annual temporal and at the coarsest 0.5° (≈50 km × 50 km) gridded spatial resolution with global coverage. The model produces location-specific groundwater supply–cost curves that provide the levelized cost to access different quantities of available groundwater. The inputs to Superwell include recent high-resolution hydrogeologic datasets of permeability, porosity, aquifer thickness, depth to water table, recharge, and hydrogeological complexity zones. It also accounts for well capital and maintenance costs, as well as the energy costs required to lift water to the surface. The model employs a Theis-based scheme coupled with an amortization-based cost accounting formulation to simulate groundwater extraction and quantify the cost of groundwater pumping. The result is a spatiotemporally flexible, physically realistic, economics-based model that produces groundwater supply–cost curves. We show examples of these supply–cost curves and the insights that can be derived from them across a set of scenarios designed to explore model outcomes. The supply–cost curves produced by the model show that most (90 %) nonrenewable groundwater in storage globally is extractable at costs lower than USD 0.57 m−3, while half of the volume remains extractable at under USD 0.108 m−3. The global unit cost is estimated to range from a minimum of USD 0.004 m−3 to a maximum of USD 3.971 m−3. We also demonstrate and discuss examples of how these cost curves could be used by linking Superwell's outputs with other models to explore coupled human–environmental system challenges, such as water resources planning and management, or broader analyses of multisectoral feedbacks.

Global Change Analysis Model (GCAM)↗

January and July global distributions of atmospheric heating for 1986, 1987, and 1988

Three-dimensional global distributions of atmospheric heating are estimated for January and July of the 3-year period 1986-88 from the European Center for Medium Weather Forecasts (ECMWF) Tropical Ocean Global Atmosphere (TOGA) assimilated datasets. Emphasis is placed on the interseasonal and interannual variability of heating both locally and regionally. Large fluctuations in the magnitude of heating and the disposition of maxima/minima in the Tropics occur over the 3-year period. This variability, which is largely in accord with anomalous precipitation expected during the El Nino-Southern Oscillation (ENSO) cycle, appears realistic. In both January and July, interannual differences of 1.0-1.5 K/day in the vertically averaged heating occur over the tropical Pacific. These interannual regional differences are substantial in comparison with maximum monthly averaged heating rates of 2.0-2.5 K/day. In the extratropics, the most prominent interannual variability occurs along the wintertime North Atlantic cyclone track. Vertical profiles of heating from selected regions also reveal large interannual variability. Clearly evident is the modulation of the heating within tropical regions of deep moist convection associated with the evolution of the ENSO cycle. The heating integrated over continental and oceanic basins emphasizes the impact of land and ocean surfaces on atmospheric energy balance and depicts marked interseasonal and interannual large-scale variability.

Schaack, Todd K.↗