Search NASA⌕ Search

SEARCH · Search NASA

Results for “Model intercomparison”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

PCMDI Metrics Package

The Program for Climate Model Diagnosis & Intercomparison (PCMDI) Metrics Package (PMP) is used to provide "quick-look" objective comparisons of Earth System Models (ESMs) with one another and available observations. The PMP provides a diverse suite of analysis utilities each of which produce summary statistics that gauge the consistency between climate model simulations and available observations. The primary application of the PMP is to evaluate simulations from the Coupled Model Intercomparison Project (CMIP). It can also be used to provide objective performance summaries during the model development process as well as selected research purposes.

Ullrich, PaulA [Lawrence Livermore National Labora↗

Intercomparison of flood inundation models across land use types and hydrological flood stages

Flood Inundation Mapping (FIM) model selection is a key operational decision because accurate, rapid mapping underpins early warning and resource allocation. FIM performance is context-dependent and can vary with hydrograph phase, land-use/land-cover (LULC), and the evaluation benchmark. Intercomparison studies typically assess a single near-peak snapshot against one reference dataset. Here, we provide a context-stratified intercomparison across (i) multiple hydrograph phases, (ii) LULC classes, and (iii) benchmark types, for five FIM approaches spanning a wide range of physical complexity and operational cost (TRITON, LISFLOOD-FP, HEC-RAS 2D, ARC-Curve2Flood, and OWP HAND-FIM). We use the Hurricane Matthew flood (2016) in the Neuse River Basin, North Carolina, USA, as a case study. Using high-resolution remote sensing-derived flood inundation maps, hand-labeled points, and building footprints, we assess model skill across two rising and two falling hydrograph limbs and across major LULC types. Results show that model rankings shift systematically across contexts: LISFLOOD-FP ranks highest in three of four flood phases, while TRITON leads during one rising limb phase; LISFLOOD-FP performs best in vegetated areas, whereas HEC-RAS improves relative performance in agricultural and urban areas; and benchmark choice influences conclusions, with LISFLOOD-FP performing best for flooded-building detection in the late falling limb, while TRITON ranks highest against hand-labeled points. We also report representative wall-clock runtimes for each workflow to provide use-case context for operational feasibility. Together, these results offer transferable guidance for model selection and for designing large-scale, benchmark-aware FIM intercomparison studies.

Nikrou, Parvaneh [University of Alabama]↗

Intercomparison of Deep Learning Model Architectures for Atmospheric River Prediction

With a rapid surge in the application of machine learning (ML) for a diverse range of tasks in climate science, the present study addresses a challenge for climate scientists when selecting the optimal ML or deep learning (DL) architecture for a given application. In particular, a DL intercomparison study was performed with a focus on forecasting the position of atmospheric rivers (ARs) on short-range time scales (up to 5-day lead times). AR predictions from multiple DL architectures, including various types of convolutional autoencoders and a vision transformer (ViT), were compared against ECMWF ERA5 reanalysis and hindcasts from a global climate model. DL models with similar trainable parameters were trained on ERA5 reanalysis data and AR positions derived from a thresholding algorithm to ensure a fair comparison among the DL models. Each model’s performance and accuracy in forecasting AR location and key input fields within a 5-day window were assessed using metrics of root-mean-square error, anomaly correlation, and mean intersection over union. The ViT architecture outperformed other autoencoder models in most of the metrics. Incorporating additional meteorological fields only yielded slight improvements in forecasting certain fields at longer lead times. The results also suggest that a smaller number of input time steps or smaller number of autoregressive steps can achieve better prediction skills, while also improving the overall computational efficiency. This research offers valuable insights into the strengths and weaknesses of different DL techniques for AR forecasting, hopefully guiding the development of improved models for forecasting this phenomenon.

54 ENVIRONMENTAL SCIENCES↗

Flower‐Type Organized Trade‐Wind Cumulus: A Multi‐Day Lagrangian Large Eddy Simulation Intercomparison Study

Shallow cumulus cloud fields in subtropical marine trade wind environments, particularly over the tropical Atlantic Ocean, show distinct organizational patterns. Among these, Flower‐type clouds are characterized by expansive stratiform cloud patches surrounded by regions of scattered convection. The objectives of this study were (a) to construct a case study of a time period during the EUREC 4 A/ATOMIC field campaign when Flower‐type organization was observed, (b) to evaluate the fidelity of a multi‐model ensemble of large eddy simulations of that case, and (c) to analyze the interaction between cloud and precipitation processes and mesoscale organization in the simulations. The simulations follow a quasi‐Lagrangian trajectory, allowing mesoscale features to develop over time in a domain that follows the boundary‐layer airmass. The results show a broad agreement in simulated thermodynamic properties across different LES codes, with Flower‐type cloud patches appearing within hours of each other. The consensus among models is consistent with observations made during the EUREC 4 A/ATOMIC field campaign on the specific day of interest. The cloud structure reveals three distinct peaks in the joint probability densities of cloud base and cloud top height, with the dominant peak at any given time influenced by the stage of cloud organization. The simulated cloud system evolution reveals consistent occurrence of maxima in liquid water path and rain rate before Flower reaches its maximum length scale. Targeted sensitivity tests reveal a weak relationship between Cloud Droplet Number concentration and the extent/degree/type of organization.

EUREC4A↗

Policy implications of net-zero emissions: A multi-model analysis of United States emissions and energy system impacts

Many countries, subnational jurisdictions, and companies are setting net-zero emissions goals; however, questions remain about strategies to reach these targets, policy measures, technology gaps, and economic impacts. Here, we investigate the potential policy implications of reaching economy-wide net-zero CO 2 emissions across the United States by 2050 using results from a multi-model comparison with 14 energy-economic models. Model results suggest that achieving net-zero CO 2 targets depends on policies that accelerate deployment of zero- and low-emitting technologies that have seen rapid cost reductions in recent years (including wind, solar, battery storage, and electric vehicles) as well as relatively nascent options (including carbon capture and storage, advanced biofuels, low-carbon hydrogen, advanced nuclear, and long-duration energy storage). While net-zero policies are likely to lower fossil fuel consumption, including considerable coal and petroleum reductions, achieving net-zero emissions does not necessarily mean phasing out all fossil fuels. Model results indicate that the Inflation Reduction Act’s energy and climate provisions amplify near-term decarbonization but that net-zero policies have larger impacts on long-run outcomes. Stringent climate policy can have large fiscal impacts on tax revenue and government spending—revenues from carbon pricing and subsidies for carbon removal range from 0.1 % to 3.7 % of GDP in 2050 across models. Each dollar per metric ton carbon price leads to a 0.06 % to 0.31 % reduction in economy-wide CO 2 emissions relative to a reference scenario with current policies. Spending on energy across the economy decreases relative to today for many models under reference and net-zero policies, especially as a share of GDP, due primarily to end-use electrification and energy efficiency.

54 ENVIRONMENTAL SCIENCES↗

Performance evaluation of CMIP6 models on the Arctic-Siberian Plain teleconnection affecting the East Asian heat waves

The frequency and intensity of summer heat waves in East Asia have increased sharply in recent decades, significantly impacting public health and the economy. The Arctic-Siberian Plain (ASP) teleconnection pattern has been identified as a key driver, with ASP warming amplifying atmospheric circulation patterns conducive to extreme temperatures. This study evaluates the ability of Coupled Model Inter-comparison Project phase 6 models to simulate the ASP pattern across interannual variability (IAV) and intra-seasonal variability (ISV) timescales using the Common Basis Function method. The multi-model mean shows statistically significant pattern correlations with ERA5 reanalysis, with correlation coefficients of 0.90 and 0.99 for IAV and ISV, respectively. While the ASP pattern is generally well captured, models exhibit substantial inter-model diversity in the intensity and position of anticyclonic anomalies over the ASP and East Asia. Models with ASP pattern variability similar to reanalysis better reproduce extreme East Asian temperatures, whereas those over- or underestimating ASP variability exhibit lower skill. These performance differences are related to differences in simulating key variables associated with the development of the ASP pattern. Our findings highlight the role of the ASP pattern in modulating extreme heat events, as models with improved ASP simulations align more closely with observed temperature extremes. Refining ASP representations in models could enhance seasonal heat wave predictions, improving climate adaptation strategies.

Arctic-Siberian Plain (ASP)↗

Demography, dynamics and data: building confidence for simulating changes in the world's forests

Vegetation demographic models (VDMs) are advanced tools for simulating forest responses to climate and land-use changes, and are essential for projecting carbon cycling and large-scale forest management strategies. Despite their increasing incorporation into Earth System Models, VDMs differ in their demographic assumptions, with no prior quantitative comparison of their performance. We benchmarked nine VDMs against observational data from boreal, temperate and tropical sites, assessing their accuracy in predicting tree growth, carbon turnover, biomass stocks and size distributions. Models were simulated under consistent climate conditions with postdisturbance recovery monitored for at least 420 yr. Postdisturbance carbon recovery trajectories showed significant variability while remaining within observational ranges. Initial regrowth rates varied substantially (0.03-0.60, 0.18-0.70 and 0.35-1.10 kgCm-2 yr-1 for boreal, temperate and tropical sites, respectively), influenced by each model's initial forest state. Models captured mature forest carbon content but showed compensating effects between overestimated growth and underestimated mortality rates. This first multi-model benchmarking identifies growth and mortality rates as critical calibration targets and highlights the need to refine postdisturbance establishment conditions for model development. We outline specific benchmarking variables needed to improve predictions of forest responses to environmental change.

demographic vegetation model benchmarking↗

Data for Bistline, et al. (2025) "Policy Implications of Net-Zero Emissions: A Multi-Model Analysis of United States Emissions and Energy System Impacts"

These files contain input assumptions, results, and figures associated with the Bistline, et al. (2025) article "Policy Implications of Net-Zero Emissions: A Multi-Model Analysis of United States Emissions and Energy System Impacts" in Energy and Climate Change as part of the Energy Modeling Forum 37 study. Please refer to the original paper for details.

climate policy↗

Radiochronometric discordance in cast uranium metal: a multi-laboratory intercomparison exercise

Here, the model age of a nuclear material is crucial in nuclear forensic analysis. Uranium metals with complex production histories often exhibit discordant model ages from the 230 Th– 234 U and 231 Pa– 235 U chronometers. Recent studies involving targeted uranium metal castings have enhanced our understanding of decay product behavior during casting, aiding nuclear forensic interpretation. Building on this prior work, forensics laboratories at Atomic Weapons Establishment (AWE), Lawrence Livermore National Laboratory (LLNL), and Los Alamos National Laboratory (LANL) conducted an interlaboratory comparison to investigate spatial heterogeneity in uranium metal cast under controlled conditions. Each laboratory measured samples of a mixed feedstock and its corresponding cast product. This work furthers our understanding of discordant model ages and the use of discordance as a signature to enhance confidence in interpretations of radiochronometric data for nuclear forensics.

230Th/234U↗

A Practical Probabilistic Benchmark for AI Weather Models

Since the weather is chaotic, it is necessary to forecast an ensemble of future states. Recently, multiple AI weather models have emerged claiming breakthroughs in deterministic skill. Unfortunately, it is hard to fairly compare ensembles of AI forecasts because variations in ensembling methodology become confounding and the baseline data volume is immense. We address this by scoring lagged initial condition ensembles—whereby an ensemble can be constructed from a library of deterministic hindcasts. This allows the first parameter‐free intercomparison of leading AI weather models' probabilistic skill against an operational baseline. Lagged ensembles of the two leading AI weather models, GraphCast and Pangu, perform similarly even though the former outperforms the latter in deterministic scoring. These results are elaborated upon by sensitivity tests showing that commonly used multiple time‐step loss functions damage ensemble calibration.

54 ENVIRONMENTAL SCIENCES↗

input4MIPs.CMIP6Plus.PCMDI.PCMDI-AMIP-1-1-9

CMIP6Plus Forcing Datasets (input4MIPs). This dataset PCMDI-AMIP-1-1-9 is part of input4MIPs under dataset_category: "['SSTsAndSeaIce']". More information about the dataset can be found at the following links: https://pcmdi.llnl.gov/mips/amip https://input4mips-controlled-vocabularies-cvs.readthedocs.io/en/stable/database-views/input4MIPs_source-id_CMIP7.html https://input4mips-controlled-vocabularies-cvs.readthedocs.io/en/stable/dataset-overviews/ The dataset is available at: https://esgf-node.ornl.gov/search/input4mips/?mip_era=CMIP6Plus&activity_id=input4MIPs&institution_id=PCMDI&source_id=PCMDI-AMIP-1-1-9

54 ENVIRONMENTAL SCIENCES↗

A protocol and analysis of year-long simulations of global storm-resolving models and beyond

We propose a protocol to evaluate and analyze year-long simulations of global storm-resolving models (GSRMs). The proposed protocol complements an earlier 40-day simulation protocol under the DYAMOND (DYnamics of the Atmospheric general circulation Modeled On Non-hydrostatic Domains) project to allow the analysis of the seasonal cycle and associated climatic relevant phenomena. This intercomparison aims to reveal how GSRMs, which can simulate mesoscale convective systems (MCSs) in the global domain, reproduce atmospheric large-scale structures related to convection beyond month-long simulations. The intercomparison for one-year simulations is conducted by either atmosphere-only models or atmosphere–ocean coupled models with atmospheric horizontal mesh sizes less than 5 km. We recommend the continuous four seasons from March 2020 to February 2021 as a target period for the intercomparison but with options for many groups to join more flexibly. The output variables are collected at 0.25° resolution, and archives of a small set of native grid variables are encouraged to analyze tropical cyclones and MCSs. Through the proposed global storm-resolving simulation, we will evaluate the climatological distributions of the atmospheric large-scale circulations, such as the Intertropical Convergence Zone (ITCZ), monsoon, midlatitude jets, their time evolution, and the upscale impacts on them. We present sample analyses from a one-year simulation using the 3.5 km mesh Nonhydrostatic Icosahedral Atmospheric Model (NICAM), revealing the realistic zonal contrast of tropical precipitation, no double ITCZ structure, the reasonable midlatitude jet position and intensity but a weak bias of storm track activities, and a warm bias over the Eurasia during boreal winter. We also clarify the cross-scale interaction, such as the effects of cold pools on mean precipitation over the Maritime Continent through the precipitation diurnal cycle and the effects of resolved gravity waves on midlatitude mean flows. The proposed one-year simulation protocol is referred to as the “Sendai Protocol.” This protocol is not unique or definite for evaluating GSRMs; we prospect a hierarchical set of experiments from short-term to multi-year simulations as GSRM intercomparisons.

54 ENVIRONMENTAL SCIENCES↗

Atmospheric River Detection Under Changing Seasonality and Mean-State Climate: ARTMIP Tier 2 Paleoclimate Experiments

Atmospheric rivers (ARs) are filamentary structures within the atmosphere that account for a substantial portion of poleward moisture transport and play an important role in Earth's hydroclimate. However, there is no one quantitative definition for what constitutes an atmospheric river, leading to uncertainty in quantifying how these systems respond to global change. This study seeks to better understand how different AR detection tools (ARDTs) respond to changes in climate states utilizing single-forcing climate model experiments under the aegis of the Atmospheric River Tracking Method Intercomparison Project (ARTMIP). We compare a simulation with an early Holocene orbital configuration and another with CO2 levels of the Last Glacial Maximum to a preindustrial control simulation to test how the ARDTs respond to changes in seasonality and mean climate state, respectively. We find good agreement among the algorithms in the AR response to the changing orbital configuration, with a poleward shift in AR frequency that tracks seasonal poleward shifts in atmospheric water vapor and zonal winds. In the low CO2 simulation, the algorithms generally agree on the sign of AR changes, but there is substantial spread in their magnitude, indicating that mean-state changes lead to larger uncertainty. This disagreement likely arises primarily from differences between algorithms in their thresholds for water vapor and its transport used for identifying ARs. These findings warrant caution in ARDT selection for paleoclimate and climate change studies in which there is a change to the mean climate state, as ARDT selection contributes substantial uncertainty in such cases.

Atmospheric river, paleoclimate↗

Mesoscale Convective Systems Tracking Method Intercomparison (MCSMIP): Application to DYAMOND Global km‐Scale Simulations

Abstract Global kilometer‐scale models represent the future of Earth system modeling, enabling explicit simulation of organized convective storms and their associated extreme weather. Here, we comprehensively evaluate tropical mesoscale convective system (MCS) characteristics in the DYAMOND (DYnamics of the atmospheric general circulation modeled on non‐hydrostatic domains) simulations for both summer and winter phases. Using 10 different feature trackers applied to simulations and satellite observations, we assess MCS frequency, precipitation, and other key characteristics. Substantial differences (a factor of 2–3) arise among trackers in observed MCS frequency and their precipitation contribution, but model‐observation differences in MCS statistics are more consistent across trackers. DYAMOND models are generally skillful in simulating tropical mean MCS frequency, with multi‐model mean biases ranging from −2%–8% over land and −8%–8% over ocean (summer vs. winter). However, most DYAMOND models underestimate MCS precipitation amount (23%) and their contribution to total precipitation (17%). Biases in precipitation contributions are generally smaller over land (13%) than over ocean (21%), with moderate inter‐model variability. While models better simulate MCS diurnal cycles and cloud shield characteristics, they overestimate MCS precipitation intensity and underestimate stratiform rain contributions (up to a factor of 2), particularly over land, albeit observational uncertainties exist. Additionally, models exhibit a wide range of precipitable water in the tropics compared to reanalysis and satellite observations, with many models showing exaggerated sensitivity of MCS precipitation intensity to precipitable water. The MCS metrics developed here provide process‐oriented diagnostics to guide future model development.

54 ENVIRONMENTAL SCIENCES↗

Implementation and Exploration of Parameterizations of Large-Scale Dynamics in NCAR's Single Column Atmosphere Model SCAM6

A single column model with parameterized large-scale (LS) dynamics is used to better understand the response of steady-state tropical precipitation to relative sea surface temperature under various representations of radiation, convection, and circulation. The large-scale dynamics are parametrized via the weak temperature gradient (WTG), damped gravity wave (DGW), and spectral weak temperature gradient (Spectral WTG) method in NCAR's Single Column Atmosphere Model (SCAM6). Radiative cooling is either specified or interactive, and the convective parameterization is run using two different values of a parameter that controls the degree of convective inhibition. Results are interpreted in the context of the Global Atmospheric System Studies -Weak Temperature Gradient (GASS-WTG) Intercomparison project. Using the same parameter settings and simulation configuration as in the GASS-WTG Intercomparison project, SCAM6 under the WTG and DGW methods produces erratic results, suggestive of numerical instability. However, when key parameters are changed to weaken the large-scale circulation's damping of tropospheric temperature variations, SCAM6 performs comparably to single column models in the GASS-WTG Intercomparison project. The Spectral WTG method is less sensitive to changes in convection and radiation than are the other two methods, performing qualitatively similarly across all configurations considered. Under all three methods, circulation strength, represented in 1D by grid-scale vertical velocity, is decreased when barriers to convection are reduced. This effect is most extreme under specified radiative cooling, and is shown to come from increased static stability in the column's reference radiative-convective equilibrium profile. This argument can be extended to interactive radiation cases as well, though perhaps less conclusively.

54 ENVIRONMENTAL SCIENCES↗

Modelling Tritium Production and Release at High-Energy Proton Accelerators

Tritium is a well-known byproduct of particle accelerator operations. To keep levels of tritium below regulatory limits, tritium production is actively monitored and managed at Fermilab. We plan to study tritium production in the targets, beamline components, and shielding elements of the Fermilab facilities such as NuMI, BNB, and MI-65. To facilitate the analysis, we construct a simple model and use three Monte Carlo radiation codes, FLUKA, MARS, and PHITS, to estimate the amount of tritium produced in these facilities. The analysis could also serve as an intercomparison between these code results related to tritium production. To assess the actual amounts of tritium that would be released from various materials, we employ a semi-empirical diffusion model. The results of this analysis are compared to experimental data whenever possible. This approach also helps to optimize proposed target materials with respect to the tritium production and release.

Georgobiani, Dali [Fermilab]↗

Report on the Atmospheric Temperature Changes and their Drivers (ATC) Activity 2025 Spring Meeting

The Atmospheric Temperature Change and their Drivers (ATC) Activity brings together experts interested in improving understanding of atmospheric temperature variability and trends and their representation in climate data records. ATC pursues this goal by fostering intercomparisons of atmospheric temperature datasets, providing and improving uncertainty information for climate data records, comparing observations with model simulations, assessing atmospheric temperature trends and their drivers, and documenting their efforts in review papers and assessment reports. The ATC activity convened at the Wegener Center for Climate and Global Change at the University of Graz in Graz, Austria over April 23 – 25. The purpose of the meeting was to provide updates on research and datasets related to atmospheric temperature change and variability, to identify areas that need further research, and to coordinate ongoing and future collaborations. Meeting themes included theoretical and simulated controls on atmospheric temperature, the development of new and improved atmospheric temperature datasets, and analysis of atmospheric temperature variability and trends. 18 activity members attended the meeting including 12 in-person attendees and 6 remote attendees. Four new early career activity members attended with support from APARC.

54 ENVIRONMENTAL SCIENCES↗

Forced Component Estimation Statistical Method Intercomparison Project (ForceSMIP)

Anthropogenic climate change is unfolding rapidly, yet its regional manifestation can be obscured by internal variability. A primary goal of climate science is to identify the externally forced climate response from among the noise of internal variability. Separating the forced response from internal variability can be addressed in climate models by using a large ensemble to average over different possible realizations of internal variability. However, with only one realization of the real world, it is a major challenge to isolate the forced response directly in observations. In the Forced Component Estimation Statistical Method Intercomparison Project (ForceSMIP), contributors used existing and newly developed statistical and machine learning methods to estimate the forced response over 1950–2022 within individual realizations of the climate system. Participants used neural networks, linear inverse models, fingerprinting methods, and low-frequency component analysis, among other approaches. These methods were trained using large ensembles from multiple climate models and then applied to observations. Here, we evaluate method performance within large ensembles and investigate the estimates of the forced response in observations. Our results show that many different types of methods are skillful for estimating the forced response in climate models, though the relative skill of individual methods varies depending on the variable and evaluation metric. Methods with comparable skill in models can give a wide range of estimates of the forced response pattern in observations, illustrating the epistemic uncertainty in forced response estimates. ForceSMIP gives new insights into the forced response in observations, its uncertainty, and methods for its estimation.

Climate attribution↗