Search NASA⌕ Search

SEARCH · Search NASA

Results for “model skill”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Can Remotely Sensed Snow Disappearance Explain Seasonal Water Supply?

Understanding the relationship between remotely sensed snow disappearance and seasonal water supply may become vital in coming years to supplement limited ground based, in situ measurements of snow in a changing climate. For the period 2001–2019, we investigated the relationship between satellite derived Day of Snow Disappearance (DSD)—the date at which snow has completely disappeared—and the seasonal water supply, i.e., the April—July total streamflow volume, for 15 snow dominated basins across the western U.S. A Monte Carlo framework was applied, using linear regression models to evaluate the predictive skill—defined here as a model’s ability to accurately predict seasonal flow volumes—of varied predictors, including DSD and in situ snow water equivalent (SWE), across a range of spring forecast dates. In all basins there is a statistically significant relationship between mean DSD and seasonal water supply (p ≤ 0.05), with mean DSD explaining roughly half of the variance. Satellite-based model skill improves later in the forecast season, surpassing the skill of in-situ-based (SWE) models in skill in 10 of the 15 basins by the latest forecast date. We found little to no correlation between model error and basin characteristics such as elevation and the ratio of snow water equivalent to total precipitation. Despite a relatively short data record, this exploratory analysis shows promise for improving seasonal water supply prediction, in particular for snow dominated basins lacking in situ observations.

snow remote sensing↗

The Subseasonal Experiment (SubX): A Multi-Model Subseasonal Prediction Experiment

SubX is a multi-model subseasonal prediction experiment designed around operational requirements with the goal of improving subseasonal forecasts. Seven global models have produced seventeen years of retrospective (re-) forecasts and more than a year of weekly real-time forecasts. The re-forecasts and forecasts are archived at the Data Library of the International Research Institute for Climate and Society, Columbia University, providing a comprehensive database for research on subseasonal to seasonal predictability and predictions. The SubX models show skill for temperature and precipitation three weeks ahead of time in specific regions. The SubX multi-model ensemble mean is more skillful than any individual model overall. Skill in simulating the Madden-Julian Oscillation (MJO) and the North Atlantic Oscillation (NAO), two sources of subseasonal predictability, is also evaluated with skillful predictions of the MJO four weeks in advance and of the NAO 2 weeks in advance. SubX is also able to make useful contributions to operational forecast guidance at the Climate Prediction Center. Additionally, SubX provides information on the potential for extreme precipitation associated with tropical cyclones which can help emergency management and aid organizations to plan for disasters. (Capsule Summary) A research to operations project in service of developing better operational subseasonal forecasts.

Precipitation↗

The Second Phase of the Global Land Atmosphere Coupling Experiment (GLACE-2): Impact of Land Initialization on Subseasonal Forecasts

The recently-completed second phase of the Global Land-Atmosphere Coupling Experiment (GLACE-2) focused on quantifying, for boreal summer, the subseasonal (out to two months) forecast skill for precipitation and air temperature that can be derived from the realistic initialization of land surface states, notably soil moisture. An overview of the multi-institutional numerical experiment is described, along with a determination and characterization of multi-model "consensus" skill. The models show modest but significant land-derived skill in predicting air temperatures out to two months, especially where the rain gauge network is dense. Given that precipitation is the chief driver of soil moisture, and thereby assuming that rain gauge density is a reasonable proxy for the adequacy of the observational network contributing to soil moisture initialization, this result indeed highlights the potential contribution of enhanced observations to prediction. Land-derived precipitation forecast skill is much weaker than that for air temperature. The skill for predicting air temperature, and to some extent precipitation, increases with the magnitude of the initial soil moisture anomaly. GLACE-2 results are examined further to provide insight into the asymmetric impacts of wet and dry soil moisture initialization on skill.

Koster, Randal↗

Effects of Resolution and Spectral Nudging in Simulation the Effects of Wintertime Atmospheric River Landfalls in the Western US

Landfalling atmospheric rivers (ARs) play a crucial role in the climate of the US Pacific coast region as they are frequently related with heavy precipitation and flash flooding events. Thus, the capability of climate models to accurately simulate AR landfalls and their key hydrologic effects is an important practical concern for WUS, from flood forecasting to future water resources projections. In order to examine the effects of model configuration, including the resolution and spectral nudging, in simulating the climatology of key weather events in the conterminous US, a NASA team has performed a hindcast experiment using the GEOS5 global and the NU-WRF regional models for Nov 1999 - Oct 2010. This study examines the skill of these hindcasts, with different models and their configurations, in simulating key footprints of landfalling ARs in the WUS region. Using an AR-landfall chronology based on the vertically-integrated water vapor flux calculated from the MERRA2 reanalysis, we have analyzed the observed and simulated precipitation and temperature anomalies associated with wintertime AR landfalls along the US Pacific coast. Model skill is measured using metrics including regional means, a skill score based on correlations and mean-square errors, and Taylor diagrams in four WUS Bukovsky regions. Results show that the AR-related anomalies of precipitation is more reliable than of surface temperatures. Model skill also varies according to regions. The AR temperature anomalies are well simulated in most of the WUS region except PNW. For precipitation, simulations with finer spatial resolution tend to generate larger spatial variability and agree better with the PRISM data in most regions. Such a resolution dependence of spatial variability is not found for temperatures; e.g., the MERRA2 reanalysis often outperforms, with similar spatial variability and higher pattern correlations with the PRISM data, finer-resolution NU-WRF runs in simulating temperature variations within subregions. Results from this study will be summarized to assist future (regional) climate experiments for climate change impact assessments and developing adaptationmitigation strategies, the key elements of the National Climate Assessment.

Simulation↗

Informing Robust Functional Relationship Benchmarks: An Evaluation of the Temperature Sensitivity of Ecosystem Respiration Across the Arctic-Boreal Region

During land model development, simulated carbon dynamics are often benchmarked against observational data sets to evaluate model performance. Functional relationship benchmarks are the relationship between a driving variable (e.g., temperature) and a response variable (e.g., ecosystem respiration) and are a promising tool for assessing model performance by evaluating modeled sensitivities to changing environmental conditions. However, observed functional relationships can be influenced by choices made during data collection and throughout the benchmarking process, impacting the inferred skill of land models. To avoid misrepresenting a model's true performance, it is necessary to systematically evaluate best practices when constructing functional relationship benchmarks. We developed a set of guidelines for constructing functional relationship benchmarks, considering the choice of data set, number of daily observations, temporal extent, and temporal resolution across Alaska and Canada over a 20-year period from 2001 to 2020. The temperature sensitivity of ecosystem respiration from observations, evaluated through an apparent Q 10 , is highly variable both spatially and as a result of the data processing approach applied in the benchmark formation. When benchmarking 13 models from the Warming Permafrost Model Intercomparison Project (WrPMIP), the range in inferred model skill is substantially impacted by the choices applied in constructing functional relationship benchmarks. The inferred performance of a given model is most sensitive to the number of daily observations and temporal extent, followed by choice of benchmark data set and temporal averaging. Results from this analysis can guide the development of consistent and robust functional relationships for future model evaluation studies.

Poe, Jeralyn [Northern Arizona University, Flagsta↗

New framework for benchmarking decadal predictions leveraging the PCMDI Metric Package with interactive visualization

Reliable climate predictions across multiple timescales are increasingly critical as climate-related risks continue to rise. With the growing number and diversity of climate prediction systems, systematic intercomparison has become essential. Here, we present a comprehensive evaluation framework based on the PCMDI Metric Package to assess the performance of multiple decadal climate prediction systems. Unlike uninitialized simulations, initialized predictions exhibit bias and predictive skill that evolve with forecast lead time. To address this, we introduce (1) model-by-lead-time portrait plots, which efficiently summarize metrics of global temperature, precipitation, and Arctic/Antarctic sea-ice extent, and (2) an HTML-based interactive visualization platform that provides detailed regional and seasonal diagnostics of model bias, skill scores, and ensemble spread for each model and lead time. Comparisons with uninitialized simulations further quantify the relative impacts of initialization and external forcing on prediction skill. The proposed framework provides a scalable and transparent approach for multi-model climate prediction assessments and can be readily extended to a wide range of operational and research forecasting systems.

54 ENVIRONMENTAL SCIENCES↗

The effect of accuracy, conservation and filtering on numerical weather forecasting

Considerations leading to the numerical design of the GLAS fourth-order global atmospheric model are discussed, including changes recently introduced into the model. The computation time and memory requirements for the fourth-order model are similar to those of the present second-order GLAS model with the same 4 deg latitude, 5 deg longitude, and 9 vertical-level resolution. However, the fourth-order model forecast skill is significantly better than that of the current GLAS model, and after three days it is comparable to the 2.5 by 3 deg version of the GLAS model in the sea level pressure maps, and has less phase errors in the 500 mb maps.

Kalnay-Rivas, E.↗

Development and Testing of the Variable Vertical Resolution Fourth Order GCM

The vertical coordinate of the Fourth Order Model has been generalized so that the model can now run with an arbitrary number of vertical layers and so that the thicknesses of these layers can be arbitrarily specified (in the sigma coordinate). This Variable Vertical Resolution (VVR) version of the Fourth Order Model will soon replace the current production model. To assess the skill of the VVR model, it has been run with 9 equally spaced layers and compared with the current production model. In two Northern Hemispheric winter cases and one summer case, the two models were virtually identical in forecast skill for 6 to 7 days. After that the VVR model was slightly better in the winter cases and the production model was slightly better in the summer case. The only exception to this was that after 2 days the production model gave slightly more skillful 500 mb forecasts in the tropics for the summer case.

Helfand, H. M.↗

Stratospheric wind errors, initial states and forecast skill in the GLAS general circulation model

Relations between stratospheric wind errors, initial states and 500 mb skill are investigated using the GLAS general circulation model initialized with FGGE data. Erroneous stratospheric winds are seen in all current general circulation models, appearing also as weak shear above the subtropical jet and as cold polar stratospheres. In this study it is shown that the more anticyclonic large-scale flows are correlated with large forecast stratospheric winds. In addition, it is found that for North America the resulting errors are correlated with initial state jet stream accelerations while for East Asia the forecast winds are correlated with initial state jet strength. Using 500 mb skill scores over Europe at day 5 to measure forecast performance, it is found that both poor forecast skill and excessive stratospheric winds are correlated with more anticyclonic large-scale flows over North America. It is hypothesized that the resulting erroneous kinetic energy contributes to the poor forecast skill, and that the problem is caused by a failure in the modeling of the stratospheric energy cycle in current general circulation models independent of vertical resolution.

Tenenbaum, J.↗

Basin Scale Estimates of Evapotranspiration Using GRACE and other Observations

Evapotranspiration is integral to studies of the Earth system, yet it is difficult to measure on regional scales. One estimation technique is a terrestrial water budget, i.e., total precipitation minus the sum of evapotranspiration and net runoff equals the change in water storage. Gravity Recovery and Climate Experiment (GRACE) satellite gravity observations are now enabling closure of this equation by providing the terrestrial water storage change. Equations are presented here for estimating evapotranspiration using observation based information, taking into account the unique nature of GRACE observations. GRACE water storage changes are first substantiated by comparing with results from a land surface model and a combined atmospheric-terrestrial water budget approach. Evapotranspiration is then estimated for 14 time periods over the Mississippi River basin and compared with output from three modeling systems. The GRACE estimates generally lay in the middle of the models and may provide skill in evaluating modeled evapotranspiration.

Rodell, M.↗

ENSO Effect on East Asian Tropical Cyclone Landfall via Changes in Tracks and Genesis in a Statistical Model

Improvements on a statistical tropical cyclone (TC) track model in the western North Pacific Ocean are described. The goal of the model is to study the effect of El Nino-Southern Oscillation (ENSO) on East Asian TC landfall. The model is based on the International Best-Track Archive for Climate Stewardship (IBTrACS) database of TC observations for 1945-2007 and employs local regression of TC formation rates and track increments on the Nino-3.4 index and seasonally varying climate parameters. The main improvements are the inclusion of ENSO dependence in the track propagation and accounting for seasonality in both genesis and tracks. A comparison of simulations of the 1945-2007 period with observations concludes that the model updates improve the skill of this model in simulating TCs. Changes in TC genesis and tracks are analyzed separately and cumulatively in simulations of stationary extreme ENSO states. ENSO effects on regional (100-km scale) landfall are attributed to changes in genesis and tracks. The effect of ENSO on genesis is predominantly a shift in genesis location from the southeast in El Nino years to the northwest in La Nina years, resulting in higher landfall rates for the East Asian coast during La Nina. The effect of ENSO on track propagation varies seasonally and spatially. In the peak activity season (July-October), there are significant changes in mean tracks with ENSO. Landfall-rate changes from genesis- and track-ENSO effects in the Philippines cancel out, while coastal segments of Vietnam, China, the Korean Peninsula, and Japan show enhanced La Nina-year increases.

simulation↗

Predicting September Arctic Sea Ice: A Multimodel Seasonal Skill Comparison

This study quantifies the state of the art in the rapidly growing field of seasonal Arctic sea ice prediction. A novel multimodel dataset of retrospective seasonal predictions of September Arctic sea ice is created and analyzed, consisting of community contributions from 17 statistical models and 17 dynamical models. Prediction skill is compared over the period 2001–20 for predictions of pan-Arctic sea ice extent (SIE), regional SIE, and local sea ice concentration (SIC) initialized on 1 June, 1 July, 1 August, and 1 September. This diverse set of statistical and dynamical models can individually predict linearly detrended pan-Arctic SIE anomalies with skill, and a multimodel median prediction has correlation coefficients of 0.79, 0.86, 0.92, and 0.99 at these respective initialization times. Regional SIE predictions have similar skill to pan-Arctic predictions in the Alaskan and Siberian regions, whereas regional skill is lower in the Canadian, Atlantic, and central Arctic sectors. The skill of dynamical and statistical models is generally comparable for pan-Arctic SIE, whereas dynamical models outperform their statistical counterparts for regional and local predictions. The prediction systems are found to provide the most value added relative to basic reference forecasts in the extreme SIE years of 1996, 2007, and 2012. SIE prediction errors do not show clear trends over time, suggesting that there has been minimal change in inherent sea ice predictability over the satellite era. Overall, this study demonstrates that there are bright prospects for skillful operational predictions of September sea ice at least 3 months in advance.

54 ENVIRONMENTAL SCIENCES↗

Subseasonal-to-Seasonal Hindcast Skill Assessment of Ridging Events Related to Drought Over the Western United States

Persistent atmospheric ridging events centered near the western United States are associated with widespread precipitation deficits and meteorological drought. Due to the relatively low skill of- dynamical models in forecasting precipitation on subseasonal‐to‐seasonal (S2S) time scales across this region, forecasts of ridging are explored in this study as a potential bridge for early warning drought prediction. To assess skill, we evaluate deterministic and probabilistic S2S hindcasts (out to 6 week lead time) of ridging events and geopotential height anomalies over the western United States in five ensemble hindcast systems. Prediction skill for ridging events is shown to be highly variable across models. For some models, longer‐time‐averaged patterns of geopotential height anomalies across the first 6 weeks can be skillfully simulated, when evaluated probabilistically against climatology. The most skillful models show modest skill in forecasting above normal ridging occurrences at lead times of Weeks 3–4 and Weeks 5–6,with some sensitivity to the specific ridge location and method for determining skill. Using the European Centre for Medium‐Range Weather Forecasts (ECMWF) model as a case study, longer lead time forecast busts are shown to often occur under extended periods of highly active La Nina‐like tropical convection. Model errors at simulating these tropical convection features may have a disproportionately large impact and degrade downstream forecasts over the western United States. Despite the documented model shortcomings, our results highlight an opportunity to target skillful ridging forecasts from dynamical models for improving early warning drought forecasting at S2S lead times.

S2S↗

Benchmarking soil moisture and its relationship to ecohydrologic variables in Earth System Models

Soil moisture (SM) is a key regulator of ecosystem biogeophysics, influencing plant water relations and land-atmosphere energy exchanges. We evaluate the representation of SM in 16 Earth System Models from the Coupled Model Intercomparison Project Phase 6 (CMIP6) using the International Land Model Benchmarking (ILAMB) framework, focusing on surface (0–5, 0–10 cm) and rootzone (0–100 cm) depths, as well as key ecohydrological variables like gross primary productivity (GPP), leaf area index (LAI), and evapotranspiration (ET), and their coupling. Models are benchmarked against multiple observational and assimilated datasets to assess both state variables and cross-variable relationships. Surface SM is generally well represented (r > 0.87), while rootzone SM variability is systematically overestimated (normalized standard deviation > 1). ET shows strong agreement with observations (r > 0.9), whereas GPP and LAI exhibit larger inter-model spread. Skill in individual variables does not guarantee realistic SM–ecohydrology coupling, which varies strongly across models and depends on the reference dataset. Köppen-based regional analyses reveal strong regime dependence, with several models performing well in Tropical and Temperate regions but degrading in Continental (high-latitude) zones. Across both global and regional benchmarks, models cluster by land surface framework, indicating that structural choices in soil hydrology and soil–plant coupling exert a first-order control on performance. These results provide process-relevant benchmarks and suggest that improving the representation of vertical soil structure, rooting depth distributions, and soil–plant hydraulic coupling will be central to advancing soil moisture realism in next-generation Earth system models.

CMIP6↗

Modeling the smoky troposphere of the southeast Atlantic: a comparison to ORACLES airborne observations from September of 2016

In the southeast Atlantic, well-defined smoke plumes from Africa advect over marine boundary layer cloud decks; both are most extensive around September, when most of the smoke resides in the free troposphere. A framework is put forth for evaluating the performance of a range of global and regional atmospheric composition models against observations made during the NASA ORACLES (ObseRvations of Aerosols above CLouds and their intEractionS) airborne mission in September 2016. A strength of the comparison is a focus on the spatial distribution of a wider range of aerosol composition and optical properties than has been done previously. The sparse airborne observations are aggregated into approximately 2° grid boxes and into three vertical layers: 3–6 km, the layer from cloud top to 3 km, and the cloud-topped marine boundary layer. Simulated aerosol extensive properties suggest that the flight-day observations are reasonably representative of the regional monthly average, with systematic deviations of 30 % or less. Evaluation against observations indicates that all models have strengths and weaknesses, and there is no single model that is superior to all the others in all metrics evaluated. Whereas all six models typically place the top of the smoke layer within 0–500 m of the airborne lidar observations, the models tend to place the smoke layer bottom 300–1400 m lower than the observations. A spatial pattern emerges, in which most models underestimate the mean of most smoke quantities (black carbon, extinction, carbon monoxide) on the diagonal corridor between 16° S, 6° E, and 10° S, 0° E, in the 3–6 km layer, and overestimate them further south, closer to the coast, where less aerosol is present. Model representations of the above-cloud aerosol optical depth differ more widely. Most models overestimate the organic aerosol mass concentrations relative to those of black carbon, and with less skill, indicating model uncertainties in secondary organic aerosol processes. Regional-mean free-tropospheric model ambient single scattering albedos vary widely, between 0.83 and 0.93 compared with in situ dry measurements centered at 0.86, despite minimal impact of humidification on particulate scattering. The modeled ratios of the particulate extinction to the sum of the black carbon and organic aerosol mass concentrations (a mass extinction efficiency proxy) are typically too low and vary too little spatially, with significant inter-model differences. Most models overestimate the carbonaceous mass within the offshore boundary layer. Overall, the diversity in the model biases suggests that different model processes are responsible. The wide range of model optical properties requires further scrutiny because of their importance for radiative effect estimates.

Yohei Shinozuka↗

Challenges of COVID-19 Case Forecasting in the US, 2020–2021

During the COVID-19 pandemic, forecasting COVID-19 trends to support planning and response was a priority for scientists and decision makers alike. In the United States, COVID-19 forecasting was coordinated by a large group of universities, companies, and government entities led by the Centers for Disease Control and Prevention and the US COVID-19 Forecast Hub ( https://covid19forecasthub.org ). We evaluated approximately 9.7 million forecasts of weekly state-level COVID-19 cases for predictions 1–4 weeks into the future submitted by 24 teams from August 2020 to December 2021. We assessed coverage of central prediction intervals and weighted interval scores (WIS), adjusting for missing forecasts relative to a baseline forecast, and used a Gaussian generalized estimating equation (GEE) model to evaluate differences in skill across epidemic phases that were defined by the effective reproduction number. Overall, we found high variation in skill across individual models, with ensemble-based forecasts outperforming other approaches. Forecast skill relative to the baseline was generally higher for larger jurisdictions (e.g., states compared to counties). Over time, forecasts generally performed worst in periods of rapid changes in reported cases (either in increasing or decreasing epidemic phases) with 95% prediction interval coverage dropping below 50% during the growth phases of the winter 2020, Delta, and Omicron waves. Ideally, case forecasts could serve as a leading indicator of changes in transmission dynamics. However, while most COVID-19 case forecasts outperformed a naïve baseline model, even the most accurate case forecasts were unreliable in key phases. Further research could improve forecasts of leading indicators, like COVID-19 cases, by leveraging additional real-time data, addressing performance across phases, improving the characterization of forecast confidence, and ensuring that forecasts were coherent across spatial scales. In the meantime, it is critical for forecast users to appreciate current limitations and use a broad set of indicators to inform pandemic-related decision making.

59 BASIC BIOLOGICAL SCIENCES↗

Evaluation of High Mountain Asia-Land Data Assimilation System (Version 1) from 2003 to 2016, Part I: A Hyper-Resolution Terrestrial Modeling System

This first paper of the two-part series focuses on demonstrating the accuracy of a hyper-resolution, offline terrestrial modeling system used for the High Mountain Asia (HMA) region. To this end, this study systematically evaluates four sets of model simulations at point scale, basin scale, and domain scale obtained from different spatial resolutions including 0.01° (∼1-km) and 0.25° (∼25-km). The assessment is conducted via comparisons against ground-based observations and satellite-derived reference products. The key variables of interest include surface net shortwave radiation, surface net longwave radiation, skin temperature, near-surface soil temperature, snow depth, snow water equivalent, and total runoff. In the evaluation against ground-based measurements, the superiority of the 0.01° estimates are mostly demonstrated across relatively complex terrain. Specifically, hyper-resolution modeling improves the skill in meteorological forcing estimates (except precipitation) by 9% relative to coarse-resolution estimates. The model forced by downscaled forcings in its entirety yields the highest skill in model output states as well as precipitation, which improves the skill obtained by coarse-resolution estimates by 7%. These findings, on one hand, corroborate the importance of employing the hyper-resolution versus coarse-resolution modeling in areas characterized by complex terrain. On the other hand, by evaluating four sets of model simulations forced with different precipitation products, this study emphasizes the importance of accurate hyper-resolution precipitation products to drive model simulations.

Yuan Xue↗

Climate change-resilient snowpack estimation in the Western United States

Abstract In the 21st century, warmer temperatures and changing atmospheric circulation will likely produce unprecedented changes in Western United States snowfall 1–3 , with impacts on the timing, amount, and spatial patterns of snowpack 4–7 . The ~900 snow pillow stations are indispensable to water resource management by measuring snow-water equivalent (SWE) 8,9 in strategic but fixed locations 10,11 . However, this network may not be impacted by climate change in the same way as the surrounding area 12 and thus fail to accurately represent unmeasured locations; climate change thereby threatens our ability to measure the effects of climate change on snow. In this work, we show that maintaining the current peak SWE estimation skill is nonetheless possible. We find that explicitly including spatial correlations—either from gridded observations or learned by the model—improves skill at predicting distributed snowpack from sparse observations by 184%. Existing artificial intelligence methods can be useful tools to harness the many available sources of snowpack information to estimate snowpack in a nonstationary climate.

54 ENVIRONMENTAL SCIENCES↗