Search NASA⌕ Search

SEARCH · Search NASA

Results for “LONG RANGE WEATHER FORECASTING”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Mesoscale Organization in Cumulus-Coupled Stratocumulus

Marine cloud systems cover a substantial portion of the world’s oceans. Most of these clouds form relatively close to the ocean surface, typically within one to two kilometers, a region referred to by meteorologists as the marine boundary layer. They are composed predominantly of liquid water, although ice particles can occur in mid- and high-latitude marine clouds during winter. In satellite imagery, these clouds appear bright against the darker ocean surface below, reflecting a large fraction of incoming sunlight back into space that would otherwise warm the ocean. Because marine boundary layer clouds cover such an extensive area of the ocean, they exert a significant influence on Earth’s overall transfer of solar energy absorbed by the surface and thermal energy emitted to space, a balance known as the planetary radiation budget. Marine boundary layer clouds are typically thin, and their formation and dissipation depend on a delicate balance between processes acting at the ocean surface below and the warm, dry air above. They are notoriously difficult to simulate accurately in weather forecast models, which often produce too few marine low clouds in midlatitudes and clouds in tropical regions that are excessively bright, meaning they reflect too much solar radiation. The marine boundary layer is frequently characterized by widespread overcast cloud cover that often transitions from a continuous, single-layer deck to more broken cloud fields toward the tropics. These transitions typically proceed through an intermediate stage in which shallow, broken clouds form beneath the overlying stratiform cloud deck. Once broken clouds develop below the overcast, they frequently self-organize into cloud clusters known as marine boundary layer convective complexes (MBLCCs), although the mechanisms governing the formation and organization of MBLCCs remain poorly understood. Accurately representing these transitions in long-range weather forecast models is essential because they influence the properties of air masses advected over the continental United States and Europe, and they become increasingly important for forecasts on seasonal and longer timescales. We employed two complementary approaches to investigate the processes controlling MBLCCs and their impact on marine cloud cover. Long-term observations from the U.S. Department of Energy’s Eastern North Atlantic (ENA) Observatory provided a unique dataset that allowed us to characterize fundamental properties of MBLCCs, including their typical size and frequency of occurrence. These observations were combined with high-resolution numerical simulations performed on supercomputers to examine the evolution of MBLCCs during cold-air outbreaks over the ENA region.

54 ENVIRONMENTAL SCIENCES↗

Evaluating the potential of short-term instrument deployment to improve distributed wind resource assessment

Distributed wind projects, which are connected at the distribution level of an electricity system or in off-grid applications to serve specific or local energy needs, often rely solely on wind resource models to establish wind speed and energy generation expectations. Historically, anemometer loan programs have provided an affordable avenue for more accurate onsite wind resource assessment, and the lowering cost of lidar systems has shown similar advantages for more recent assessments. While a full 12 months of onsite wind measurement is the standard for correcting model-based long-term wind speed estimates for utility-scale wind farms, the time and capital investment involved in gathering onsite measurements must be reconciled with the energy needs and funding opportunities that drive expedient deployment of distributed wind projects. Much literature exists to quantify the performance of correcting long-term wind speed estimates with 1 or more years of observational data, but few studies explore the impacts of correcting with months-long observational periods. This study aims to answer the question of how short you can go in terms of the observational time period needed to make impactful improvements to model-based long-term wind speed estimates. Three algorithms, multivariable linear regression, adaptive regression splines, and regression trees, are evaluated for their skill at correcting long-term wind resource estimates from the European Centre for Medium-Range Weather Forecasts Reanalysis version 5 (ERA5) using months-long periods of observational data from 66 locations across the US. On average, correction with even 1 month of observations provides significant improvement over the baseline ERA5 wind speed estimates and produces median bias magnitudes and relative errors within 0.22 m s −1 and 4 percentage points of the median bias magnitudes and relative errors achieved using the standard 12 months of data for correction. However, in cases when the shortest observational periods (1 to 2 months) used for correction are not well correlated with the overlapping ERA5 reference, the resultant long-term wind speed errors are worse than those produced using ERA5 without correction. Summer months, which are characterized by weaker relative wind speeds and standard deviations for most of the evaluation sites, tend to produce the worst results for long-term correction using months-long observations. The three tested algorithms perform similarly for long-term wind speed bias; however, regression trees perform notably worse than multivariable linear regression and adaptive regression splines in terms of correlation when using 6 months or less of observational data for correction. Translating the analysis to wind energy, median relative errors in the capacity factor are on average within 10 % using 1 month of training. If the observation period used for correction is not well correlated with the reference data, however, misrepresentation of the observed capacity factor can be substantial. The risk associated with poor correlation between the observed and reference datasets decreases with increasing training period length. In the worst-correlation scenarios, the median capacity factor relative errors from using 1, 3, and 6 months are within 47 %, 26 %, and 16 %, respectively.

17 WIND ENERGY↗

Data Assimilation with Machine Learning Surrogate Models: A Case Study with FourCastNet

Modern data-driven surrogate models for weather forecasting provide accurate short-term predictions but inaccurate and nonphysical long-term forecasts. This paper investigates online weather prediction using machine learning surrogates supplemented with partial and noisy observations. We empirically demonstrate and theoretically justify that, despite the long-time instability of the surrogates and the sparsity of the observations, filtering estimates can remain accurate in the long-time horizon. As a case study, we integrate the Fourier Forecasting Neural Network (FourCastNet), a weather surrogate model, within a variational data assimilation framework using partial, noisy ERA5 global reanalysis data from the European Centre for Medium-Range Weather Forecasts (ECMWF). Here, our results show that filtering estimates remain accurate over a year-long assimilation window and provide effective initial conditions for forecasting tasks, including extreme event prediction.

Data assimilation↗

Numerical case study of the aerosol–cloud interactions in warm boundary layer clouds over the eastern North Atlantic with an interactive chemistry module

The presence of warm boundary layer stratiform clouds over the eastern North Atlantic (ENA) region is commonly influenced by the Azores High, especially during the summer season. To investigate comprehensive aerosol–cloud interactions, this study employs the Weather Research and Forecasting model coupled with a chemistry component (WRF-Chem), incorporating aerosol chemical components that are relevant to the formation of cloud condensation nuclei (CCN) and accounting for aerosol spatiotemporal variation. This study focuses on aerosol indirect effects, particularly the long-range transport of aerosols, in the ENA region under three different weather regimes: a ridge with a surface high-pressure system, a post-trough with a surface high-pressure system, and a weak trough. The WRF-Chem simulations conducted at a near-large-eddy scale offer valuable insights into the model's performance, especially in terms of its ability to use high spatial resolution to capture mesoscale cloud features across various weather regimes. Our result shows that introducing 5 times more aerosols to either non-precipitating or precipitating clouds significantly increases ambient CCN numbers, resulting in, to varying degrees, higher liquid water path (LWP) values. The substantial aerosol–cloud interaction especially occurs in the precipitating clouds and demonstrates the susceptibility of the LWP to changes in CCN under different regimes. Conversely, thin, non-rain clouds at the edges of a cloud system are prone to evaporation, exhibiting an aerosol drying effect. The aerosols released during this process transition back to the accumulation mode, facilitating future activation. This dynamic behavior is not adequately represented in prescribed-aerosol simulations.

54 ENVIRONMENTAL SCIENCES↗

WTK-LED: The WIND Toolkit Long-Term Ensemble Dataset

To satisfy a wide group of stakeholders across various wind energy disciplines, including but not limited to stakeholders in the distributed and utility scale wind industry, the new emerging airborne wind energy field, grid integration, power systems modeling, environmental modeling, and researchers in academia, and to close some of the gaps that current public datasets have, we aimed at developing an updated version of the meteorological WIND Toolkit, named WIND Toolkit Long-term Ensemble Dataset (WTK-LED), which is a meteorological dataset providing time series every 5 min and 2 km, including model uncertainty of wind speed at every modeling grid point so that users are provided with a range of possible wind speeds every 2 km. The data were produced using the Weather Research and Forecasting Model (WRF). The vertical grid used in WTK-LED includes many vertical layers in the atmospheric boundary layer to provide information of atmospheric quantities across the rotor layer of utility scale and distributed wind turbines. The WTK-LED includes: 1) Numerical simulations covering the continental United States, Alaska, and Hawaii, with high-resolution data being available for 3 years (2018-2020). 2) Climate simulations from Argonne National Laboratories covering the North American continent, including Alaska, Canada, and most of Mexico and the Caribbean Islands. These simulations complement the new WTK-LED to offer a 4-km dataset covering 20 years, from 2001-2020. 3) Specific long-term,high-resolution offshore simulations have been conducted separately for the US coasts, Hawaii, and the Great Lakes, leading to the 2023 National Offshore Wind data set. This report focuses on a description of the land-based WTK-LED for CONUS, Hawaii, and Alaska, for the 3-year 2-km/5-min dataset and the 20-year 4-km/hourly dataset, as well as the uncertainty quantification method. We also provide limited validation results. Based on our results to date, we suggest use cases and applications for each dataset of the WTK-LED.

17 WIND ENERGY↗

Evaluation of a high-resolution regional climate simulation for surface and hub-height wind climatology over North America

Assessing the availability of key wind resources requires augmenting observations to support the implementation of wind energy infrastructure. However, observations are limited, necessitating the development of high-resolution, long-term gridded datasets. This study presents a robust, dynamically downscaled climatological dataset, offering 20 years of hourly wind data at a 4 km spatial resolution across North America, and evaluates its performance against observations, including meteorological towers and automated surface-observing system (ASOS) stations, as well as coarse-resolution reanalysis data (the European Centre for Medium-Range Weather Forecasts (ECMWF) reanalysis version 5 (ERA5)). Results demonstrate that the downscaled high-resolution wind data outperform ERA5 in regions of complex terrain and coastal areas, with improved overlap coefficients for wind data distributions and reduced root mean square errors (RMSEs) for hub-height and near-surface diurnal wind patterns. The downscaled simulation also captures the synoptic drivers of seasonal wind direction patterns reasonably well, indicated by high wind rose similarity indices. This study also provides an analysis of interannual variability, utilizing the dataset's full 20-year period, and model uncertainty, generated by varying model initial conditions and physics parameterizations across 1-year ensemble members, which are key considerations for wind resource assessment in wind farm development.

17 WIND ENERGY↗

Forecast of Wildfire Potential Across California USA Using a Transformer

Wildfires are a major issue facing the United States, a matter further exacerbated by an ever-changing climate. In California alone, wildfires are responsible for billions of dollars in damages and take lives each year. Accurately predicting fire danger conditions allows preparation awareness before wildfires start. Transformers are a class of deep learning models designed to identify patterns in sequential datasets. In recent years, transformers have gained popularity through their impressive performance in natural language processing and other applications of signal recognition. This analysis demonstrates the ability of a transformer with a residual connection to forecast fire danger potential over the state of California. Wildland fire potential index (WFPI) maps collected from the US Geological Survey database from January 1st 2020 to December 31st 2023 were used to tune, train and evaluate the transformer. Meteorological inputs (provided by Daymet daily weather and climatological summaries), the normalized difference vegetation index (NDVI) (calculated from the Moderate Resolution Imaging Spectroradiometer (MODIS)), and outputs from the Scott and Burgman fire behavior fuel models (to characterize maps of fuel types), were used as inputs. Our results show that a transformer can effectively emulate the US Forest Service modeled WFPI maps of California USA for four week long forecasts over the month of July, 2023, with correlations ranging from 0.85 – 0.98.

Limber, Russell [ORNL]↗

Differences in cluster and internal wake effects from mesoscale and large-eddy simulations off the US East Coast

Mesoscale simulations are increasingly used to estimate wake effects within and between large wind farms, despite limited validation for large-scale wake effects. This study evaluates the capabilities and limitations of mesoscale simulations in capturing wake-induced impacts on wind turbine power production through a direct comparison with large-domain large-eddy simulations (LESs) for three planned offshore wind farms under realistic atmospheric conditions and a range of atmospheric stabilities. We assess mesoscale performance in replicating wake characteristics behind single and multiple turbine clusters and quantify the resulting variability in mean turbine power. Results show that mesoscale Weather Research and Forecasting simulations with the Fitch wind farm parameterization capture key features of the velocity deficit downstream of both single and multiple wind farms, with mean root-mean-square errors near 5 % and good agreement with stability-driven wake behavior. However, in these simulations, the mesoscale Fitch parameterization underestimates power losses from internal wake effects, particularly when turbines align with the prevailing wind direction or under stable stratification. In these conditions, individual wakes persist and dominate downstream power deficits. The coarse resolution of the mesoscale simulations limits their ability to resolve individual wind turbine wakes that drive power fluctuations within wind farms. Nonetheless, mesoscale simulations can yield accurate estimates of combined wake losses from internal and cluster effects across some wind direction sectors, where errors in wake representation may cancel each other out. These findings underscore the strengths of mesoscale simulations for capturing broader wake patterns while highlighting their limitations for modeling turbine-level power losses. Future work should explore hybrid modeling approaches to capture both long-range cluster wake propagation and localized internal wake dynamics.

17 WIND ENERGY↗

Multi-Scale Integrated Monitoring System for Enhancing Methane Emission Detection, Quantification & Prediction

This report details the progress and findings of a comprehensive study on reviewing existing solutions, identifying technology gaps, and formulating an “all-in-one” integrated strategy for developing the next-generation multiscale methane monitoring and modeling platform, conducted under grant number DE-FE0032292. Co-led by Dr. David Ebert, Dr. Binbin Weng, and Dr. Chenghao Wang at the University of Oklahoma, the project’s goal was to develop an integrated approach for building this engineering platform to detect, quantify, and mitigate methane emissions across various temporal scale, spatial scales, and sectors. The planning grant study began with an extensive review of various methane sensing and monitoring technologies and systems, surveying over 100 technology providers globally. This review revealed the prevalence of optical methods over chemical methods in commercially available sensors, with Non-Dispersive Infrared (NDIR), Tunable Diode Laser Absorption Spectroscopy (TDLAS), and Optical Gas Imaging (OGI) cameras being the most prevalent options. A trend towards more advanced optical techniques was observed, driven by increased regulatory focus and technological advancements. The technical evaluation of these sensing technologies provided crucial insights into their capabilities and limitations. The study examined emerging technologies such as Differential Absorption LiDAR (DIAL), which show promise for high-precision and long-range detection. The team then investigated the features and application bandwidth of various sensing platforms, including handheld, fixed/stationary, mobile, aerials, and spaceborne monitors. Pilot field studies were conducted to assess the capabilities of solutions for different emission scenarios. Field work with sensor deployments was conducted at three distinct site types: an oil & gas industry site, a cattle ranching operation, and a waste processing facility. The team also conducted a thorough review of methane flux inverse modeling approaches, focused on physically based methods. These approaches were categorized into simple, intermediate, and advanced methods. A realtime WRF-GHG (Weather Research and Forecasting-Greenhouse Gas) modeling system was developed and applied, incorporating multiple data sources to guide field experiments and inform methane plume detection. The project identified and analyzed numerous categories of methane data sources, including satellite measurements, ground-based sensors, and inventory databases. Key platforms examined include EDGAR, EPA GHGI, NASA TROPOMI, Carbon Mapper, and Climate TRACE, among others. The team proposed an architecture for a comprehensive methane monitoring platform. This system incorporates multi-source data acquisition, advanced data processing and assimilation, interactive visualization tools, and analytical capabilities for emissions forecasting and scenario analysis. The proposed platform aims to provide a user-friendly interface catering to various stakeholders, from researchers to policymakers. The architecture includes sophisticated data ingestion methods, a centralized data warehouse, and advanced analytical tools for data fusion and interpretation. To ensure the relevance and effectiveness of the proposed system, a comprehensive survey was conducted to gather stakeholder input on system requirements. Key findings include a strong need for integrating various data types and formats, a preference for real-time data updates and advanced visualization tools, and a demand for user-friendly interfaces catering to different expertise levels.

03 NATURAL GAS↗

Performance of reanalysis and mesoscale models off the coast of Hawai'i

The eastern Hawai'i coast in the United States is characterized by considerable wind resource fuelled by persistent trade winds, making it an important area for energy research. The need is strong for reanalyses and higher-resolution regional simulations where observations have been historically limited, such as Hawai'i's offshore environments. However, studies using offshore observations in other parts of the world have shown that significant errors can occur in reanalyses and wind datasets, which can lead to inaccurate estimates of wind energy generation, payback periods, and extreme weather risks at project locations. The degree of such errors is influenced by a number of factors, including spatial resolution and the handling of processes within the planetary boundary layer (PBL). In this work, we provide a wind resource characterization from year-long lidar buoy measurements off the eastern coast of O'ahu, Hawai'i, an environment previously unobserved at the rotor level, and use the characterization to evaluate the performance of two simulation datasets. The O'ahu deployment location is meteorologically unique and less complex than land-based wind resource characterizations, being strongly characterized by trade winds with minimal land–atmosphere interaction influences. Despite the unique and fairly consistent meteorological conditions, we hypothesize that distinct simulation datasets will exhibit diverse ranges of errors similar to those that have been seen for other offshore locations. We find the European Centre for Medium-Range Weather Forecasts (ECMWF) Reanalysis version 5 (ERA5) to strongly underestimate observed wind speeds at the O'ahu location (bias = −1.54 m s −1 at a height of 140 m above sea level), while a regional Weather Research and Forecasting Model (WRF) simulation produced by the University of Hawai'i (UH-WRF) provides a significantly smaller wind speed bias (−0.25 m s −1 ), highlighting the value of running regional, higher-resolution simulations. The large bias noted for ERA5 is driven by significant underestimation of fast wind speeds (>9 m s −1 ), which the study site is largely characterized by, along with discontinuities in the ERA5 diurnal cycle. We also speculate that the relative sparsity of observations for data assimilation in this remote part of the world could influence the performance of ERA5 and that challenges with characterizing island effects could impact the performance of both datasets.

17 WIND ENERGY↗

Huge ensembles – Part 1: Design of ensemble weather forecasts using spherical Fourier neural operators

Abstract. Simulating low-likelihood high-impact extreme weather events in a warming world is a significant and challenging task for current ensemble forecasting systems. While these systems presently use up to 100 members, larger ensembles could enrich the sampling of internal variability. They may capture the long tails associated with climate hazards better than traditional ensemble sizes. Due to computational constraints, it is infeasible to generate huge ensembles (comprised of 1000–10 000 members) with traditional, physics-based numerical models. In this two-part paper, we replace traditional numerical simulations with machine learning (ML) to generate hindcasts of huge ensembles. In Part 1, we construct an ensemble weather forecasting system based on spherical Fourier neural operators (SFNOs), and we discuss important design decisions for constructing such an ensemble. The ensemble represents model uncertainty through perturbed-parameter techniques, and it represents initial condition uncertainty through bred vectors, which sample the fastest-growing modes of the forecast. Using the European Centre for Medium-Range Weather Forecasts Integrated Forecasting System (IFS) as a baseline, we develop an evaluation pipeline composed of mean, spectral, and extreme diagnostics. With large-scale, distributed SFNOs with 1.1 billion learned parameters, we achieve calibrated probabilistic forecasts. As the trajectories of the individual members diverge, the ML ensemble mean spectra degrade with lead time, consistent with physical expectations. However, the individual ensemble members' spectra stay constant with lead time. Therefore, these members simulate realistic weather states during the rollout, and the ML ensemble passes a crucial spectral test in the literature. The IFS and ML ensembles have similar extreme forecast indices, and we show that the ML extreme weather forecasts are reliable and discriminating. These diagnostics ensure that the ensemble can reliably simulate the time evolution of the atmosphere, including low-likelihood high-impact extremes. In Part 2, we generate a huge ensemble initialized each day in summer 2023, and we characterize the simulations of extremes.

Mahesh, Ankur↗

Advancing ocean monitoring and knowledge for societal benefit: the urgency to expand Argo to OneArgo by 2030

The ocean plays an essential role in regulating Earth’s climate, influencing weather conditions, providing sustenance for large populations, moderating anthropogenic climate change, encompassing massive biodiversity, and sustaining the global economy. Human activities are changing the oceans, stressing ocean health, threatening the critical services the ocean provides to society, with significant consequences for human well-being and safety, and economic prosperity. Effective and sustainable monitoring of the physical, biogeochemical state and ecosystem structure of the ocean, to enable climate adaptation, carbon management and sustainable marine resource management is urgently needed. The Argo program, a cornerstone of the Global Ocean Observing System (GOOS), has revolutionized ocean observation by providing real-time, freely accessible global temperature and salinity data of the upper 2,000m of the ocean (Core Argo) using cost-effective simple robotics. For the past 25 years, Argo data have underpinned many ocean, climate and weather forecasting services, playing a fundamental role in safeguarding goods and lives. Argo data have enabled clearer assessments of ocean warming, sea level change and underlying driving processes, as well as scientific breakthroughs while supporting public awareness and education. Building on Argo’s success, OneArgo aims to greatly expand Argo’s capabilities by 2030, expanding to full-ocean depth, collecting biogeochemical parameters, and observing the rapidly changing polar regions. Providing a synergistic subsurface and global extension to several key space-based Earth Observation missions and GOOS components, OneArgo will enable biogeochemical and ecosystem forecasting and new long-term climate predictions for which the deep ocean is a key component. Driving forward a revolution in our understanding of marine ecosystems and the poorly-measured polar and deep oceans, OneArgo will be instrumental to assess sea level change, ocean carbon fluxes, acidification and deoxygenation. Emerging OneArgo applications include new views of ocean mixing, ocean bathymetry and sediment transport, and ecosystem resilience assessment. Implementing OneArgo requires about $100 million annually, a significant increase compared to present Argo funding. OneArgo is a strategic and cost-effective investment which will provide decision-makers, in both government and industry, with the critical knowledge needed to navigate the present and future environmental challenges, and safeguard both the ocean and human wellbeing for generations to come.

ARGO↗

High-Resolution South American Wind Resource Data Downscaled with Generative Machine Learning Conditioned on Near-Surface Observations

High-resolution historical wind data was developed for the entirety of South America using the innovative Super-Resolution for Renewable Resource Data (sup3r) machine learning framework. The publicly available Sup3rWind South America dataset represents a significant advancement in wind resource data generation, leveraging generative machine learning conditioned on near-surface observations from the Meteorological Assimilation Data Ingest System (MADIS) to efficiently and accurately downscale coarse reanalysis data from the European Centre for Medium-Range Weather Forecasts (ERA5). This approach produces fine-scale, spatially and temporally coherent wind and meteorological fields hundreds of times more computationally efficient than traditional numerical weather modeling methods, enabling access to high-fidelity wind information across both continental and offshore regions. Sup3rWind South America builds on the earlier Sup3rWind Ukraine dataset through improvements in model architecture and outputs conditioned on near-surface observation inputs. As with the Ukraine data release, this dataset includes wind speed, wind direction, temperature, relative humidity, and pressure at a horizontal resolution of ~2 km, representing a 15x spatial enhancement relative to the 31 km ERA5 grid. Wind speed and direction are provided at 5-minute resolution, a 12x temporal refinement compared to the hourly ERA5 data, while temperature, relative humidity, and pressure remain at hourly resolution. The data covers all years from 2005 to 2024. Before downscaling, ERA5 inputs were bias-corrected using long-term monthly means and a limited number of quality-controlled observations to align large-scale statistics with regional conditions. The resulting dataset is the first publicly available high-resolution timeseries wind record that provides full spatial coverage of South America. Model validation demonstrates strong agreement with observations across several statistical metrics, consistent with other state-of-the-art high-resolution wind resource datasets. The potential applications of Sup3rWind South America span renewable energy resource assessment, energy system modeling, and grid resilience analysis. The 20-year record and high spatial and temporal resolution support accurate estimation of long-term energy yield and the economic feasibility of potential wind development sites. Continuous coverage across both continental and offshore regions enables comprehensive site prospecting within exclusive economic zones. The 2 km, 5-minute resolution data provide the spatial and temporal variability required for power system simulation, operational planning, and regional risk assessments.

17 WIND ENERGY↗

Estimating soybean yields from high-temporal-resolution multi-source data using deep learning

Accurate and timely crop yield prediction is crucial for ensuring food security and maintaining stable agricultural markets. In recent years, there has been a surge in interest in leveraging high-temporal-resolution, multi-source data for effective crop growth monitoring and yield estimation. A notable challenge arises from the difficulty in capturing the intricate interactions between variables across different time steps within these high-temporal-resolution time series datasets. This complexity hinders the reliable extraction of yield information from voluminous and often noisy datasets, especially during periods of extreme weather events. Here, in this study, we propose an Attention and Graph Isomorphism Network-enhanced Bi-directional Long Short-Term Memory network (AGB-LSTM) for estimating county-level soybean yield in the United States. This model integrates a diverse set of remote sensing data, including Near-Infrared Reflectance of Vegetation (NIRv), Sun-Induced chlorophyll Fluorescence (SIF), and Gross Primary Productivity (GPP), along with environmental covariates. The AGB-LSTM effectively leverages information related to crop yield from high-temporal-resolution time series data (5-days), achieving an accuracy of R²= 0.67 and rRMSE = 14.46%. This approach significantly outperforms traditional machine learning methods such as Random Forest (RF) (R²= 0.52, rRMSE = 17.36%) and Bi-LSTM (R²= 0.58, rRMSE = 16.17%). Sensitivity experiments with different time steps and ranges demonstrated that our model could accurately and stably predict yields 1 to 2 months before harvest. Moreover, data with a finer temporal resolution consistently improved prediction performance, resulting in an approximately 20% increase in and an approximately 20% decrease in rRMSE compared to using monthly composites. We also evaluated the robustness of the model under extreme climate events and observed strong performance (R²= 0.50, rRMSE = 21.32%). Finally, yield mapping for major soybean-producing regions in North America in 2023 revealed spatial patterns that closely matched USDA yield reports. Our findings suggest that the AGB-LSTM model is a promising and effective method for estimating yield and has notable potential for global crop yield forecasting.

Deep learning↗

Towards provision of regularly updated climate data from the Coupled Model Intercomparison Project

The Coupled Model Intercomparison Project (CMIP) is a flagship of the World Climate Research Programme (WCRP). CMIP has become a recognised ‘brand’ in climate circles evolving over the last thirty years from a targeted research activity by a small number of climate modelling centres intercomparing their Earth System Model (ESM) simulations to a broad international coordinated research effort (Durack et al, 2025). CMIP is organized as a research activity leveraging funded and in-kind contributions from experts within modelling centres and the broader scientific community supported more recently by a fully-funded International Project Office. Within CMIP, Model Intercomparison Projects (MIPs) are community-designed to understand past, present and future climate. CMIP data provides a valuable resource for climate research and is routinely used to assess model representation of climate processes and test scientific hypotheses in the context of model uncertainty and (forced and internal) variability as evident from its prolific use in scientific publications1 . The impact relies on enabling infrastructure (most prominently via the Earth System Grid Federation (ESGF)), which allows sharing of simulation output, provision of the boundary conditions used in each simulation, and definition of the data standards that are essential to facilitating wide use of the data. The impact is supplemented by the wide-ranging scrutiny to which model simulations are subjected. Beyond its use in research, CMIP data is a key resource for communities producing derived climate information from downscaling and impact studies, such as the Coordinated Regional Downscaling Experiment (CORDEX; Gutowski et al., 2016) and the Intersectoral Impacts MIP (ISIMIP; Frieler et al., 2024). Government, academic and commercial entities also increasingly rely on CMIP and its downstream data for climate risk assessments and climate services (for example, Copernicus Climate Change Service and World Bank portal). This means that, although CMIP is a research activity, it increasingly serves a secondary and very relevant role as a provider of climate data – a long-recognised dichotomy (Stevens, 2024). Research and applications have distinct needs, with the former requiring flexibility and generality and the latter consistency. Here we explain how the design of the research activity has been adapted to reduce the burdens imposed by applications and how the research infrastructure might evolve to further enable scientific inquiry. We propose one possible approach to consistently providing model information and projections for applications in the future.

Environmental sciences↗