Crop yield prediction based on reanalysis and crop phenology data in the agroclimatic zones
Not provided.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not provided.
Explore the source record for details and available documents.
Accurate and timely crop yield prediction is crucial for ensuring food security and maintaining stable agricultural markets. In recent years, there has been a surge in interest in leveraging high-temporal-resolution, multi-source data for effective crop growth monitoring and yield estimation. A notable challenge arises from the difficulty in capturing the intricate interactions between variables across different time steps within these high-temporal-resolution time series datasets. This complexity hinders the reliable extraction of yield information from voluminous and often noisy datasets, especially during periods of extreme weather events. Here, in this study, we propose an Attention and Graph Isomorphism Network-enhanced Bi-directional Long Short-Term Memory network (AGB-LSTM) for estimating county-level soybean yield in the United States. This model integrates a diverse set of remote sensing data, including Near-Infrared Reflectance of Vegetation (NIRv), Sun-Induced chlorophyll Fluorescence (SIF), and Gross Primary Productivity (GPP), along with environmental covariates. The AGB-LSTM effectively leverages information related to crop yield from high-temporal-resolution time series data (5-days), achieving an accuracy of R²= 0.67 and rRMSE = 14.46%. This approach significantly outperforms traditional machine learning methods such as Random Forest (RF) (R²= 0.52, rRMSE = 17.36%) and Bi-LSTM (R²= 0.58, rRMSE = 16.17%). Sensitivity experiments with different time steps and ranges demonstrated that our model could accurately and stably predict yields 1 to 2 months before harvest. Moreover, data with a finer temporal resolution consistently improved prediction performance, resulting in an approximately 20% increase in and an approximately 20% decrease in rRMSE compared to using monthly composites. We also evaluated the robustness of the model under extreme climate events and observed strong performance (R²= 0.50, rRMSE = 21.32%). Finally, yield mapping for major soybean-producing regions in North America in 2023 revealed spatial patterns that closely matched USDA yield reports. Our findings suggest that the AGB-LSTM model is a promising and effective method for estimating yield and has notable potential for global crop yield forecasting.
Abstract. Irrigated cultivation exerts a significant influence on the local climate and the hydrological cycle. The North China Plain (NCP) is known for its intricate agricultural system, marked by expansive cropland, high productivity, compact rotation, a semi-arid climate, and intensive irrigation practices. As a result, there has been considerable attention on the potential impact of this intensive irrigated agriculture on the local climate. However, studying the irrigation impact in this region has been challenging due to the lack of an accurate simulation of crop phenology and irrigation practices within the climate model. By incorporating double cropping with interactive irrigation, our study extends the capabilities of the Weather Research Forecast (WRF) model, which has previously demonstrated commendable performance in simulating single-cropping scenarios. This allows for two-way feedback between irrigated crops and climate, further enabling the inclusion of irrigation feedback from both ground and vegetation perspectives. The improved crop modeling system shows significant enhancement in capturing vegetation and irrigation patterns, which is evidenced by its ability to identify crop stages, estimate field biomass, predict crop yield, and project monthly leaf area index. The improved simulation of large-scale irrigated crops in the NCP can further enhance our understanding of the intricate relationship between agricultural development and climate change.
In both plant breeding and crop management, interpretability plays a crucial role in instilling trust in AI-driven approaches and enabling the provision of actionable insights. The primary objective of this research is to explore and evaluate the potential contributions of deep learning network architectures that employ stacked LSTM for end-of-season maize grain yield prediction. A secondary aim is to expand the capabilities of these networks by adapting them to better accommodate and leverage the multi-modality properties of remote sensing data. In this study, a multi-modal deep learning architecture that assimilates inputs from heterogeneous data streams, including high-resolution hyperspectral imagery, LiDAR point clouds, and environmental data, is proposed to forecast maize crop yields. The architecture includes attention mechanisms that assign varying levels of importance to different modalities and temporal features that, reflect the dynamics of plant growth and environmental interactions. The interpretability of the attention weights is investigated in multi-modal networks that seek to both improve predictions and attribute crop yield outcomes to genetic and environmental variables. This approach also contributes to increased interpretability of the model's predictions. The temporal attention weight distributions highlighted relevant factors and critical growth stages that contribute to the predictions. The results of this study affirm that the attention weights are consistent with recognized biological growth stages, thereby substantiating the network's capability to learn biologically interpretable features. Accuracies of the model's predictions of yield ranged from 0.82-0.93 R 2 ref in this genetics-focused study, further highlighting the potential of attention-based models. Further, this research facilitates understanding of how multi-modality remote sensing aligns with the physiological stages of maize. The proposed architecture shows promise in improving predictions and offering interpretable insights into the factors affecting maize crop yields, while demonstrating the impact of data collection by different modalities through the growing season. By identifying relevant factors and critical growth stages, the model's attention weights provide valuable information that can be used in both plant breeding and crop management. The consistency of attention weights with biological growth stages reinforces the potential of deep learning networks in agricultural applications, particularly in leveraging remote sensing data for yield prediction. To the best of our knowledge, this is the first study that investigates the use of hyperspectral and LiDAR UAV time series data for explaining/interpreting plant growth stages within deep learning networks and forecasting plot-level maize grain yield using late fusion modalities with attention mechanisms.
Switchgrass is a potential crop for bioenergy or carbon capture schemes, but further yield improvements through selective breeding are needed to encourage commercialization. To identify promising switchgrass germplasm for future breeding efforts, we conducted multisite and multitrait genomic prediction with a diversity panel of 630 genotypes from 4 switchgrass subpopulations (Gulf, Midwest, Coastal, and Texas), which were measured for spaced plant biomass yield across 10 sites. Our study focused on the use of genomic prediction to share information among traits and environments. Specifically, we evaluated the predictive ability of cross-validation (CV) schemes using only genetic data and the training set (cross-validation 1: CV1), a subset of the sites (cross-validation 2: CV2), and/or with 2 yield surrogates (flowering time and fall plant height). We found that genotype-by-environment interactions were largely due to the north–south distribution of sites. The genetic correlations between the yield surrogates and the biomass yield were generally positive (mean height r = 0.85; mean flowering time r = 0.45) and did not vary due to subpopulation or growing region (North, Middle, or South). Genomic prediction models had CV predictive abilities of –0.02 for individuals using only genetic data (CV1), but 0.55, 0.69, 0.76, 0.81, and 0.84 for individuals with biomass performance data from 1, 2, 3, 4, and 5 sites included in the training data (CV2), respectively. To simulate a resource-limited breeding program, we determined the predictive ability of models provided with the following: 1 site observation of flowering time (0.39); 1 site observation of flowering time and fall height (0.51); 1 site observation of fall height (0.52); 1 site observation of biomass (0.55); and 5 site observations of biomass yield (0.84). The ability to share information at a regional scale is very encouraging, but further research is required to accurately translate spaced plant biomass to commercial-scale sward biomass performance.
Agriculture, a cornerstone of human civilization, faces rising challenges from climate change, resource limitations, and stagnating yields. Precise crop production forecasts are crucial for shaping trade policies, development strategies, and humanitarian initiatives. This study introduces a comprehensive machine learning framework designed to predict crop production. We leverage CMIP5 climate projections under a moderate carbon emission scenario to evaluate the future suitability of agricultural lands and incorporate climatic data, historical agricultural trends, and fertilizer usage to project yield changes. Our integrated approach forecasts significant regional variations in crop production across Southeast Asia by 2028, identifying potential cropland utilization. Specifically, the cropland area in Indonesia, Malaysia, Philippines, and Viet Nam is projected to decline by more than 10% if no action is taken, and there is potential to mitigate that loss. Moreover, rice production is projected to decline by 19% in Viet Nam and 7% in Thailand, while the Philippines may see a 5% increase compared to 2021 levels. Our findings underscore the critical impacts of climate change and human activities on agricultural productivity, offering essential insights for policy-making and fostering international cooperation.
As climate change expands the world’s arid and semiarid regions, sustainable systems that integrate food and energy production are becoming increasingly critical. Agrivoltaics—co-locating crops with photovoltaic (PV) panels—offers a dual land-use strategy that mitigates environmental stress by shading crops, conserving soil moisture, and enhancing PV efficiency. While climate-smart crops like the tepary bean ( Phaseolus acutifolius ) are well adapted to heat and drought, little is known about how these crops and their associated soil microbiomes respond to the unique microclimates created by PV shading. This study evaluated tepary bean performance and plant–microbial interactions under PV-shade vs. no shade across three soil amendment treatments at two experimental sites. We assessed plant traits including germination, phenology, biomass, height, as well as yield and bean morphology, alongside shifts in soil microbial composition and functional potential. Plants grown under PV-shade were generally taller, with extended reproductive periods and higher yields: 42% of shaded plants produced beans compared to only 8% under full sun. Shaded plants also produced rounder, higher-quality beans, whereas non-shaded plants yielded flatter, less developed beans. Microbial community composition was more strongly influenced by amendment and site conditions than by shading alone. Key microbial taxa (e.g., Glomeromycetes, Desulfobacterota ) and predicted functions (e.g., denitrification, nitrogen-respiration, sulfate reduction) were associated with differences in plant performance. Finally, combining agrivoltaic systems with targeted soil amendments can enhance crop yield and soil microbial functionality—offering a promising strategy for sustainable agriculture in arid landscapes.
This dataset contains biomass yield measurements and associated vegetation index data collected from commercial Miscanthus × giganteus fields in eastern Iowa during the 2022–2023 growing seasons. The data support the analyses presented in the article: “Yield From Iowa's First Commercial Miscanthus Fields: Implications of Spatial Variability for Productivity and Sustainability Beyond Research Plots.” We collected 105 ground-truth biomass samples from four mature commercial fields (>4 years old) covering 92.81 ha. Samples were taken from 3 m² quadrats that were hand-harvested in alignment with commercial harvest timing. Stem biomass (excluding leaves) was weighed, moisture-corrected, and converted to dry-matter yield expressed in Mg DM ha⁻¹. Sampling locations were selected to capture spatial variability visible in aerial imagery and were recorded using RTK GPS. Each biomass observation was paired with vegetation indices derived from high-resolution PlanetScope satellite imagery (3 m resolution). Images were acquired throughout the growing season, and indices were calculated to evaluate their ability to predict end-of-season biomass yield. Statistical and machine learning approaches were used to identify key predictors, and a linear regression model based on end-of-July Green Normalized Difference Vegetation Index (GNDVI) was developed and evaluated. This repository includes the data used in that modeling workflow. Management practices, economic data, full imagery time series, and additional methodological details are described in the associated publication and are not included here. The dataset consists of three comma-separated value (CSV) files: 1. Combine_Groundtruth_Yield_VI_22_23.csv This file contains ground-truth biomass yield measurements and associated key vegetation index values collected during the 2022 and 2023 growing seasons. Rows: 105 observations Columns: Year — Year of observation (2022 or 2023) Field — Field location identifier Sample_number — Unique sample identifier GNDVI_End_Jul — Green Normalized Difference Vegetation Index calculated at end of July GNDVI_End_Aug — Green Normalized Difference Vegetation Index calculated at end of August NDRE_End_Aug — Normalized Difference Red Edge index calculated at end of August Biomass_Stem_Yield_MgDM/ha — Measured stem biomass yield (megagrams dry matter per hectare) 2. trainData_GNDVI.csv This file contains the subset of observations used to train the predictive relationship between July GNDVI and biomass yield. Rows: 76 observations Columns: Unnamed: 0 — Row index retained from the original data processing workflow GNDVI_End_Jul — GNDVI at end of July Stem_Yield_MgDM/ha — Observed stem biomass yield (Mg DM ha⁻¹) 3. testData_GNDVI.csv This file contains the test dataset used to evaluate model performance. Rows: 29 observations Columns: Unnamed: 0 — Row index retained from the original data processing workflow GNDVI_End_Jul — GNDVI at end of July Predicted_Yield_MgDM/ha — Model-predicted stem biomass yield (Mg DM ha⁻¹) Observed_Yield_MgDM/ha — Measured stem biomass yield (Mg DM ha⁻¹)
The stability of food supply and prices may depend more on annual changes in yields from year-to-year variability in weather than on longer-term average changes from changing climatic conditions. However, the absence of high-quality data on crop yields at fine spatial resolutions in many regions of the world makes it challenging to statistically model their response to interannual variability in weather patterns. Therefore, there is a need for empirical methods that can project annual crop yield changes even in limited data regions. Here, we propose a transfer learning algorithm that uses high spatial resolution data from one region to project yields in another region with more limited data. The goal of our work is to understand what data types can be beneficial for transferring learning from a source region to a very different target region with more limited data. We utilize Long Short-Term Memory to develop a transfer learning model that is trained on historical county-level corn yield in the United States and predicts district-level corn yield variations in India. Even using smaller amounts of data in India, simulating a data-scarce region, we achieve an average root mean square error of 0.48 bu acre−1 in predicting interannual yield variations. Using Shapley values to interpret results, we explore the contribution of the different weather parameters to interannual yield variability and find a larger influence of precipitation-related variables. Our study demonstrates the usefulness of this method for transferring models of weather impacts on crop yields trained on a data-rich country to one with more limited data. It suggests the potential of applying the transfer learning model to mitigate the need for extensive raw data globally.
Abstract Annual wheat yields have steadily risen over the past century, but harvests remain highly variable and dependent on myriad weather conditions during a long growing season. In Kansas, for example, the 2014 crop year brought the lowest average yield in decades at 28 bushels per acre, while in 2016 farmers in the Wheat State, as Kansas is often called, enjoyed a historic high of 57 bushels per acre. It is broadly known that remote forces like El Niño–Southern Oscillation contribute to meteorological outcomes across North America, including in the wheat-growing regions of the U.S. Midwest, but the differential imprints of ENSO phases and flavors have not been well explored as leading indicators for harvest outcomes in highly specific agricultural regions, such as the more than 7 million acres upon which wheat is grown in Kansas. Here, we demonstrate a strong, steady, and long-term association between a simple “wheat yield index” and sea surface temperature anomalies, more than a year earlier, in the East Pacific, potentially offering insights into forthcoming harvest yields several seasons before planting commences.
Phenotyping high-biomass perennial crops is laborious and the rate of genetic gain in conventional perennial crop breeding programs is typically low. So, it is especially important to identify methods that produce efficiency gains in the breeding process. Miscanthus is a C4 perennial grass with favorable characteristics for producing biomass as a feedstock for biofuels and diverse bio-based products. Increasing biomass yield will increase profitability and environmental benefits, so it is a key target for Miscanthus breeding. In addition, the identification of well-adapted genotypes across a wide range of environmental conditions requires the establishment of multi-environment trials (METs). Sparse testing is a genomic prediction-based strategy that reduces the phenotyping costs in METs by selecting a subset of genotypes to evaluate in a subset of environments and then predicts the performance of the unobserved genotype-environment combinations. A Miscanthus sacchariflorus (MSA) population comprising 336 genotypes observed across three environments was analyzed implementing sparse testing designs. Three prediction models considering main effects (environments, genotypes, genomic) and interaction effects (genotype-by-environment; G×E interaction) were implemented for forecasting dry biomass yield (YDY), total culm (TCM), average internode length (AIL), and culm node number (CNN). Multiple calibration sets based on different compositions and sizes were considered to evaluate performance in terms of the predictive ability (PA) and the mean square error (MSE) for a fixed testing set size. The training set size ranged from 52 to 112 to predict a fixed set of 224 unobserved genotypes across all three environments. The results showed that the model accounting for G×E interaction consistently presented the highest PA and the lowest MSE: for CNN (PA: ~0.77, MSE: ~0.5) and YDY (PA: ~0.70, MSE: ~1.3) while for TCM and AIL these ranged from ~0.28 to 0.41 and ~1.3 to 4.3, respectively. Overall, varying training sets and allocation strategies did not affect PA and MSE, with 52 non-overlapping and 0 overlapping genotypes per environment as the optimal cost-effective allocation framework. This suggests that implementing sparse testing designs could significantly reduce phenotyping costs by fivefold, without compromising PA in breeding programs for perennial crops such as Miscanthus.
ABSTRACT Agriculture is crucial for global food supply and dominates the Earth's land surface. It is unknown, however, how slow but relentless changes in climate mean state, versus random extreme conditions arising from changing variability , will affect agroecosystems' carbon fluxes, energy fluxes, and crop production. We used an advanced weather generator to partition changes in mean climate state versus variability for both temperature and precipitation, producing forcing data to drive factorial‐design simulations of US Midwest agricultural regions in the Energy Exascale Earth System Model. We found that an increase in temperature mean lowers stored carbon, plant productivity, and crop yield, and tends to convert agroecosystems from a carbon sink to a source, as expected; it also can cause local to regional cooling in the earth system model through its effects on the Bowen Ratio. The combined effect of mean and variability changes on carbon fluxes and pools was nonlinear, that is, greater than each individual case. For instance, gross primary production reduces by 9%, 1%, and 13% due to change in mean temperature, change in temperature variability, and change in both temperature mean and variability, respectively. Overall, the scenario with change in both temperature and precipitation means leads to the largest reduction in carbon fluxes (−16% gross primary production), carbon pools (−35% vegetation carbon), and crop yields (−33% and −22% median reduction in yield for corn and soybean, respectively). By unambiguously parsing the effects of changing climate mean versus variability and quantifying their nonadditive impacts, this study lays a foundation for more robust understanding and prediction of agroecosystems' vulnerability to 21st‐century climate change.
The cultivation of sterile giant miscanthus (Miscanthus × giganteus, M × g) for bioenergy and bioproducts has expanded into grain-cropped land in the United States (US) as local markets developed for this high-yielding perennial grass (10–30 Mg DM ha −1 ). However, the magnitude of spatial and temporal variability in yield within US Corn Belt fields, along with impacts on economic return and sustainable land management, is poorly understood. This study established a diagnostic model relating remote sensing-derived vegetation indices to ground truth data from 105 hand-harvested stem biomass samples, which were strategically selected to represent the full range of vegetation index observations. The high-resolution satellite-sensed vegetation indices captured > 90% of the yield variation measured within fields. This model was then used to predict yield variability and assess economic performance across four of the first commercial M × g fields in the Corn Belt state of Iowa, US. Significant spatial variability in biomass dry matter (DM) yields (9.3–18.1 Mg DM ha −1 ) and net profits ($\$$83 to $\$$1211.5 ha −1 ) was observed. All fields were profitable in all site-years. When low profit occurred, it was explained by limited management experience of the crop in Iowa. The breakeven yield at a selling price of $\$$130 Mg −1 varied from 9.0–12.1 Mg ha −1 at 15% moisture content (7.6–10.3 Mg DM ha −1 ). Breakeven prices ranged from $\$$73 to $\$$122.4 Mg −1 , matching ranges used in the Department of Energy Billion Ton Report (US Department of Energy, 2023). Notably, M × g yield and profits were commensurate with grain crops particularly with favorable precipitation. This study provides insight on the M × g management “learning curve”, performance on marginal land and in drought conditions, and demonstrates that addressing yield gaps, reducing costs, and implementing precision agriculture strategies can enhance profitability. These findings emphasize the value of remote sensing technologies in guiding sustainable and competitive commercial-scale M × g production.
Many winter annual crops, such as pennycress ( Thlaspi arvense L.), are subjected to heavy precipitation events during their growing season. Therefore, it is essential to identify pennycress accessions with natural variation in flooding resilience. We used climate modeling data to assess spring soil moisture levels in the geographic origins of 471 natural pennycress accessions. We selected 34 accessions with variation in predicted soil moisture and tested survivability under prolonged waterlogging at the rosette stage. It took seven weeks for the first accessions to die, indicating that pennycress is hardy to prolonged waterlogging at the vegetative stage. Furthermore, we chose ‘susceptible’ and ‘tolerant’ accessions to waterlog for one week at the reproductive stage, the growth stage aligned with spring rainfall. Six accessions had significantly reduced seed weight at maturity and two had minimal impacts on growth and seed yield after waterlogging and can be further explored for adaptive traits.
Nutrient exports from agricultural lands in the Great Lakes Region pose significant threats to water quality and ecological health through eutrophication, hypoxia, and harmful algal blooms. Climate change and agricultural adaptation practices complicate future nutrient loading due to intensified hydrologic cycles and land use decisions. Our research focuses on evaluating the Soil and Water Assessment Tool (SWAT) plus model parametric uncertainties for human and natural outcomes across different scales. These factors are integral to ensuring a balance between productive agricultural practices and maintaining the health of watershed hydrology. However, uncertainties in modeling such complex interactions pose significant challenges, limiting our ability to precisely determine critical factors that influence crop yield and soil moisture. Our analysis employs Sobol global sensitivity analysis to evaluate first-order, second order, and total-order indices for SWAT crop growth parameter, ensuring comprehensive assessment of individual and interactive effects on model outputs. The objective is to identify the parameters that significantly affect model outputs for crop yield and soil moisture and improve our understanding of their interactions at the basin and hydrological response unit (HRU) scale. Our case study, the Portage River Watershed, which drains into Lake Erie, is chosen to better capture finer scale interactions crucial for predicting nutrient loading under future climate scenarios. This foundational work is aimed at setting the stage for the future development of an agent-based model (ABM). The ABM model would incorporate SWAT outputs to dynamically simulate decision-making processes.
Abstract The chemical composition of growing media is a key factor for plant growth, impacting agricultural yield and sustainability. However, there is a lack of affordable chemical sensors for ubiquitous nutrient ion monitoring in agricultural applications. This work investigates using fully printed ion‐sensor arrays to measure the concentrations of nitrate, ammonium, and potassium in mixed‐electrolyte media. Ion sensor arrays composed of nitrate, ammonium, and potassium ion‐selective electrodes and a printed silver‐silver chloride (Ag/AgCl) reference electrode are fabricated and characterized in aqueous solutions in a range of concentrations that encompass what is typical for agricultural growing media (0.01 m m –1 m ). The sensors are also tested in mixed‐electrolyte solutions of NaNO 3 , NH 4 Cl, and KCl of varying concentrations, and the recorded potentials are input into Nernstian and artificial neural network models to compare the prediction accuracy of the models against ground truth. The artificial neural network models demonstrated higher accuracy over the Nernstian model, and the model using only ion‐sensor inputs is 7.5% more accurate than the Nernstian model under the same conditions. By enabling more precise and efficient fertilizer application, these sensor arrays coupled to computational models can help increase crop yields, optimize resource use, and reduce environmental impact.