Search NASA⌕ Search

SEARCH · Search NASA

Results for “variable”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28

Enhancing Interpretability in Generative Modeling: Statistically Disentangled Latent Spaces Guided by Generative Factors in Scientific Datasets

This study addresses the challenge of statistically extracting generative factors from complex, high-dimensional datasets in unsupervised or semi-supervised settings. We investigate encoder-decoder-based generative models for nonlinear dimensionality reduction, focusing on disentangling low-dimensional latent variables corresponding to independent physical factors. Introducing Aux-VAE, a novel architecture within the classical Variational Autoencoder framework, we achieve disentanglement with minimal modifications to the standard VAE loss function by leveraging prior statistical knowledge through auxiliary variables. These variables guide the shaping of the latent space by aligning latent factors with learned auxiliary variables. We validate the efficacy of Aux-VAE through comparative assessments on multiple datasets, including astronomical simulations.

97 MATHEMATICS AND COMPUTING↗

Second-generation downscaled earth system model data using generative machine learning

The second-generation Sup3rCC dataset provides high-resolution meteorological data generated through the downscaling of multiple earth system models (ESMs) from the Coupled Model Intercomparison Project Phase 6 (CMIP6). This downscaling is performed through application of a generative machine learning approach called Super-Resolution for Renewable Resource Data (sup3r). This dataset builds on the first-generation Sup3rCC data by applying improved bias correction methods and adding downscaled precipitation to the output variables. As with the first Sup3rCC version, the data still include temperature, wind speed and direction at multiple heights, pressure, three components of downwelling solar radiation, and relative humidity—all at 4-kilometer (km) hourly resolution over the contiguous United States. This is a 25x spatial enhancement and 24x temporal enhancement of the source 100-km daily-average ESM data. This extension of the Sup3rCC dataset includes data from six ESMs from two shared socioeconomic pathways (SSPs) totaling 400 years of data with multiple future projections of changing meteorological conditions. The scenario selection was based on a structured evaluation of historical ESM skill and comprehensive representation of possible trajectories of future climate change in temperature, humidity, precipitation, solar irradiance, and near-surface wind speeds. The inclusion of multiple future projections is intended to enable users to assess key drivers of un 36 certainty and variability. All data are double-bias corrected, resulting in a product that can be used out-of-the-box for energy system analysis with minimal historical bias. The potential applications of Sup3rCC data extend to various topics in renewable energy resource assessment, energy systems modeling, and grid resilience studies. High-resolution future meteorological projections are critical for evaluating the effects of changing meteorological conditions on renewable energy generation, energy demand, and for optimizing energy storage and grid infrastructure. The 4-km hourly resolution of the downscaled data enables understanding of spatial and temporal variability at the scales necessary for energy system operational planning. In addition, the dataset can support risk assessments by providing detailed information on possible future extreme weather events and long-term meteorological variability at scales relevant to energy infrastructure. By offering an enhanced representation of possible future meteorological conditions, the second-generation Sup3rCC dataset enables more precise modeling of energy resilience and adaptation strategies in response to changing meteorological conditions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Simulating water dynamics related to pedogenesis across space and time: Implications for four-dimensional digital soil mapping

Digital soil mapping (DSM) relies on machine-learning and geostatistics to represent soil property observations across space. DSM techniques are powerful but often empirical, being limited to the quality and density of point samples. Water dynamics are closely related to soil variability, and the physics that govern water movement are well known. Hydrological properties can hence be simulated by physical models through space and time, unveiling key characteristics about soils. We propose the use of hydrologic models to map soils across the surface (2D), depth (1D), and time (1D)–which provides a 4D approach to digital soil mapping (4DSM). The Distributed Hydrology Soil Vegetation Model (DHSVM) was applied to a watershed currently under pasture. Moisture sensors and wells were installed at different depths in the watershed on summit, sideslope and toeslope positions to validate the model. DHSVM simulations of soil moisture distribution and depth to saturation were performed during the hydrological year (October 2008-September 2009). Clusters of similar pixels based on soil moisture values were determined using Dynamic Time Warping (DTW) to align temporal data and K-means. Clustering was performed both seasonally and for the entire year. Temporal patterns simulated by DHSVM matched measurements given by moisture sensors and wells. Seasonal clusters differed from the annual cluster. Distinct clusters were observed for each season and with depth, showing that spatiotemporal soil variability is lost when statically assessing soils. Spatiotemporal clusters corroborated field observations of fragipan occurrence not explicitly spatially mapped by Soil Survey Geographic Database (SSURGO). If a connection can be made between water and soils, static and dynamic soil variability can be predicted using physically based hydrologic models. Hydrologic models can benefit soil mapping by enabling reliable 4D simulation of water dynamics, which are fundamental to soil variability and soil classification and directly relate to biological, physical and chemical soil processes not captured by typical soil sampling protocols.

54 ENVIRONMENTAL SCIENCES↗

Understanding the structural and morphological effects of synthesis route on NpO 2

The availability of actinide standard materials for use in nuclear safeguard applications is critical, as is thorough characterization thereof. Although accurate trace element compositions and isotopic considerations are paramount for deployment of reference standards, structural characterization is also essential towards accurately describing the chemical form and potential matrix effects in candidate materials. Here, to this end, samples of NpO 2 were synthesized via a direct denitration (DD) method and probed with powder X-ray diffraction (PXRD), Raman spectroscopy, and scanning electron microscopy (SEM) for structural and morphological characterization and comparison with NpO 2 materials produced via modified direct denitration (MDD). PXRD confirmed the bulk identity of NpO 2 , and no additional phases were identified using this method. Analysis of Raman data collected using a 532 nm excitation wavelength indicates that samples are mostly phase pure; however, some variability in spectral features is observed. Analysis of additional spectroscopic data collected with a 785 nm excitation wavelength revealed variability in the relative intensity of spectral features. Raman spectroscopy indicates that the sample is primarily NpO 2 ; however, additional signals indicate possible structural disorder, oxidized species, or potential contributions from other Np phases. To further investigate the possibility of additional phase contributions within the sample of NpO 2 , Raman spectroscopic mapping was employed to examine the homogeneity of the sample produced via DD. From this analysis, we determined that despite variability in the intensity of Raman-active vibrational modes, consistent spectra are obtained throughout the area of the sample investigated. SEM images show aggregates with variable sizes and shapes, with rounded, primary particles possessing an average diameter of approximately 100 nm. Comparison of the results of these multimodal analyses to the literature indicates that the crystal chemical, spectroscopic, and microstructural properties of NpO 2 vary based on synthesis method, even if X-ray diffraction data indicate that the bulk phase is NpO 2 .

Direct denitration↗

A Bayesian inferencing framework for ultrasound wave speed measurements in metal additive manufacturing

Process-related changes during metal additive manufacturing introduce microstructural variability in the material properties of printed parts, directly affecting component reliability. Accurate estimation of these property variations with part performance are essential for quality assurance. Ultrasound testing offers a non-destructive means to estimate mechanical properties and detect defects; however, conventional analysis methods often neglect the influence of microstructural variability, limiting their effectiveness. Here, this research presents a Bayesian inference technique for quantifying wave speed uncertainty from ultrasound measurements of metal additive manufactured parts. By integrating prior ultrasound data with a Bayesian model, the proposed approach generates posterior density estimates of wave speed that systematically account for manufacturing-induced variability and uncertainty. The novelty of this research lies in applying a Bayesian framework to analyze experimental ultrasound measurements within the context of metal additive manufacturing variability. The method enhances the accuracy of wave speed estimation by 64%, defect position by 50% and increases confidence associated with wave speed variance by 30% across different porosity levels, thereby providing a robust foundation for improved decision-making and increased reliability in additively manufactured components.

Additive manufacturing↗

Formation of Linear Plasmonic Heterotrimers Using Nanoparticle Docking to DNA Origami Cages

The fabrication of complex assemblies with interesting collective properties from plasmonic nanoparticles (NPs) is often challenging. While DNA-directed self-assembly has emerged as one of the most promising approaches to forming such complex assemblies, the resulting structures tend to have large variability in gap sizes and shapes, as the DNA strands used to organize these particles are flexible, and the polydispersity of the NPs leads to variability in these critical structural features. Here, we use a new strategy termed docking to DNA origami cages (D-DOC) to organize spherical NPs into a linear heterotrimer with a precisely defined geometrical arrangement. Instead of binding NPs to the exterior of the DNA templates, D-DOC binds the NPs to either the interior or the opening of a 3D cage, which significantly reduces the variability of critical structural features by incorporating multiple diametrically arranged capture strands to tether NPs. Additionally, such a spatial arrangement of the capture strand can work synergistically with shape complementarity to achieve tighter confinement. To assemble NPs via D-DOC, we developed a multistep assembly process that first encapsulates an NP inside a cage and then binds two other NPs to the openings. Microscopic characterization shows low variability in the bond angles and gap sizes. Both UV–vis absorption and surface-enhanced Raman scattering (SERS) measurements showed strong plasmonic coupling that aligned with predictions by electrodynamic simulations, further confirming the precision of the assembly. These results suggest D-DOC could open new opportunities in biomolecular sensing, SERS and fluorescence spectroscopies, and energy harvesting through the self-assembly of NPs into more complex 3D assemblies.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Expanded Understanding of the Western Antarctic Peninsula Sea‐Ice Environment Through Local and Regional Observations at Palmer Station

Abstract The Western Antarctic Peninsula (WAP) has been experiencing rapid regional warming since at least the 1950s, however, the impacts of this warming at the local scale are variable and nuanced. Previous studies that have linked sea‐ice variability to biogeochemical cycles and food web dynamics often combine local‐scale biogeochemical data with coarse‐resolution regional satellite sea‐ice data, which may not adequately capture local sea‐ice conditions. In this study, we analyzed local‐scale in situ sea‐ice observations collected as part of a 28‐year record (1992–2020) from the Palmer Long‐Term Ecological Research site at Anvers Island, mid‐WAP, in conjunction with isotopically‐derived sea‐ice meltwater (SIM) fractions and satellite‐derived sea‐ice motion and concentration, to quantify the variability and long‐term trends in local sea‐ice behavior. In situ sea ice observations at Palmer Station displayed higher variability than satellite observations and showed no significant declines over this time, despite region‐wide declines identified in prior studies. Higher spring SIM fractions were attributed to strong northward sea‐ice motion throughout the winter. Applying these local‐scale sea‐ice insights to similarly scaled stratification and chlorophyll‐ a measurements, we found that a longer‐lasting, more consistent sea‐ice pack led to greater water column stratification following the spring sea‐ice retreat. Greater sea‐ice persistence and stronger stratification led to larger peaks in chlorophyll‐ a , though sea‐ice metrics did not explain the positive temporal trends in either stratification strength or chlorophyll‐ a . Through this study, we identify how local sea‐ice observations and meltwater data can enhance satellite data to build an understanding of the intricate connections between ice, water column dynamics, and phytoplankton.

Goodell, E.↗

Responses of Marginal and Intrinsic Water-Use Efficiency to Changing Aridity Using FLUXNET Observations

According to classic stomatal optimization theory, plant stomata are regulated to maximize carbon assimilation for a given water loss. A key component of stomatal optimization models is marginal water-use efficiency (mWUE), the ratio of the change of transpiration to the change in carbon assimilation. Although the mWUE is often assumed to be constant, variability of mWUE under changing hydrologic conditions has been reported. However, there has yet to be a consensus on the patterns of mWUE variabilities and their relations with atmospheric aridity. We investigate the dynamics of mWUE in response to vapor pressure deficit (VPD) and aridity index using carbon and water fluxes from 115 eddy covariance towers available from the global database FLUXNET. We demonstrate a non-linear mWUE-VPD relationship at a sub-daily scale in general; mWUE varies substantially at both low and high VPD levels. However, mWUE remains relatively constant within the mid-range of VPD. Despite the highly non-linear relationship between mWUE and VPD, the relationship can be informed by the strong linear relationship between ecosystem-level inherent water-use efficiency (IWUE) and mWUE using the slope, m *. We further identify site-specific m * and its variability with changing site-level aridity across six vegetation types. We suggest accurately representing the relationship between IWUE and VPD using Michaelis–Menten or quadratic functions to ensure precise estimation of mWUE variability for individual sites.

54 ENVIRONMENTAL SCIENCES↗

Assessing Radiative Feedbacks and Their Contribution to the Arctic Amplification Measured by Various Metrics

Arctic amplification (AA), characterized by a more rapid surface air temperature (SAT) warming in the Arctic than the global average, is a major feature of global climate warming. Various metrics have been used to quantify AA based on SAT anomalies, trends, or variability, and they can yield quite different conclusions regarding the magnitude and temporal patterns of AA. This study examines and compares various AA metrics for their temporal consistency in the region north of 70°N from the early twentieth to the early 21st century using observational data and reanalysis products. We also quantify contributions of different radiative feedback mechanisms to AA based on short-term climate variability in reanalysis and model data using the Kernel-Gregory approach. Albedo and lapse rate feedbacks are positive and comparable, with albedo feedback being the leading contributor for all AA metrics. The net cloud feedback, which has large uncertainties, depends strongly on the data sets and AA metrics used. By quantifying the influence of internal variability on AA and related feedbacks based on global climate model ensemble simulations, we find that water vapor and cloud feedbacks are most heavily affected by internal variability.

54 ENVIRONMENTAL SCIENCES↗

Climate, Hydrology, and Nutrients Control the Seasonality of Si Concentrations in Rivers

Abstract The seasonal behavior of fluvial dissolved silica (DSi) concentrations, termedDSi regime, mediates the timing of DSi delivery to downstream waters and thus governs river biogeochemical function and aquatic community condition. Previous work identified five distinct DSi regimes across rivers spanning the Northern Hemisphere, with many rivers exhibiting multiple DSi regimes over time. Several potential drivers of DSi regime behavior have been identified at small scales, including climate, land cover, and lithology, and yet the large‐scale spatiotemporal controls on DSi regimes have not been identified. We evaluate the role of environmental variables on the behavior of DSi regimes in nearly 200 rivers across the Northern Hemisphere using random forest models. Our models aim to elucidate the controls that give rise to (a) average DSi regime behavior, (b) interannual variability in DSi regime behavior (i.e., Annual DSi regime), and (c) controls on DSi regime shape (i.e., minimum and maximum DSi concentrations). Average DSi regime behavior across the period of record was classified accurately 59% of the time, whereas Annual DSi regime behavior was classified accurately 80% of the time. Climate and primary productivity variables were important in predicting Average DSi regime behavior, whereas climate and hydrologic variables were important in predicting Annual DSi regime behavior. Median nitrogen and phosphorus concentrations were important drivers of minimum and maximum DSi concentrations, indicating that these macronutrients may be important for seasonal DSi drawdown and rebound. Our findings demonstrate that fluctuations in climate, hydrology, and nutrient availability of rivers shape the temporal availability of fluvial DSi.

Environmental Sciences & Ecology↗

Comparison of Global Aboveground Biomass Estimates From Satellite Observations and Dynamic Global Vegetation Models

The global forest carbon stocks represent the amount of carbon stored in woody vegetation and are important for quantifying the ability of the global forests to sequester atmospheric CO 2 and to provide ecosystem services (e.g., timber) under climate change. The forest ecosystem carbon pool estimates are highly variable and poorly quantified in areas lacking forest inventory estimates. Here, we compare and analyze aboveground biomass (AGB) estimates from five satellite-based global data sets and nine dynamic global vegetation models (DVGMs). We find that across the data sets, mean AGB exhibits the largest variability around the tropical area. In addition, AGB shows a similar latitudinal trend but large variability among the data sets. Satellite-based AGB estimates are lower than those simulated by DVGMs. The divergence among the satellite-based AGB estimates can be driven by the methodology, input satellite products, and the forested areas used to estimate AGB. The modeled NPP, autotrophic respiration, and carbon allocation mostly drive the variability of AGB simulated by DGVMs. The future availability of a high-quality global forest area map is anticipated to improve AGB estimate accuracy and to reduce the discrepancies among different satellite- and model-based AGB estimates. Furthermore, we suggest the carbon-modeling community reexamine the methodology used to estimate AGB and forested areas for a more robust global forest carbon stock estimation.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning Calibration of Groundwater Table Depth in ELM: Impact on Land Surface Hydrology and Land‐Atmosphere Fluxes

Accurate representation of groundwater table depth (GWTD) is crucial for simulating hydrological cycling in Earth system models (ESM). Nevertheless, there is a notable gap in the literature regarding the validation of GWTD simulations in ESMs and their subsequent impact on downstream hydrological components. This study explores the calibration of parameterization of global GWTD using machine learning within the Energy Exascale Earth System Model (E3SM) Land Model (ELM). Despite achieving significant gains in simulating GWTD through calibration, offline ELM simulations unexpectedly show that these improvements do not translate to substantial enhancements in model performance for other key hydrological variables, including soil moisture (SM), runoff, groundwater contribution to runoff or base flow index (BFI), and evapotranspiration and its partitioning. The performance in SM and runoff was even degraded in some regions, while BFI was mostly overestimated. Although there is significant improvement in GWTD within the critical range of 1–5 m, where groundwater traditionally influences land surface energy fluxes, these improvements occurred mostly in humid areas where the impact of GWTD on surface processes is minimal. Although the impacts of model calibration are generally small in offline ELM simulations, coupled land-atmosphere simulations exhibit much stronger responses to GWTD calibration, highlighting the role of land-atmosphere feedbacks in Earth system modeling. These findings underscore the need for integrated calibration strategies that simultaneously optimize multiple hydrological variables. However, if a single-variable approach is necessary, it is crucial to establish clear priorities for calibration, identifying the most critical variables that have the greatest impact on overall model performance.

Fang, Yilin [Pacific Northwest National Laboratory↗

Unified 0.25-degree gridded infrastructure-critical extreme weather for the United States from 1979 to 2100

Extreme weather events can severely disrupt critical infrastructure, triggering cascading effects on power, transportation, and essential services. However, standard weather and climate datasets often lack specialized variables necessary for hazard assessments. We present a unified dataset of infrastructure-critical weather and climate variables across the United States at 0.25° resolution, covering daily or sub-daily intervals from 1979 to 2100. The dataset includes temperature, dew point, wind gusts, precipitation partitioned by rain, snow, and freezing rain or ice pellets, lightning, and wildfire metrics. Historical conditions (1979-2023) are synthesized from observations and reanalysis products, while future projections are derived from 14 CMIP6 global climate models (historical, SSP245, and SSP585 experiments). Physically based and data-driven methods are used to estimate variables not directly provided by existing models. By integrating these variables into a single unified dataset, we enable consistent, high-resolution assessments of weather-related infrastructure risks across past and future periods, supporting wide-ranging applications in energy, transportation, water resources, emergency management, and beyond.

Climate and Earth system modelling↗

Accelerating the design of lattice structures using machine learning

Lattices remain an attractive class of structures due to their design versatility; however, rapidly designing lattice structures with tailored or optimal mechanical properties remains a significant challenge. With each added design variable, the design space quickly becomes intractable. To address this challenge, research efforts have sought to combine computational approaches with machine learning (ML)-based approaches to reduce the computational cost of the design process and accelerate mechanical design. While these efforts have made substantial progress, significant challenges remain in (1) building and interpreting the ML-based surrogate models and (2) iteratively and efficiently curating training datasets for optimization tasks. Here, we address the first challenge by combining ML-based surrogate modeling and Shapley additive explanation (SHAP) analysis to interpret the impact of each design variable. We find that our ML-based surrogate models achieve excellent prediction capabilities (R 2 > 0.95) and SHAP values aid in uncovering design variables influencing performance. We address the second challenge by utilizing active learning-based methods, such as Bayesian optimization, to explore the design space and report a 5 × reduction in simulations relative to grid-based search. Collectively, these results underscore the value of building intelligent design systems that leverage ML-based methods for uncovering key design variables and accelerating design.

36 MATERIALS SCIENCE↗

Challenges and alternatives to empirical orthogonal functions for earth system data

Empirical orthogonal functions (EOFs) applied to gridded Earth system data enables users to diagnose modes of variability with relative ease. Yet, many challenges to interpretation exist such that they must be used with awareness and intention when applied to gridded climate data, especially with large ensembles. Utilizing data from two different Earth system modelling large ensemble frameworks, the Energy Exoscale Earth System Model and the Community Earth System Model, as well as reanalysis data, common EOF pitfalls are summarized and discussed. Challenges include erroneous mode swapping, sign flipping, and the temporal variability of the centers of action. For modes of variability with similar contribution to variance, mode swapping is not uncommon. Sign flipping can occur with almost any mode where the pattern is correct, but the sign is arbitrary. Although the variability of the center of action is not necessarily problematic, it potentially complicates interpretation over multi-century timescales. A wide variety of alternative methods to EOFs exist, but fitness-for-purpose must be evaluated. Additionally, illustrations of alternative methods and examples of proper use are provided. Alternative methods fit into three categories: EOF variants, linear methods, and multilinear methods.

54 ENVIRONMENTAL SCIENCES↗

Codiscovering graphical structure and functional relationships within data: A Gaussian Process framework for connecting the dots

Most problems within and beyond the scientific domain can be framed into one of the following three levels of complexity of function approximation. Type 1: Approximate an unknown function given input/output data. Type 2: Consider a collection of variables and functions, some of which are unknown, indexed by the nodes and hyperedges of a hypergraph (a generalized graph where edges can connect more than two vertices). Given partial observations of the variables of the hypergraph (satisfying the functional dependencies imposed by its structure), approximate all the unobserved variables and unknown functions. Type 3: Expanding on Type 2, if the hypergraph structure itself is unknown, use partial observations of the variables of the hypergraph to discover its structure and approximate its unknown functions. These hypergraphs offer a natural platform for organizing, communicating, and processing computational knowledge. While most scientific problems can be framed as the data-driven discovery of unknown functions in a computational hypergraph whose structure is known (Type 2), many require the data-driven discovery of the structure (connectivity) of the hypergraph itself (Type 3). We introduce an interpretable Gaussian Process (GP) framework for such (Type 3) problems that does not require randomization of the data, access to or control over its sampling, or sparsity of the unknown functions in a known or learned basis. Its polynomial complexity, which contrasts sharply with the super-exponential complexity of causal inference methods, is enabled by the nonlinear ANOVA capabilities of GPs used as a sensing mechanism.

Science & Technology - Other Topics↗

Data-based filtered dissipation rate modelling for multi-modal turbulent combustion: evaluating a priori model generalizability

Manifold-based models offer a computationally efficient alternative to directly transporting the thermochemical state in computational simulations of turbulent reacting flows, projecting the high-dimensional thermochemical state-space onto a low-dimensional manifold. Recent efforts have yielded a manifold-based model applicable to multi-modal combustion, enabling reconstruction of the thermochemical state from solutions to two-dimensional manifold equations in mixture fraction and generalized progress variable that are parameterised by three scalar dissipation rates. In coarse-grained simulations such as Large Eddy Simulation (LES), closure of the multi-modal manifold equations and subfilter variances/covariance requires closure of three filtered scalar dissipation rates. Here, the present work adopts a data-based approach, providing closure for the three filtered scalar dissipation rates via deep neural networks (DNNs). High-fidelity datasets corresponding to an autoigniting n-dodecane jet flame and a bluff body swirl-stabilized confined lifted spray flame of two aviation fuels (Jet-A and C1) with different ignition propensities are leveraged to generate training data that spans a diverse range of thermodynamic conditions and combustion modes, including low- and high-temperature ignition regimes in addition to premixed and nonpremixed behaviour. A final DNN model is trained to enforce inherent physical constraints by learning nonlinear functional transformations of the three filtered scalar dissipation rates. The generalizability of this constrained DNN model is demonstrated a priori via conditional statistics evaluated on the lifted spray flame with C1–a configuration that had not been included in the training data. Excellent DNN agreement with conditional DNS statistics is observed, and integrated gradients are computed to identify the most sensitive input variables. The similarity of the marginal PDFs of the most informative input variables and outputs across configurations are quantified via the Wasserstein metric, demonstrating that data-based models may successfully generalize to unseen parametric conditions so long as the most informative input variables share similar distributions across training and testing datasets.

Data-based modelling↗

Predicting weather impacts on corn production in a data-limited region using a transfer learning approach

The stability of food supply and prices may depend more on annual changes in yields from year-to-year variability in weather than on longer-term average changes from changing climatic conditions. However, the absence of high-quality data on crop yields at fine spatial resolutions in many regions of the world makes it challenging to statistically model their response to interannual variability in weather patterns. Therefore, there is a need for empirical methods that can project annual crop yield changes even in limited data regions. Here, we propose a transfer learning algorithm that uses high spatial resolution data from one region to project yields in another region with more limited data. The goal of our work is to understand what data types can be beneficial for transferring learning from a source region to a very different target region with more limited data. We utilize Long Short-Term Memory to develop a transfer learning model that is trained on historical county-level corn yield in the United States and predicts district-level corn yield variations in India. Even using smaller amounts of data in India, simulating a data-scarce region, we achieve an average root mean square error of 0.48 bu acre−1 in predicting interannual yield variations. Using Shapley values to interpret results, we explore the contribution of the different weather parameters to interannual yield variability and find a larger influence of precipitation-related variables. Our study demonstrates the usefulness of this method for transferring models of weather impacts on crop yields trained on a data-rich country to one with more limited data. It suggests the potential of applying the transfer learning model to mitigate the need for extensive raw data globally.

Vishwakarma, Srishti [ORNL] (ORCID:000000031674419↗