Search NASA⌕ Search

SEARCH · Search NASA

Results for “multiple regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Clear-sky detection for PV degradation analysis using multiple regression

A method is presented to detect clear-sky periods for plane-of-array irradiance time-averaged data that is based on the algorithm originally described by Reno and Hansen. Here we show this new method improves the state-of-the-art by providing accurate detection at longer data averaging intervals. Moreover, our new method detects clear periods in plane-of-array data, which is novel. The new method is developed by applying a Design of Experiment approach to optimize the parameters used in the Reno method, and Monte Carlo simulations are used to understand the robustness of the found parameters. Clear-sky detection accuracy is compared among four methods: the Reno method, the default clear-sky filter in RdTools, the Ellis method, and the method outlined in this work, using a hand-labeled two-year data set of 1-min plane-of-array irradiance for a fixed tilt system. The RdTools clear-sky filter is marred by excessive false positives. The other methods all perform well at 1-min data intervals; the method developed here provides more accurate detection at longer data averaging intervals. We show that the parameters are directly linked to the data frequency in the hope that these input variables may not have to be optimized for every data frequency and location. However, only a single fixed system in one location was carefully examined. Finally, we illustrate how accurate determination of clear-sky conditions helps to eliminate data noise and bias in the assessment of long-term performance of PV plants.

14 SOLAR ENERGY↗

Mountain Basin Controls on the Snow-to-Streamflow Signal: An AIC-Weighted Multiple Linear Regression Framework

A regression-based analysis quantifies how basin characteristics modulate the snow-to-streamflow signal. First, we use the ERA5-Land reanalysis gridded product (European Centre for Medium Range Weather Forecasts reanalysis 5 -Land component) for 4,655 hydrologic unit code - 10 (HUC10) mountain basins across the western United States (US) for water years 1987–2024. Linear regressions are performed for peak snow water equivalent (SWE) and annual streamflow for each mountain basin. Models use ordinary least squares in Python’s statsmodels package. After which, an Akaike Information Criterion (AIC)–weighted ensemble multiple linear regression (MLR) framework with 47 watershed traits is used to predict the linear regression coefficient of determination (r-squared) defining the ability of peak SWE to predict annual streamflow across all mountain basin. Predictor sets are constrained to avoid multicollinearity by excluding models with variance inflation factors (VIF) greater than 5. Mountain basin traits included in the MLR include seasonal climate, topography, vegetation type and structure, and bedrock geology. Accepted models are considered if their AIC is within 2.0 of the model with the minimum AIC, or best model. To compare predictor influence across acceptable models, we computed standardized regression coefficients. To evaluate structural redundancy among models, we constructed binary inclusion vectors for each acceptable model, denoting whether a predictor was present (1) or absent (0). Core predictor variables are defined as occurring in at least 67% of the acceptable models. For this regional analysis, only one model was found acceptable, with higher snow-to-streamflow translation (higher r-squared) occurring in colder mountain basins with higher relative winter precipitation, more snow accumulation and a lower fraction of annual precipitation that falls in the spring and summer. The second component of the data package uses previously published, high-resolution output from an integrated hydrological model of the East River watershed using the U.S. Geological Survey Groundwater and Surface water Flow model (GSFLOW, doi:10.15485/1998576). East River MLR expands upon the approach described above to explore the response of five streamflow metrics—annual streamflow, runoff efficiency, 7-day minimum flow, low-flow duration, and non-perennial stream fraction to snow system indicators including peak SWE, snow-covered area, snow disappearance date, and the fraction of basin area characterized by low-to-no snow, as well as seasonal precipitation and temperature, and annual hydrologic variables representing soil moisture, evapotranspiration (ET), the partitioning of incoming precipitation to evapotranspiration (ET/P), groundwater storage, and groundwater inflow to streams. MLR was done on all water years (P0: 1987-2024) and for each period as determined in the split analysis using pooled regression techniques (P1: 1987-2011 and P2: 2012-2024) to evaluate shifting predictor variable emphasis on streamflow generation. Results indicate that since 2012, peak SWE has lost statistical strength in its prediction of annual streamflow and runoff efficiency, and the indirect influence of spring temperature has emerged as critically important. Low-flow metrics remain largely influenced by soil moisture, vegetation water use and groundwater inflows with summer precipitation becoming a direct influence on minimum summer flow. Together, these data and Python-based analysis tools provide a framework for identifying the key watershed characteristics that control how streamflow responds to snow from year to year. The package also helps quantify uncertainty in statistical models and assess how snow–streamflow relationships vary across regions and over time. This dataset contains comma-separated values files (.csv), text files (.txt), python code files (.py), figure files (.png), and shapefiles (.cpg, .dbf, .prj, .sbn, .sbx, .shp, .xml). Further details on file contents and MLR execution can be found in the readme file and the FLMD files. Work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Statistical analysis and degradation pathway modeling of photovoltaic minimodules with varied packaging strategies

Degradation pathway models constructed using network structural equation modeling (netSEM) are used to study degradation modes and pathways active in photovoltaic (PV) system variants in exposure conditions of high humidity and temperature. This data-driven modeling technique enables the exploration of simultaneous pairwise and multiple regression relationships between variables in which several degradation modes are active in specific variants and exposure conditions. Durable and degrading variants are identified from the netSEM degradation mechanisms and pathways, along with potential ways to mitigate these pathways. A combination of domain knowledge and netSEM modeling shows that corrosion is the primary cause of the power loss in these glass/backsheet PV minimodules. We show successful implementation of netSEM to elucidate the relationships between variables in PV systems and predict a specific service lifetime. The results from pairwise relationships and multiple regression show consistency. This work presents a greater opportunity to be expanded to other materials systems.

electrical measurements↗

Multiclass Classification Using Bayesian Multivariate Adaptive Regression Splines

We present a new Bayesian model for the problem of multiclass classification. In this model, the probabilities of class membership of a given observation are determined by the mean of a latent Gaussian distribution. The mean functions of this latent distribution consist of combinations of highly flexible basis functions of the inputs: multivariate adaptive regression splines (MARS), first developed for multiple regression. We use reversible jump Markov chain Monte Carlo to make inference on the classification model, including the number of basis functions. We compare the probabilistic classification performance of our proposed approach to existing methods on simulated and benchmark data, and compare uncertainty estimates on simulated data. Our proposed method compares favorably with existing Bayesian and frequentist multiclass classification methods in out-of-sample probabilistic classification, and uncertainty estimation of these probabilistic classifications. We examine the fit of the proposed method to a data set of hurricane storm surge levels near Delaware Bay, US, and conclude that sea level rise is a key contributor to damage delivered by storm surge.

97 MATHEMATICS AND COMPUTING↗

RxnRover/amlro

AMLRO (Active Machine Learning Reaction Optimizer) is an open-source framework designed to accelerate chemical reaction optimization using active learning with classical machine learning regression models. AMLRO integrates space-filling sampling strategies (e.g., Sobol and Latin Hypercube sampling) with iterative model training, prediction, and experiment selection to efficiently navigate complex reaction spaces. The platform supports multiple regression models, flexible multi-objective definitions, and user-defined parameter bounds, enabling data-efficient optimization from small initial datasets. AMLRO is designed for ease of use by experimentalists and can operate as a standalone decision-support tool or be integrated into closed-loop automated experimentation workflows.

Kulathunga, Dulitha Prasanna [Iowa State Universit↗

The executive disruption model of tinnitus distress: Model validation in two independent datasets using factor score regression

This study presents the executive disruption model (EDM) of tinnitus distress and subsequently validates it statistically using two independent datasets (the Construction Dataset: n = 96 and the Validation Dataset: n = 200). The conceptual EDM was first operationalised as a structural causal model (construction phase). Then multiple regression was used to examine the effect of executive functioning on tinnitus-related distress (validation phase), adjusting for the additional contributions of hearing threshold and psychological distress. For both datasets, executive functioning negatively predicted tinnitus distress score by a similar amount (the Construction Dataset: β = −3.50, p = 0.13 and the Validation Dataset: β = −3.71, p = 0.02). Theoretical implications and applications of the EDM are subsequently discussed; these include the predictive nature of executive functioning in the development of distressing tinnitus, and the clinical utility of the EDM.

Clarke, Nathan A.↗

Jensen–Shannon divergence based novel loss functions for Bayesian neural networks

Bayesian neural networks (BNNs) are state-of-the-art machine learning methods that can naturally regularize and systematically quantify uncertainties using their stochastic parameters. Kullback–Leibler (KL) divergence-based variational inference used in BNNs suffer from unstable optimization and challenges in approximating light-tailed posteriors due to the unbounded nature of the KL divergence. To resolve these issues, we formulate a novel loss function for BNNs based on a new modification to the generalized Jensen–Shannon (JS) divergence, which is bounded. In addition, we propose a Geometric JS divergence-based loss, which is computationally efficient since it can be evaluated analytically. We found that the JS divergence-based variational inference is intractable, and hence employed a constrained optimization framework to formulate these losses. Our theoretical analysis and empirical experiments on multiple regression and classification data sets suggest that the proposed losses perform better than the KL divergence-based loss, especially when the data sets are noisy or biased. Specifically, there are approximately 5% and 8% improvements in accuracy for a noise-added CIFAR-10 dataset and a regression dataset, respectively. There is about 13% reduction in false negative predictions of a biased histopathology dataset. Additionally, we quantify and compare the uncertainty metrics for the regression and classification tasks.

97 MATHEMATICS AND COMPUTING↗

Semi-empirical model for Henry’s law constant of noble gases in molten salts

Henry’s law constant, which describes the proportionality of dissolved gas to partial pressure of free gas in liquid–gas equilibrium systems, can also be applied to mass transport applications. In this work, we investigated an approach for determining the solubility of noble gases in a molten salt liquid utilizing the equilibrium concept of Henry’s gas constant. Henry’s gas constant is described as a mathematical function dependent on the van der Waals radius of the noble gas and the temperature of the molten salt. The alteration in Gibbs free energy encompasses contributions from both surface and volume energies. Enthalpy and entropy are deduced from these surface and volume energies in the Gibbs free energy formulation. A comparative analysis was conducted between the conventional method and our proposed model. Moreover, useful chemical properties can be determined from examination of surface and volume energies. Our findings provide an accurate and general theory of Gibbs free energy that can be validated experimentally based on the model proposed herein. This work unifies the prediction of Henry gas constant and subsequently the entropy and enthalpy calculation for noble gases in a molten salt solution to a single functional form using van der Waals radius of the gas and temperature of the system. This functional form is then used to perform a multiple regression method to find two parameters corresponding to the surface energy and volume energy. These two parameters are consistent between all combinations of noble gas and molten salt.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Identification of Distorted Gamma-Ray Signature Patterns Using Digital Filtering and Auto-Associative Memory Implemented with a Hopfield Neural Network

The detection and identification of radioactive sources in search applications involve analyzing passive gamma-ray emissions from high-level radioactive materials. This process uses a mobile detector-spectrometer in a complex field test environment. Recently, the use of artificial intelligence for gamma-ray spectrum analysis has shown promising results. However, challenges persist in identifying isotopic signatures from spectral measurements that may be distorted due to source shielding, random variations in natural radioactive background, or insufficient measurement time to obtain clear spectral lines. Here, this paper presents a novel intelligent signature recognition method that combines digital filtering techniques with an artificial Hopfield Neural Network (HNN). The HNN leverages auto-associative memory to store training sample patterns and match them with incoming gamma spectra from distorted sources. It restores the testing sources’ measurements by finding the closest matching signature patterns in the spectral library. Before HNN recognition, the measured spectrum undergoes preprocessing with a digital image filter to reduce fluctuations. Performance of the proposed method is evaluated using a set of gamma-ray spectra measured with a sodium iodide detector. The data collected include measurements from six pure samples: 241 Am, 60 Co, 137 Cs, 192 Ir, 239 Pu, and 235 U, which are used for training and validation (i.e. six cases). Additionally, the data set contains 24 distorted synthesized sources with various fluctuating backgrounds. Test results demonstrate the potential of the proposed method to accurately recognize the correct isotope with high precision, achieving an accuracy rate exceeding 85%. Furthermore, the proposed method exhibits superior performance compared to the conventional multiple regression fitting and simple feedforward neural network methods.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Nitrogen fertilization effects on aged Miscanthus × giganteus stands: Exploring biomass yield, yield components, and biomass prediction using in–season morphological traits

For sustainable biomass production of Miscanthus × giganteus (hereafter miscanthus), understanding the impact of stand age and nitrogen (N) fertilization on biomass yield is crucial. This study investigated the effects of varying N fertilization rates (0, 56, 112, and 168 kg N ha –1 ) on yield components (tiller height, density, and weight) and their correlations with end-of-season biomass yield in miscanthus. We also explored end-of-season biomass yield prediction using in-season traits (canopy height, leaf area index, and leaf chlorophyll content [LCC]). The study was conducted at two sites in Illinois: a previously unfertilized 10-year-old miscanthus research stand at Urbana and a 16-year-old commercial stand at Pesotum with a history of annual 56N application. Results from 2018 to 2021 in Urbana and 2020 to 2021 in Pesotum showed increased biomass yields with N fertilization, varying by rate, year, and location. Biomass yield in Pesotum peaked at 56N, while in Urbana, it increased significantly at 112 kg N ha –1 . Biomass yield was strongly correlated with tiller height and weight measured at Urbana across N rates. Morphological traits measured every 2–3 weeks during the 2020 and 2021 growing seasons showed that canopy height was the strongest single predictor of miscanthus biomass yield, followed by LCC. Mid-August to September measurements of these traits were the best predictors of biomass yield. Multiple regressions involving the canopy height and LCC further improved yield predictions. We conclude that while N enhances biomass yields of aging miscanthus, the optimum rate depends on the site, environmental conditions, and management history.

59 BASIC BIOLOGICAL SCIENCES↗

Morphological traits for allometric scaling of the European Sea Bass Dicentrarchus labrax (Linnaeus, 1758) from Southern Portugal population

Abstract The present study aimed to determine the allometric scaling among a selection of morphological traits in European sea bass ( Dicentrarchus labrax ) to estimate fish body weight. A set of morphological traits (fish body weight, length, height, and width) were directly measured in 146 fish of a recirculating aquaculture system, with body weights ranging from 17.11 to 652.21 g. In addition, a collection of digital imagery of each anesthetized fish from the side and top views were used to estimate other traits (indirect measures). Multiple regression analysis and regression coefficients were calculated using all possible combinations of biometric data (predictors) to estimate fish body weight, applying different numerical fitting models (linear, log‐linear, quadratic, exponential). The results showed that the best combination of traits for estimating fish body weight were fish body width, length and height, collected from direct measure ( R 2 = 0.995), for a log‐linear model fitting, which revealed more accurate determinations than the most commonly used length–weight relationship. Nevertheless, other combinations of morphological traits and fitting models were also found to be suitable in successfully predict fish body weight, with variability ranging between 92.5% and 98.5%. For indirect measures, the best predictor was a combination of traits from top view (width, eye distance and area without fins) fitted with a log‐linear function. These results comprise a relevant baseline in supporting the high potential of noninvasive methods to accurately follow the growth of European sea bass juveniles, recurring to imagery analysis of anesthetized fish. It has major potential applications in feeding consumption trials and fish growth models, as it allows for continuously following up fish growth under different experimental conditions without therein distress derived from manipulation.

Azevedo, Ana↗

The Role of Data Filtering in Open Source Software Ranking and Selection

Faced with more than 100M open source projects, a more manageable small subset is needed for most empirical investigations. More than half of the research papers in leading venues investigated filtering projects by some measure of popularity with explicit or implicit arguments that unpopular projects are not of interest, may not even represent "real" software projects, or that less popular projects are not worthy of study. However, such filtering may have enormous effects on the results of the studies if and precisely because the sought-out response or prediction is in any way related to the filtering criteria.This paper exemplifies the impact of this common practice on research outcomes, specifically how filtering of software projects on GitHub based on inherent characteristics affects the assessment of their popularity. Using a dataset of over 100,000 repositories, we used multiple regression to model the number of stars -a commonly used proxy for popularity- based on factors such as the number of commits, the duration of the project, the number of authors and the number of core developers. Our control model included the entire dataset, while a second filtered model considered only projects with ten or more authors. The results indicated that while certain characteristics of the repository consistently predict popularity, the filtering process significantly alters the relationships between these characteristics and the response. We found that the number of commits exhibited a positive correlation with popularity in the control sample but showed a negative correlation in the filtered sample. These findings highlight the potential biases introduced by data filtering and emphasize the need for careful sample selection in empirical research of mining software repositories. We recommend that empirical work should either analyze complete datasets such as World of Code, or employ stratified random sampling from a complete dataset to ensure that filtering is not biasing the results.

Malviya Thakur, Addi↗

Investigating the ecological fallacy through sampling distributions constructed from finite populations

Correlation coefficients and linear regression values computed from group averages can differ from correlation coefficients and linear regression values computed using individual scores. This observation known as the ecological fallacy often assumes that all the individual scores are available from a population. In many situations, one must use a sample from the larger population. In such cases, the computed correlation coefficient and linear regression values will depend on the sample that is chosen and the underlying sampling distribution. The sampling distribution of correlation coefficients and linear regression values for group averages will be identical to the sampling distribution for individuals for normally distributed variables for random samples drawn from infinitely large continuous distributions. However, data that is acquired in practice is often acquired when sampling without replacement from a finite population. Our objective is to demonstrate through Monte Carlo simulations that the sampling distributions for correlation and linear regression will also be similar for individuals and group averages when sampling without replacement from normally distributed variables. These simulations suggest that when a random sample from a population is selected, the correlation coefficients and linear regression values computed from individual scores will not be more accurate in estimating the entire population values compared to samples when group averages are used as long as the sample size is the same.

97 MATHEMATICS AND COMPUTING↗

Social support and cognitive function in Chinese older adults who experienced depressive symptoms: is there an age difference?

Objective This study examined the moderating effect of overall social support and the different types of social support on cognitive functioning in depressed older adults. We also investigated whether the moderating effect varied according to age. Methods A total of 2,500 older adults (≥60 years old) from Shanghai, China were enrolled using a multistage cluster sampling method. Weighted linear regression and multiple linear regression was utilized to analyze the moderating effect of social support on the relationship between depressive symptoms and cognitive function and to explore its differences in those aged 60–69, 70–79, and 80 years and above. Results After adjusting for covariates, the results indicated that overall social support (β = 0.091, p = 0.043) and support utilization (β = 0.213, p < 0.001) moderated the relationship between depressive symptoms and cognitive function. Support utilization reduced the possibility of the cognitive decline in depressed older adults aged 60–69 years (β = 0.310, p < 0.001) and 80 years and above (β = 0.199, p < 0.001), while objective support increased the possibility of cognitive decline in depressed older people aged 70–79 years (β = −0.189, p < 0.001). Conclusion Our findings highlight the buffering effects of support utilization on cognitive decline in depressed older adults. We suggest that age-specific measures should be taken when providing social support to depressed older adults in order to reduce the deterioration of cognitive function.

Jing, Yurong↗

Data from: Biomass yield, yield components and growing season phenotypic measurements of Miscanthus

For sustainable biomass production of Miscanthus × giganteus (hereafter miscanthus), understanding the impact of stand age and nitrogen (N) fertilization on biomass yield is crucial. This study investigated the effects of varying N fertilization rates (0, 56, 112, and 168 kg N ha-1) on yield components (tiller height, density, and weight) and their correlations with end-of-season biomass yield in miscanthus. We also explored end-of-season biomass yield prediction using in-season traits (canopy height, leaf area index (LAI), and leaf chlorophyll content (LCC)). The study was conducted at two sites in Illinois: a previously unfertilized 10-year-old miscanthus research stand at Urbana and a 16-year-old commercial stand at Pesotum with a history of annual 56N application. Results from 2018-2021 in Urbana and 2020-2021 in Pesotum showed increased biomass yields with N fertilization, varying by rate, year, and location. Biomass yield in Pesotum peaked at 56N, while in Urbana, it increased significantly at 112 kg N ha-1. Biomass yield was strongly correlated with tiller height and weight measured at Urbana across N rates. Morphological traits measured every 2-3 weeks during the 2020 and 2021 growing seasons showed that canopy height was the strongest single predictor of miscanthus biomass yield, followed by LCC. Mid-August to September measurements of these traits were the best predictors of biomass yield. Multiple regressions involving the canopy height and LCC further improved yield predictions. We conclude that while N enhances biomass yields of aging miscanthus, the optimum rate depends on the site, environmental conditions, and management.

Aging↗

Knowledge of lactation amenorrhea method among postpartum women in Ethiopia: a facility-based cross-sectional study

While the importance of knowledge about contraceptives in improving their utilization and thereby reducing the risk of unintended pregnancies is well documented, there are limited studies documented about the Lactational Amenorrhea Method (LAM). Thus, understanding the knowledge of postpartum mothers about LAM is essential for designing tailored interventions. This study assessed the level of knowledge about LAM and its associated factors among postpartum mothers in Ethiopia. A facility-based cross-sectional study was conducted among 3148 randomly selected postpartum participants. The study utilized multistage sampling approach in hospitals located across five regions and one city administration in Ethiopia. Data were collected using face-to-face interviews at discharge. A participant was categorized as having knowledge of LAM if she correctly answered the three LAM criteria: amenorrhea, the first 6 months, and exclusive breast feeding. A binary logistic regression model was used to identify factors associated with knowledge of LAM. Variables with p < 0.25 in the binary logistic regression were included in the multiple logistic regression. Then, associations were described using the adjusted odds ratio (AOR) along with the 95% confidence interval (CI), and statistical significance was declared at p < 0.05. Only four in 10 participants (40.6%; 95% CI 38.9–42.3) had knowledge of LAM. Participants who attended college or above educational level (AOR = 2.1, 95% CI 1.5–2.8), those with parity of two (AOR = 2.3; 95% CI 1.6–3.6) or more than two (AOR = 2.4; 95% CI 1.5–4.0), those who expressed a desire for further fertility (AOR = 1.3; 95% CI 1.1–1.5), individuals who received counselling on LAM (AOR = 3.0; 95% CI 2.6–3.7), and those who gave birth in hospital (AOR = 2.6; 95% CI 1.4–2.6) had higher odds of knowledge about LAM, compared to their counter parts. In contrary, participants resided far away from health facilities had 30% lower odd of knowledge about LAM compared to those resided near the health facilities (AOR = 0.70; 95% CI 0.6–0.8). The proportion of participants who had knowledge of LAM was low. Strengthening counseling about LAM during antenatal care and delivery with due attention to women with limited access to health facilities should be considered for increasing their level of knowledge on LAM.

60 APPLIED LIFE SCIENCES↗

DeepPhenoMem V1.0: deep learning modelling of canopy greenness dynamics accounting for multi-variate meteorological memory effects on vegetation phenology

Abstract. Vegetation phenology plays a key role in controlling the seasonality of ecosystem processes that modulate carbon, water and energy fluxes between the biosphere and atmosphere. Accurate modelling of vegetation phenology in the interplay of Earth's surface and the atmosphere is thus crucial to understand how the coupled system will respond to and shape climatic changes. Phenology is controlled by meteorological conditions at different timescales: on the one hand, changes in key meteorological variables (temperature, water, radiation) can have immediate effects on the vegetation development; on the other hand, phenological changes can be driven by past environmental conditions, known as memory effects. However, the processes governing meteorological memory effects on phenology are not completely understood, resulting in their limited performance of vegetation phenology represented in land surface models. A deep learning model, specifically a long short-term memory network (LSTM), has the potential to capture and model the meteorological memory effects on vegetation phenology. Here, we apply the LSTM to model the vegetation phenology using meteorological drivers and high-temporal-resolution canopy greenness observations through digital repeat photography by the PhenoCam network. We compare a multiple linear regression model, a no-memory-effect LSTM model and a full-memory-effect LSTM model to predict the whole seasonal greenness trajectory and the corresponding phenological transition dates across 50 sites and 317 site years during 2009–2018, covering deciduous broadleaf forests, evergreen needleleaf forests and grasslands. Results show that the deep learning model outperforms the multiple linear regression model, and the full-memory-effect LSTM model performs better than the no-memory-effect model for all three plant function types (median R2 of 0.878, 0.957 and 0.955 for broadleaf forests, evergreen needleleaf forests and grasslands). We also find that the full-memory-effect LSTM model is capable of predicting the seasonal dynamic variations of canopy greenness and reproducing trends in shifting phenological transition dates. We also performed a sensitivity analysis of the full-memory-effect LSTM model to assess its plausibility, revealing its coherence with established knowledge of vegetation phenology sensitivity to meteorological conditions, particularly changes in temperature. Our study highlights that (1) multi-variate meteorological memory effects play a crucial role in vegetation phenology, and (2) deep learning opens up new avenues for improving the representation of vegetation phenological processes in land surface models via a hybrid modelling approach.

Geology↗