Search NASA⌕ Search

SEARCH · Search NASA

Results for “Regression Analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Risk assessment of wellbore leakage during underground hydrogen storage

The expansion of renewable energy sources would require large-scale energy storage options to overcome the intermittent nature of these sources. Underground hydrogen storage (UHS) in depleted hydrocarbon reservoirs offers a scalable and practical energy storage solution. These reservoirs are chosen for their availability and large capacity, but the unique properties of hydrogen raise concerns about potential leakage pathways, particularly through wellbores. In this study, we develop and apply, for the first time, reduced-order models (ROMs) specifically designed for efficient leakage risk prediction in UHS systems operating in depleted hydrocarbon reservoirs. Using 3,000 high-fidelity simulation scenarios, we examine the influence of 11 key parameters, including reservoir and aquifer depths, wellbore permeability and porosity, initial saturations of water, oil and gas fractions (hydrogen, light, intermediate, and heavy hydrocarbons), reservoir pressure multiplier, and the aquifer-to-reservoir volume ratio, to simulate leakage behavior over a 1,000-year timescale. We train ROMs using a two-step classification-regression approach, achieving R 2 values exceeding 99 % across all targets. These ROMs effectively capture the leakage evolution and identify critical controls of leakage, guiding the design of mitigation strategies. Results indicate that gas leakage occurs in about 27 % of scenarios as early as five years post-operation, reaching volumes of up to 106 ft3. Oil leakage is less frequent (~17 %) and typically begins decades later. Our findings also show that hydrogen often migrates first, owing to its smaller molecular size and higher buoyancy, followed by heavier hydrocarbons. Over time, these heavier components contribute significantly to the total leaked volume, reinforcing the need for targeted monitoring and remediation strategies. Our analysis highlights that deeper storage reservoirs, shallower aquifers, and low-permeability wellbores significantly reduce leakage risks. In conclusion, this work offers a robust framework for risk-informed UHS deployment, supporting energy security through reliable large-scale hydrogen storage while safeguarding environmental integrity.

08 HYDROGEN↗

Latent Stochastic Differential Equations for Modeling Quasar Variability and Inferring Black Hole Properties

Quasars are bright and unobscured active galactic nuclei (AGN) thought to be powered by the accretion of matter around supermassive black holes at the centers of galaxies. The temporal variability of a quasar’s brightness contains valuable information about its physical properties. The UV/optical variability is thought to be a stochastic process, often represented as a damped random walk described by a stochastic differential equation (SDE). Upcoming wide-field telescopes such as the Rubin Observatory Legacy Survey of Space and Time (LSST) are expected to observe tens of millions of AGN in multiple filters over a ten year period, so there is a need for efficient and automated modeling techniques that can handle the large volume of data. Latent SDEs are machine learning models well suited for modeling quasar variability, as they can explicitly capture the underlying stochastic dynamics. In this work, we adapt latent SDEs to jointly reconstruct multivariate quasar light curves and infer their physical properties such as the black hole mass, inclination angle, and temperature slope. Our model is trained on realistic simulations of LSST ten year quasar light curves, and we demonstrate its ability to reconstruct quasar light curves even in the presence of long seasonal gaps and irregular sampling across different bands, outperforming a multioutput Gaussian process regression baseline. Our method has the potential to provide a deeper understanding of the physical properties of quasars and is applicable to a wide range of other multivariate time series with missing data and irregular sampling.

79 ASTRONOMY AND ASTROPHYSICS↗

Taming nuclear mass models with Gaussian processes

We propose a new set of nuclear mass predictions based on multiple theoretical mass models. By employing Gaussian process regression with the Matérn kernel, we achieved root-mean-square (rms) deviations below 100 keV for the training dataset. The best-performing mass models achieved rms deviations below 150 keV for the new precise mass data from AME2020, whereas the ensemble average showed robust performance across the nuclear chart. Our approach uniquely combines: (1) systematic refinement of eight mass models through their residuals, (2) physics-informed features, including magic numbers, nucleon parity numbers, neutron excess, and nuclear collectivity, and (3) theory-to-theory validation demonstrating robust extrapolation capability. We find that the Matérn kernel provides superior uncertainty quantification compared to the RBF kernel, with a length-scale analysis revealing enhanced inter-nuclei correlations. We provide complete mass predictions for all unknown nuclides in AME2020, offering valuable constraints for nuclear structure studies and astrophysical modeling when used with proper uncertainty propagation.

Gaussian processes↗

Silver diamine fluoride differentially affects dentin and hypomineralized enamel permeabilities

OBJECTIVES: To investigate the physicochemical effect of silver diamine fluoride (SDF) by correlating permeability with mineral density and elemental composition of hypomineralized enamel and carious dentin. METHODS: Enamel and dentin from human carious primary teeth with and without SDF treatment in-vivo, and hypomineralized enamel from permanent molars with and without SDF treatment in-vitro were scanned using micro X-ray computed tomography. Spatial maps of biometals (calcium, zinc), phosphorus, and silver were generated using X-ray fluorescence microprobe. Permeabilities were computed using Porous Microstructure Analysis software. RESULTS: The intrinsic permeability of SDF-treated carious dentin was 14.3 % lower than untreated sound dentin (6.39e-15 ± 3.01e-15 m² vs 7.46e-15 ± 1.82e-15 m²; P < 0.0001), while untreated carious dentin was 98.4 % higher (1.48e-14 ± 7.11e-15 m²; P < 0.0001). SDF-treated and untreated transparent dentin showed similar reduced permeabilities (75.6 % and 78.4 % lower than untreated sound dentin, respectively; P = 0.93). Severely hypomineralized enamel showed permeability reaching 108.1 % of adjacent sound dentin (5.71e-15 ± 2.04e-15 m² vs 5.28e-15 ± 1.30e-15 m²; P = 0.1409) and was significantly higher than mildly hypomineralized enamel (1.39e-15 ± 1.04e-15 m²; P < 0.0001). SDF treatment did not significantly impact the permeability of severely hypomineralized enamel (12.4 % reduction; P = 0.07). Principal component regression identified Zn level as a significant effector of tissue permeabilities in carious primary teeth (P < 0.0001). SIGNIFICANCE: This study introduces a computational method to measure dental tissue permeability, and demonstrates that SDF significantly reduces permeability in carious dentin but not intact hypomineralized enamel. The study reveals biometal Zn localization can alter dentin and enamel permeabilities, providing new insights into pathobiological mechanisms underlying caries and hypomineralization.

Chou, Conrad↗

Effect of Environmental and Socioeconomic Factors on Increased Early Childhood Blood Lead Levels: A Case Study in Chicago

This study analyzes the prevalence of elevated blood lead levels (BLLs) in children across Chicagoland zip codes from 2019 to 2021, linking them to socioeconomic, environmental, and racial factors. Wilcoxon tests and generalized additive model (GAM) regressions identified economic hardship, reflected in per capita income and unemployment rates, as a significant contributor to increased lead poisoning (LP) rates. Additionally, LP rates correlate with the average age of buildings, particularly post the 1978 lead paint ban, illustrating policy impacts on health outcomes. The study further explores the novel area of land surface temperature (LST) effects on LP, finding that higher nighttime LST, indicative of urban heat island effects, correlates with increased LP. This finding gains additional significance in the context of anthropogenic climate change. When these factors are combined with the ongoing expansion of urban territories, a significant risk exists of escalating LP rates on a global scale. Racial disparity analysis revealed that Black and Hispanic/Latino populations face higher LP rates, primarily due to unemployment and older housing. The study underscores the necessity for targeted public health strategies to address these disparities, emphasizing the need for interventions that cater to the unique challenges of these at-risk communities.

Lee, Jangho (ORCID:0000000289421092)↗

Machine Learning Correlation of Electron Micrographs and ToF-SIMS for the Analysis of Organic Biomarkers in Mudstone

The spatial distribution of organics in geological samples can be used to determine when and how these organics were incorporated into the host rock. Mass spectrometry (MS) imaging can rapidly collect a large amount of data, but ions produced are mixed without discrimination, resulting in complex mass spectra that can be difficult to interpret. Here, we apply unsupervised and supervised machine learning (ML) to help interpret spectra from time-of-flight-secondary ion mass spectrometry (ToF-SIMS) of an organic-carbon-rich mudstone of the Middle Jurassic of England (UK). It was previously shown that the presence of sterane molecular biomarkers in this sample can be detected via ToF-SIMS (Pasterski, M. J. et al., Astrobiology 2023, 23, 936). We use unsupervised ML on scanning electron microscopy–electron dispersive spectroscopy (SEM-EDS) measurements to define compositional categories based on differences in elemental abundances. We then test the ability of four ML algorithms─k-nearest neighbors (KNN), recursive partitioning and regressive trees (RPART), eXtreme gradient boost (XGBoost), and random forest (RF)─to classify the ToF-SIM spectra using (1) the categories assigned via SEM-EDS, (2) organic and inorganic labels assigned via SEM-EDS, and (3) the presence or absence of detectable steranes in ToF-SIMS spectra. In terms of predictive accuracy and balanced accuracy, KNN was the best performing model and RPART the worst. The feature importance, or the specific features of the ToF-SIM spectra used by the models to make classifications, cannot be determined for KNN, preventing posthoc model interpretation. Nevertheless, the feature importance extracted from the other models was useful for interpreting spectra. In conclusion, we determined that some of the organic ions used to classify biomarker containing spectra may be fragment ions derived from kerogen which is abundant in this mudstone sample.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Investigating genomic prediction strategies for grain carotenoid traits in a tropical/subtropical maize panel

Abstract Vitamin A deficiency remains prevalent on a global scale, including in regions where maize constitutes a high percentage of human diets. One solution for alleviating this deficiency has been to increase grain concentrations of provitamin A carotenoids in maize (Zea mays ssp. mays L.)—an example of biofortification. The International Maize and Wheat Improvement Center (CIMMYT) developed a Carotenoid Association Mapping panel of 380 inbred lines adapted to tropical and subtropical environments that have varying grain concentrations of provitamin A and other health-beneficial carotenoids. Several major genes have been identified for these traits, 2 of which have particularly been leveraged in marker-assisted selection. This project assesses the predictive ability of several genomic prediction strategies for maize grain carotenoid traits within and between 4 environments in Mexico. Ridge Regression-Best Linear Unbiased Prediction, Elastic Net, and Reproducing Kernel Hilbert Spaces had high predictive abilities for all tested traits (β-carotene, β-cryptoxanthin, provitamin A, lutein, and zeaxanthin) and outperformed Least Absolute Shrinkage and Selection Operator. Furthermore, predictive abilities were higher when using genome-wide markers rather than only the markers proximal to 2 or 13 genes. These findings suggest that genomic prediction models using genome-wide markers (and assuming equal variance of marker effects) are worthwhile for these traits even though key genes have already been identified, especially if breeding for additional grain carotenoid traits alongside β-carotene. Predictive ability was maintained for all traits except lutein in between-environment prediction. The TASSEL (Trait Analysis by aSSociation, Evolution, and Linkage) Genomic Selection plugin performed as well as other more computationally intensive methods for within-environment prediction. The findings observed herein indicate the utility of genomic prediction methods for these traits and could inform their resource-efficient implementation in biofortification breeding programs.

59 BASIC BIOLOGICAL SCIENCES↗

Standardising the “Gregory method” for calculating equilibrium climate sensitivity

The equilibrium climate sensitivity (ECS) – the equilibrium global mean temperature response to a doubling of atmospheric CO 2 – is a high-profile metric for quantifying the Earth system's response to human-induced climate change. A widely applied approach to estimating the ECS is the “Gregory method” (Gregory et al., 2004), which uses an ordinary least squares (OLS) regression between the net radiative flux, N, and surface air temperature anomalies, ΔT, from a 150 year experiment in which atmospheric CO 2 concentrations are quadrupled. The ECS is determined by extrapolating the linear fit to N=0, i.e. the ΔT-intercept, indicating the point at which the system is back in equilibrium. This method has been used to compare ECS estimates across the CMIP5 and CMIP6 ensembles and will likely be a key diagnostic for CMIP7. Despite its widespread application, there is little consistency or transparency between studies in how the climate model data is processed prior to the regression, leading to potential discrepancies in ECS estimates. We identify 32 alternative data processing pathways, varying by differences in global mean weighting, net radiative flux variable, anomaly calculation method, and linear regression fit. Using 44 CMIP6 models, we systematically assess the impact of these choices on ECS estimates and calculate uncertainty ranges using two bootstrap approaches. While the inter-model ECS range is insensitive to the data processing pathway, individual outlier models exhibit notable differences. Approximating a model's native grid cell area (if irregular) with cosine of the latitude can decrease the ECS by 11 %, the choice of N-variable can change the ECS by 6 %, and some anomaly calculation methods can introduce spurious temporal correlations in the processed data. Beyond data processing choices, we also evaluate an alternative linear regression method – total least squares (TLS) – which has a more statistically robust basis than OLS. However, for consistency with previous literature, and given TLS may reduce the ECS compared to OLS (by up to 24 %), thereby making a known bias in the Gregory method worse, we do not feel there is sufficient clarity to recommend a transition to TLS in all cases. To improve reproducibility and comparability in future studies, we recommend a standardised Gregory method: weighting the global mean by cell area, using the top of the atmosphere (as opposed to the top of model) N-variable, and calculating anomalies by first applying a rolling average to the preindustrial control timeseries then subtracting from the raw CO 2 quadrupling experiment. This approach accounts for model drift while reducing noise in the data to best meet the pre-conditions of the linear regression. While CMIP6 results of the multi-model mean ECS appear insensitive to these processing choices, similar assumptions may not hold for CMIP7, underscoring the need for standardised data preparation in future climate sensitivity assessments.

Geosciences↗

Catalytic Reduction of Esters over Zirconia-Supported Metal Catalysts

Esters are often produced as unwanted byproducts during the catalytic upgrading of ethanol to diesel fuel precursors through Guerbet coupling. Removal of esters from the product stream is important to prevent the loss of downstream catalyst activity from ester-derived carboxylic acids. In this work, we studied ester hydrogenolysis to the parent alcohols as a viable route for enhanced diesel fuel production. Specifically, we investigated the reduction of hexyl acetate in butanol over ZrO 2 -supported Ni, Co, Cu, Rh, Pd, and Pt catalysts, where Cu/ZrO 2 was the most selective catalyst for the hydrogenolysis of hexyl acetate into hexanol and ethanol. Thermodynamic analysis reveals that a 90% alcohol yield can be obtained at 200 °C, 30 bar, and a relatively high H 2 :hexyl acetate molar ratio of 480:1. Experimentally, an alcohol yield of 88% yield was obtained with a 10 wt % Cu/ZrO 2 catalyst at these conditions with a residence time of 5.4 h kg cat kmol gas –1 . Catalytic tests on the support revealed that ZrO 2 catalyzes the transesterification reaction between hexyl acetate and butanol. However, only the Cu sites can catalyze the hydrogenolysis of the esters into the final alcohols. We developed a kinetic model for our experimental results, which shows that the transesterification and hydrogenolysis reactions run at two different timescales, the former being 10 times faster than the latter. Data regression has been used to develop a model to predict the mole fraction distribution of ester hydrogenolysis products over a wide range of contact times. Cu/ZrO 2 loses half its catalytic activity after 80 h of time on stream. Modeling of deactivation data reveals that the ZrO 2 support conserves a residual activity due to external active sites, while active sites over the Cu surface deactivate at different rates. Furthermore, the catalytic conversion of esters into their parent alcohols is relevant to the production of surrogate liquid fuels since alcohols can be bimolecularly dehydrated to produce a blend of ethers with diesel fuel-like properties.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Enhanced accuracy through ensembling of randomly initialized auto-regressive models for dynamical systems

Computational mechanics simulations using traditional finite element methods (FEM) require prohibitively expensive computational resources for real-time engineering applications, design optimization, and digital twin implementations. While machine learning (ML) surrogate models offer significant computational speedups, autoregressive ML models for time-dependent mechanical systems suffer from error accumulation that compromises long-term prediction reliability - a critical concern for engineering applications where accuracy over extended time horizons is essential for safety and performance assessments. Here, we propose a deep ensemble framework specifically designed to address this challenge in computational mechanics applications, where multiple ML surrogate models with random weight initializations are trained in parallel and their predictions aggregated during inference. This approach leverages statistical diversity to maximize information gain from a fixed set of training data and to mitigate error propagation, while maintaining the computational efficiency that makes ML surrogates attractive for engineering practice. We validate the framework on three representative problems spanning critical areas of computational mechanics: stress field evolution in heterogeneous microstructures under complex loading (relevant to advanced materials design and composite analysis), planetary-scale shallow water dynamics (applicable to environmental and geotechnical engineering), and Gray-Scott reaction-diffusion systems (relevant to mass transport and chemical process engineering). Across all test cases, the ensemble approach demonstrates consistent error reduction of 15-33% compared to individual models. The codes for this work are available on GitHub (https://github.com/Graham-Brady-Research-Group/AutoregressiveEnsemble_SpatioTemporal_Evolution).

autoregressive prediction↗

Information-incorporated gene network construction with FDR control

Abstract Motivation Large-scale gene expression studies allow gene network construction to uncover associations among genes. To study direct associations among genes, partial correlation-based networks are preferred over marginal correlations. However, FDR control for partial correlation-based network construction is not well-studied. In addition, currently available partial correlation-based methods cannot take existing biological knowledge to help network construction while controlling FDR. Results In this paper, we propose a method called Partial Correlation Graph with Information Incorporation (PCGII). PCGII estimates partial correlations between each pair of genes by regularized node-wise regression that can incorporate prior knowledge while controlling the effects of all other genes. It handles high-dimensional data where the number of genes can be much larger than the sample size and controls FDR at the same time. We compare PCGII with several existing approaches through extensive simulation studies and demonstrate that PCGII has better FDR control and higher power. We apply PCGII to a plant gene expression dataset where it recovers confirmed regulatory relationships and a hub node, as well as several direct associations that shed light on potential functional relationships in the system. We also introduce a method to supplement observed data with a pseudogene to apply PCGII when no prior information is available, which also allows checking FDR control and power for real data analysis. Availability and implementation R package is freely available for download at https://cran.r-project.org/package=PCGII.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative Analysis of HEATNETS for Geothermal Network Performance: Preprint

Thermal energy networks (TENs), also known as 5th generation district energy systems, or more specifically geothermal networks when exchanging heat with geothermal boreholes, are an important technology for decarbonization. In these networks an ambient loop connects buildings and thermal sources, such as a borehole field, to exchange energy and maintain a desired loop temperature. Water-source heat pumps are used at the buildings to connect to the ambient or thermal loop to meet to the building heating and cooling loads and maintain comfort. A semi-transient, reduced-order technical model and techno-economic model, called HEATNETS, has been developed at NREL that captures the flow of energy around a TEN. In this work, a comparison of the HEATNETS technical model and a well-known coding platform used for modeling geothermal networks, TRNSYS, has been completed for a proposed geothermal network as a verification process. Hourly data provided from the TRNSYS simulation included building loads, pumping power, heat pump power, temperature entering and leaving the borehole field, and mass flow rates. The hourly borehole temperatures were used to create a linear regression model utilized in HEATNETS to estimate the borehole field heat exchange. The building loads and mass flow rates were direct inputs to HEATNETS while the pumping power, heat pump power, borehole temperatures, and coefficients of performance were all simulated and calculated by HEATNETS, allowing for direct comparison of the thermal energy transfer, rather than also comparing control systems responses. HEATNETS considers the full process from design inputs to economic outputs and can provide modeling options for high-level initial system design and operational optimization. This study shows that HEATNETS, while not intended to replace other modeling tools, can be a unique modeling tool for the performance of a full geothermal network system.

15 GEOTHERMAL ENERGY↗

Comparative Analysis of HEATNETS for Geothermal Network Performance

Thermal energy networks (TENs), also known as 5th generation district energy systems, or more specifically geothermal networks when exchanging heat with geothermal boreholes, are an important technology for decarbonization. In these networks an ambient loop connects buildings and thermal sources, such as a borehole field, to exchange energy and maintain a desired loop temperature. Water-source heat pumps are used at the buildings to connect to the ambient or thermal loop to meet to the building heating and cooling loads and maintain comfort. A semi-transient, reduced-order technical model and techno-economic model, called HEATNETS, has been developed at NREL that captures the flow of energy around a TEN. In this work, a comparison of the HEATNETS technical model and a well-known coding platform used for modeling geothermal networks, TRNSYS, has been completed for a proposed geothermal network as a verification and validation process. Hourly data provided from the TRNSYS simulation included building loads, pumping power, heat pump power, temperature entering and leaving the borehole field, and mass flow rates. The hourly borehole temperatures were used to create a linear regression model utilized in HEATNETS to estimate the borehole field heat exchange. The building loads and mass flow rates were direct inputs to HEATNETS while the pumping power, heat pump power, borehole temperatures, and coefficients of performance were all simulated and calculated by HEATNETS, allowing for direct comparison of the thermal energy transfer HEATNETS considers the full process from design inputs to economic outputs and can provide modeling options for high-level initial system design and operational optimization. This study focuses on a validation of HEATNETS using results from TRNSYS. HEATNETS is not intended to replace other modeling tools, but this work demonstrates, via a comparison with an industry standard code, that HEATNETS can be a unique, high-level and rapid modeling tool for estimating the performance of a full geothermal network system.

15 GEOTHERMAL ENERGY↗

Constraining Galaxy-Halo connection using machine learning

We investigate the potential of machine learning (ML) methods to model small-scale galaxy clustering for constraining Halo Occupation Distribution (HOD) parameters. Our analysis reveals that while many ML algorithms report good statistical fits, they often yield likelihood contours that are significantly biased in both mean values and variances relative to the true model parameters. This highlights the importance of careful data processing and algorithm selection in ML applications for galaxy clustering, as even seemingly robust methods can lead to biased results if not applied correctly. ML tools offer a promising approach to exploring the HOD parameter space with significantly reduced computational costs compared to traditional brute-force methods if their robustness is established. Using our ANN-based pipeline, we successfully recreate some standard results from recent literature. Properly restricting the HOD parameter space, transforming the training data, and carefully selecting ML algorithms are essential for achieving unbiased and robust predictions. Among the methods tested, artificial neural networks (ANNs) outperform random forests (RF) and ridge regression in predicting clustering statistics, when the HOD prior space is appropriately restricted. We demonstrate these findings using the projected two-point correlation function (w p (r p )), angular multipoles of the correlation function (ξ ℓ (r)), and the void probability function (VPF) of Luminous Red Galaxies from Dark Energy Spectroscopic Instrument mocks. Our results show that while combining w p (r p ) and VPF improves parameter constraints, adding the multipoles ξ 0 , ξ 2 , and ξ 4 to w p (r p ) does not significantly improve the constraints.

cosmology↗

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES↗

Artificial neural networks estimate evapotranspiration for Miscanthus × giganteus as effectively as empirical model but with fewer inputs

Estimating actual evapotranspiration (ET) is particularly crucial for addressing how vegetation affects the water balance of ecosystems. ET estimation can be complex with empirical models due to their many parameters and reliance on aridity. In contrast, artificial neural networks (ANNs) could potentially estimate ET with fewer and more common meteorological parameters. In this study, we trained two ANNs, one using a feed-forward approach (FFN) and the other a nonlinear auto-regressive network (NARX), to predict ET and compared them to the commonly used empirical model Granger and Gray (GG). We trained our models on a nine-year eddy covariance (EC) dataset for Miscanthu s × giganteus ( M . × giganteus ) from Illinois (UIEF), then tested them using out-of-sample data from both UIEF and a different location in Iowa (SABR) to compare the accuracy of FFN, NARX, and GG models in estimating daily ET. A combination of air temperature (T a ) and solar radiation (R s ) was chosen as inputs due to the highest R 2 for FFN (R 2 = 0.79, 0.81, and 0.79 for training, testing, and validation, respectively) and only T a for NARX (R 2 = 0.70 for out-of-sample validation). The predictive power of the FFN model was superior to the NARX and GG models at the UIEF site (R 2 = 0.84, 0.70, and 0.83 for out-of-sample validation, respectively). Our analysis showed that ANN approaches are as accurate as empirical approaches for estimating ET but use fewer inputs.

54 ENVIRONMENTAL SCIENCES↗

Using Machine Learning to Predict Cloud Turbulent Entrainment–Mixing Processes

Different turbulent entrainment–mixing mechanisms between clouds and environment are essential to cloud–related processes; however, accurate representation of entrainment–mixing in weather/climate models still poses a challenge. This study exploits the use of machine learning (ML) to address this challenge. Four ML (Light Gradient Boosting Machine [LGB], eXtreme Gradient Boosting, Random Forest, and Support Vector Regression) are examined and compared. It is found that LGB performs best, and thus is selected to understand the impact of entrainment–mixing on microphysics using simulation data from Explicit Mixing Parcel Model. Compared with traditional parameterizations, the trained LGB provides more accurate microphysical properties (number concentration and cloud droplet spectral dispersion). The partial dependences of predicted microphysics on features exhibit a strong alignment with physical mechanisms and expectations, as determined by the interpreting method, thus overcoming the limitations of the “black box” scheme. The underlying mechanisms are that the smaller number concentration and larger spectral dispersion correspond to more inhomogeneous entrainment–mixing. Specifically, number concentration after entrainment–mixing is positively correlated with adiabatic number concentration and liquid water content affected by entrainment–mixing, and inversely correlated with adiabatic volume mean radius. Spectral dispersion after entrainment–mixing is negatively correlated with liquid water content affected by entrainment–mixing, turbulent dissipation rate and relative humidity of entrained air. Sensitivity analysis further suggests that number concentration is mainly determined by cloud microphysical properties whereas spectral dispersion is influenced by both cloud microphysical properties and environmental variables. The results indicate that the LGB scheme has the potential to enhance the representation of entrainment–mixing in weather/climate models.

54 ENVIRONMENTAL SCIENCES↗

Bayesian inference of structured latent spaces from neural population activity with the orthogonal stochastic linear mixing model

The brain produces diverse functions, from perceiving sounds to producing arm reaches, through the collective activity of populations of many neurons. Determining if and how the features of these exogenous variables (e.g., sound frequency, reach angle) are reflected in population neural activity is important for understanding how the brain operates. Often, high-dimensional neural population activity is confined to low-dimensional latent spaces. However, many current methods fail to extract latent spaces that are clearly structured by exogenous variables. This has contributed to a debate about whether or not brains should be thought of as dynamical systems or representational systems. Here, we developed a new latent process Bayesian regression framework, the orthogonal stochastic linear mixing model (OSLMM) which introduces an orthogonality constraint amongst time-varying mixture coefficients, and provide Markov chain Monte Carlo inference procedures. We demonstrate superior performance of OSLMM on latent trajectory recovery in synthetic experiments and show superior computational efficiency and prediction performance on several real-world benchmark data sets. We primarily focus on demonstrating the utility of OSLMM in two neural data sets: μ ECoG recordings from rat auditory cortex during presentation of pure tones and multi-single unit recordings form monkey motor cortex during complex arm reaching. We show that OSLMM achieves superior or comparable predictive accuracy of neural data and decoding of external variables (e.g., reach velocity). Most importantly, in both experimental contexts, we demonstrate that OSLMM latent trajectories directly reflect features of the sounds and reaches, demonstrating that neural dynamics are structured by neural representations. Together, these results demonstrate that OSLMM will be useful for the analysis of diverse, large-scale biological time-series datasets.

59 BASIC BIOLOGICAL SCIENCES↗