Search NASA⌕ Search

SEARCH · Search NASA

Results for “regression analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Uncertainty quantification and sensitivity analysis of a nuclear thermal propulsion reactor startup sequence

The research presented in this article describes progress in applying stochastic methods, uncertainty quantification, parametric studies, and variance-based sensitivity analysis (also known as Sobol sensitivity analysis) to a full-core model of a nuclear thermal propulsion (NTP) system simulated via the radiation transport code Griffin to simulate neutronics. Our goal is to develop a reduced-order (surrogate) model that can be rapidly sampled with perturbations to multiple input parameters. In this NTP system, reactivity and power feedback affect the rotation of control drums (CDs), which is itself controlled by a hybrid proportional-integral-derivative (PID) controller actuated by the power demand and reactivity feedback from the numerical model. This model uses reactor kinetic feedback (mean generation time [Λ] and effective delayed neutron fraction [ β eff ] from a transient Griffin simulation executed via Griffin’s improved quasi-static solver to provide the kinetic parameters) as inputs to functions that control the CD rotation angle. By investigating numerous stochastic approaches, we developed a dual-purpose surrogate model of the NTP system, using polynomial regression in the Multiphysics Object-Oriented Simulation Environment (MOOSE) Stochastic Tools Module (STM). The trained model can be rapidly sampled while simultaneously perturbing various input parameters, such as coefficients on the PID control or temperature (directly affecting the neutron cross section). The surrogate model delivers accurate (within 5%) results at speeds orders of magnitude faster (minutes, not days of computational time) than the base model. Once the surrogate model has been trained, distributions of the uncertain parameters can be changed at will to investigate the effects of perturbing multiple inputs as well as the effects of these inputs on the model output. For example, coefficients used in the PID control system may vary due to some type of physical interference, or uncertainty may exist in the temperature of the neutron cross sections in various regions of the reactor. A distribution can be placed on these parameters, and operational boundaries can be determined. The goal of this work is to support development of an advanced control system for operating CDs in a functioning NTP system. This work is a scoping study of the MOOSE STM.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN↗

Machine Learning Analysis of Temperature-Strain Relationships for Structural Health Monitoring of Pipes: Self-powered wireless sensor system for health monitoring of liquid-sodium cooled fast reactors

This report presents machine learning (ML) analysis of temperature-strain relationships for structural health monitoring of nuclear reactor stainless steel (SS) pipes with the strain gauge sensor directly printed on the pipe with a 3D conformal aerosol jet printer. We investigate correlations for two sensor pairs installed on the same SS304 pipe: commercial K-type thermocouple with a printed gold strain gauge (TC3-SG3), and commercial K-type thermocouple with commercial Kyowa strain gauge (TC0-SG0). The temperature ranges for the sensor pairs TC0-SG0 and TC3-SG3 are 20.00°C to 266.37°C and 39.95°C to 219.28°C respectively. ML algorithms in this study include Linear Regression (baseline method), Ridge Regression, Lasso Regression, and Gradient Boosting. Performance evaluation metrics include Root Mean Square Error (RMSE), Mean Square Error (MSE), Mean Absolute Error (MAE), R 2 Score, and Explained Variance. Using advanced feature engineering techniques, we extracted 27 temperature-based features and 30 strategic inclusion features. The best performance was obtained with the Gradient Boosting method, which achieves prediction accuracy of R 2 = 0.9999 and RMSE = 7.69 μStrain for TC0-SG0, and R 2 = 0.9998 and RMSE = 18.03 μStrain for TC3-SG3. While the temperature-strain correlations are weaker for the gauge directly printed on the pipe than for the commercial strain gauge, deployment-ready performance exceeding industry standards is achieved for both sensor pairs.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Risk assessment of wellbore leakage during underground hydrogen storage

The expansion of renewable energy sources would require large-scale energy storage options to overcome the intermittent nature of these sources. Underground hydrogen storage (UHS) in depleted hydrocarbon reservoirs offers a scalable and practical energy storage solution. These reservoirs are chosen for their availability and large capacity, but the unique properties of hydrogen raise concerns about potential leakage pathways, particularly through wellbores. In this study, we develop and apply, for the first time, reduced-order models (ROMs) specifically designed for efficient leakage risk prediction in UHS systems operating in depleted hydrocarbon reservoirs. Using 3,000 high-fidelity simulation scenarios, we examine the influence of 11 key parameters, including reservoir and aquifer depths, wellbore permeability and porosity, initial saturations of water, oil and gas fractions (hydrogen, light, intermediate, and heavy hydrocarbons), reservoir pressure multiplier, and the aquifer-to-reservoir volume ratio, to simulate leakage behavior over a 1,000-year timescale. We train ROMs using a two-step classification-regression approach, achieving R 2 values exceeding 99 % across all targets. These ROMs effectively capture the leakage evolution and identify critical controls of leakage, guiding the design of mitigation strategies. Results indicate that gas leakage occurs in about 27 % of scenarios as early as five years post-operation, reaching volumes of up to 106 ft3. Oil leakage is less frequent (~17 %) and typically begins decades later. Our findings also show that hydrogen often migrates first, owing to its smaller molecular size and higher buoyancy, followed by heavier hydrocarbons. Over time, these heavier components contribute significantly to the total leaked volume, reinforcing the need for targeted monitoring and remediation strategies. Our analysis highlights that deeper storage reservoirs, shallower aquifers, and low-permeability wellbores significantly reduce leakage risks. In conclusion, this work offers a robust framework for risk-informed UHS deployment, supporting energy security through reliable large-scale hydrogen storage while safeguarding environmental integrity.

08 HYDROGEN↗

Taming nuclear mass models with Gaussian processes

We propose a new set of nuclear mass predictions based on multiple theoretical mass models. By employing Gaussian process regression with the Matérn kernel, we achieved root-mean-square (rms) deviations below 100 keV for the training dataset. The best-performing mass models achieved rms deviations below 150 keV for the new precise mass data from AME2020, whereas the ensemble average showed robust performance across the nuclear chart. Our approach uniquely combines: (1) systematic refinement of eight mass models through their residuals, (2) physics-informed features, including magic numbers, nucleon parity numbers, neutron excess, and nuclear collectivity, and (3) theory-to-theory validation demonstrating robust extrapolation capability. We find that the Matérn kernel provides superior uncertainty quantification compared to the RBF kernel, with a length-scale analysis revealing enhanced inter-nuclei correlations. We provide complete mass predictions for all unknown nuclides in AME2020, offering valuable constraints for nuclear structure studies and astrophysical modeling when used with proper uncertainty propagation.

Gaussian processes↗

Silver diamine fluoride differentially affects dentin and hypomineralized enamel permeabilities

OBJECTIVES: To investigate the physicochemical effect of silver diamine fluoride (SDF) by correlating permeability with mineral density and elemental composition of hypomineralized enamel and carious dentin. METHODS: Enamel and dentin from human carious primary teeth with and without SDF treatment in-vivo, and hypomineralized enamel from permanent molars with and without SDF treatment in-vitro were scanned using micro X-ray computed tomography. Spatial maps of biometals (calcium, zinc), phosphorus, and silver were generated using X-ray fluorescence microprobe. Permeabilities were computed using Porous Microstructure Analysis software. RESULTS: The intrinsic permeability of SDF-treated carious dentin was 14.3 % lower than untreated sound dentin (6.39e-15 ± 3.01e-15 m² vs 7.46e-15 ± 1.82e-15 m²; P < 0.0001), while untreated carious dentin was 98.4 % higher (1.48e-14 ± 7.11e-15 m²; P < 0.0001). SDF-treated and untreated transparent dentin showed similar reduced permeabilities (75.6 % and 78.4 % lower than untreated sound dentin, respectively; P = 0.93). Severely hypomineralized enamel showed permeability reaching 108.1 % of adjacent sound dentin (5.71e-15 ± 2.04e-15 m² vs 5.28e-15 ± 1.30e-15 m²; P = 0.1409) and was significantly higher than mildly hypomineralized enamel (1.39e-15 ± 1.04e-15 m²; P < 0.0001). SDF treatment did not significantly impact the permeability of severely hypomineralized enamel (12.4 % reduction; P = 0.07). Principal component regression identified Zn level as a significant effector of tissue permeabilities in carious primary teeth (P < 0.0001). SIGNIFICANCE: This study introduces a computational method to measure dental tissue permeability, and demonstrates that SDF significantly reduces permeability in carious dentin but not intact hypomineralized enamel. The study reveals biometal Zn localization can alter dentin and enamel permeabilities, providing new insights into pathobiological mechanisms underlying caries and hypomineralization.

Chou, Conrad↗

Machine Learning Correlation of Electron Micrographs and ToF-SIMS for the Analysis of Organic Biomarkers in Mudstone

The spatial distribution of organics in geological samples can be used to determine when and how these organics were incorporated into the host rock. Mass spectrometry (MS) imaging can rapidly collect a large amount of data, but ions produced are mixed without discrimination, resulting in complex mass spectra that can be difficult to interpret. Here, we apply unsupervised and supervised machine learning (ML) to help interpret spectra from time-of-flight-secondary ion mass spectrometry (ToF-SIMS) of an organic-carbon-rich mudstone of the Middle Jurassic of England (UK). It was previously shown that the presence of sterane molecular biomarkers in this sample can be detected via ToF-SIMS (Pasterski, M. J. et al., Astrobiology 2023, 23, 936). We use unsupervised ML on scanning electron microscopy–electron dispersive spectroscopy (SEM-EDS) measurements to define compositional categories based on differences in elemental abundances. We then test the ability of four ML algorithms─k-nearest neighbors (KNN), recursive partitioning and regressive trees (RPART), eXtreme gradient boost (XGBoost), and random forest (RF)─to classify the ToF-SIM spectra using (1) the categories assigned via SEM-EDS, (2) organic and inorganic labels assigned via SEM-EDS, and (3) the presence or absence of detectable steranes in ToF-SIMS spectra. In terms of predictive accuracy and balanced accuracy, KNN was the best performing model and RPART the worst. The feature importance, or the specific features of the ToF-SIM spectra used by the models to make classifications, cannot be determined for KNN, preventing posthoc model interpretation. Nevertheless, the feature importance extracted from the other models was useful for interpreting spectra. In conclusion, we determined that some of the organic ions used to classify biomarker containing spectra may be fragment ions derived from kerogen which is abundant in this mudstone sample.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Standardising the “Gregory method” for calculating equilibrium climate sensitivity

The equilibrium climate sensitivity (ECS) – the equilibrium global mean temperature response to a doubling of atmospheric CO 2 – is a high-profile metric for quantifying the Earth system's response to human-induced climate change. A widely applied approach to estimating the ECS is the “Gregory method” (Gregory et al., 2004), which uses an ordinary least squares (OLS) regression between the net radiative flux, N, and surface air temperature anomalies, ΔT, from a 150 year experiment in which atmospheric CO 2 concentrations are quadrupled. The ECS is determined by extrapolating the linear fit to N=0, i.e. the ΔT-intercept, indicating the point at which the system is back in equilibrium. This method has been used to compare ECS estimates across the CMIP5 and CMIP6 ensembles and will likely be a key diagnostic for CMIP7. Despite its widespread application, there is little consistency or transparency between studies in how the climate model data is processed prior to the regression, leading to potential discrepancies in ECS estimates. We identify 32 alternative data processing pathways, varying by differences in global mean weighting, net radiative flux variable, anomaly calculation method, and linear regression fit. Using 44 CMIP6 models, we systematically assess the impact of these choices on ECS estimates and calculate uncertainty ranges using two bootstrap approaches. While the inter-model ECS range is insensitive to the data processing pathway, individual outlier models exhibit notable differences. Approximating a model's native grid cell area (if irregular) with cosine of the latitude can decrease the ECS by 11 %, the choice of N-variable can change the ECS by 6 %, and some anomaly calculation methods can introduce spurious temporal correlations in the processed data. Beyond data processing choices, we also evaluate an alternative linear regression method – total least squares (TLS) – which has a more statistically robust basis than OLS. However, for consistency with previous literature, and given TLS may reduce the ECS compared to OLS (by up to 24 %), thereby making a known bias in the Gregory method worse, we do not feel there is sufficient clarity to recommend a transition to TLS in all cases. To improve reproducibility and comparability in future studies, we recommend a standardised Gregory method: weighting the global mean by cell area, using the top of the atmosphere (as opposed to the top of model) N-variable, and calculating anomalies by first applying a rolling average to the preindustrial control timeseries then subtracting from the raw CO 2 quadrupling experiment. This approach accounts for model drift while reducing noise in the data to best meet the pre-conditions of the linear regression. While CMIP6 results of the multi-model mean ECS appear insensitive to these processing choices, similar assumptions may not hold for CMIP7, underscoring the need for standardised data preparation in future climate sensitivity assessments.

Geosciences↗

Catalytic Reduction of Esters over Zirconia-Supported Metal Catalysts

Esters are often produced as unwanted byproducts during the catalytic upgrading of ethanol to diesel fuel precursors through Guerbet coupling. Removal of esters from the product stream is important to prevent the loss of downstream catalyst activity from ester-derived carboxylic acids. In this work, we studied ester hydrogenolysis to the parent alcohols as a viable route for enhanced diesel fuel production. Specifically, we investigated the reduction of hexyl acetate in butanol over ZrO 2 -supported Ni, Co, Cu, Rh, Pd, and Pt catalysts, where Cu/ZrO 2 was the most selective catalyst for the hydrogenolysis of hexyl acetate into hexanol and ethanol. Thermodynamic analysis reveals that a 90% alcohol yield can be obtained at 200 °C, 30 bar, and a relatively high H 2 :hexyl acetate molar ratio of 480:1. Experimentally, an alcohol yield of 88% yield was obtained with a 10 wt % Cu/ZrO 2 catalyst at these conditions with a residence time of 5.4 h kg cat kmol gas –1 . Catalytic tests on the support revealed that ZrO 2 catalyzes the transesterification reaction between hexyl acetate and butanol. However, only the Cu sites can catalyze the hydrogenolysis of the esters into the final alcohols. We developed a kinetic model for our experimental results, which shows that the transesterification and hydrogenolysis reactions run at two different timescales, the former being 10 times faster than the latter. Data regression has been used to develop a model to predict the mole fraction distribution of ester hydrogenolysis products over a wide range of contact times. Cu/ZrO 2 loses half its catalytic activity after 80 h of time on stream. Modeling of deactivation data reveals that the ZrO 2 support conserves a residual activity due to external active sites, while active sites over the Cu surface deactivate at different rates. Furthermore, the catalytic conversion of esters into their parent alcohols is relevant to the production of surrogate liquid fuels since alcohols can be bimolecularly dehydrated to produce a blend of ethers with diesel fuel-like properties.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Enhanced accuracy through ensembling of randomly initialized auto-regressive models for dynamical systems

Computational mechanics simulations using traditional finite element methods (FEM) require prohibitively expensive computational resources for real-time engineering applications, design optimization, and digital twin implementations. While machine learning (ML) surrogate models offer significant computational speedups, autoregressive ML models for time-dependent mechanical systems suffer from error accumulation that compromises long-term prediction reliability - a critical concern for engineering applications where accuracy over extended time horizons is essential for safety and performance assessments. Here, we propose a deep ensemble framework specifically designed to address this challenge in computational mechanics applications, where multiple ML surrogate models with random weight initializations are trained in parallel and their predictions aggregated during inference. This approach leverages statistical diversity to maximize information gain from a fixed set of training data and to mitigate error propagation, while maintaining the computational efficiency that makes ML surrogates attractive for engineering practice. We validate the framework on three representative problems spanning critical areas of computational mechanics: stress field evolution in heterogeneous microstructures under complex loading (relevant to advanced materials design and composite analysis), planetary-scale shallow water dynamics (applicable to environmental and geotechnical engineering), and Gray-Scott reaction-diffusion systems (relevant to mass transport and chemical process engineering). Across all test cases, the ensemble approach demonstrates consistent error reduction of 15-33% compared to individual models. The codes for this work are available on GitHub (https://github.com/Graham-Brady-Research-Group/AutoregressiveEnsemble_SpatioTemporal_Evolution).

autoregressive prediction↗

Comparative Analysis of HEATNETS for Geothermal Network Performance: Preprint

Thermal energy networks (TENs), also known as 5th generation district energy systems, or more specifically geothermal networks when exchanging heat with geothermal boreholes, are an important technology for decarbonization. In these networks an ambient loop connects buildings and thermal sources, such as a borehole field, to exchange energy and maintain a desired loop temperature. Water-source heat pumps are used at the buildings to connect to the ambient or thermal loop to meet to the building heating and cooling loads and maintain comfort. A semi-transient, reduced-order technical model and techno-economic model, called HEATNETS, has been developed at NREL that captures the flow of energy around a TEN. In this work, a comparison of the HEATNETS technical model and a well-known coding platform used for modeling geothermal networks, TRNSYS, has been completed for a proposed geothermal network as a verification process. Hourly data provided from the TRNSYS simulation included building loads, pumping power, heat pump power, temperature entering and leaving the borehole field, and mass flow rates. The hourly borehole temperatures were used to create a linear regression model utilized in HEATNETS to estimate the borehole field heat exchange. The building loads and mass flow rates were direct inputs to HEATNETS while the pumping power, heat pump power, borehole temperatures, and coefficients of performance were all simulated and calculated by HEATNETS, allowing for direct comparison of the thermal energy transfer, rather than also comparing control systems responses. HEATNETS considers the full process from design inputs to economic outputs and can provide modeling options for high-level initial system design and operational optimization. This study shows that HEATNETS, while not intended to replace other modeling tools, can be a unique modeling tool for the performance of a full geothermal network system.

15 GEOTHERMAL ENERGY↗

Comparative Analysis of HEATNETS for Geothermal Network Performance

Thermal energy networks (TENs), also known as 5th generation district energy systems, or more specifically geothermal networks when exchanging heat with geothermal boreholes, are an important technology for decarbonization. In these networks an ambient loop connects buildings and thermal sources, such as a borehole field, to exchange energy and maintain a desired loop temperature. Water-source heat pumps are used at the buildings to connect to the ambient or thermal loop to meet to the building heating and cooling loads and maintain comfort. A semi-transient, reduced-order technical model and techno-economic model, called HEATNETS, has been developed at NREL that captures the flow of energy around a TEN. In this work, a comparison of the HEATNETS technical model and a well-known coding platform used for modeling geothermal networks, TRNSYS, has been completed for a proposed geothermal network as a verification and validation process. Hourly data provided from the TRNSYS simulation included building loads, pumping power, heat pump power, temperature entering and leaving the borehole field, and mass flow rates. The hourly borehole temperatures were used to create a linear regression model utilized in HEATNETS to estimate the borehole field heat exchange. The building loads and mass flow rates were direct inputs to HEATNETS while the pumping power, heat pump power, borehole temperatures, and coefficients of performance were all simulated and calculated by HEATNETS, allowing for direct comparison of the thermal energy transfer HEATNETS considers the full process from design inputs to economic outputs and can provide modeling options for high-level initial system design and operational optimization. This study focuses on a validation of HEATNETS using results from TRNSYS. HEATNETS is not intended to replace other modeling tools, but this work demonstrates, via a comparison with an industry standard code, that HEATNETS can be a unique, high-level and rapid modeling tool for estimating the performance of a full geothermal network system.

15 GEOTHERMAL ENERGY↗

Constraining Galaxy-Halo connection using machine learning

We investigate the potential of machine learning (ML) methods to model small-scale galaxy clustering for constraining Halo Occupation Distribution (HOD) parameters. Our analysis reveals that while many ML algorithms report good statistical fits, they often yield likelihood contours that are significantly biased in both mean values and variances relative to the true model parameters. This highlights the importance of careful data processing and algorithm selection in ML applications for galaxy clustering, as even seemingly robust methods can lead to biased results if not applied correctly. ML tools offer a promising approach to exploring the HOD parameter space with significantly reduced computational costs compared to traditional brute-force methods if their robustness is established. Using our ANN-based pipeline, we successfully recreate some standard results from recent literature. Properly restricting the HOD parameter space, transforming the training data, and carefully selecting ML algorithms are essential for achieving unbiased and robust predictions. Among the methods tested, artificial neural networks (ANNs) outperform random forests (RF) and ridge regression in predicting clustering statistics, when the HOD prior space is appropriately restricted. We demonstrate these findings using the projected two-point correlation function (w p (r p )), angular multipoles of the correlation function (ξ ℓ (r)), and the void probability function (VPF) of Luminous Red Galaxies from Dark Energy Spectroscopic Instrument mocks. Our results show that while combining w p (r p ) and VPF improves parameter constraints, adding the multipoles ξ 0 , ξ 2 , and ξ 4 to w p (r p ) does not significantly improve the constraints.

cosmology↗

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES↗

Artificial neural networks estimate evapotranspiration for Miscanthus × giganteus as effectively as empirical model but with fewer inputs

Estimating actual evapotranspiration (ET) is particularly crucial for addressing how vegetation affects the water balance of ecosystems. ET estimation can be complex with empirical models due to their many parameters and reliance on aridity. In contrast, artificial neural networks (ANNs) could potentially estimate ET with fewer and more common meteorological parameters. In this study, we trained two ANNs, one using a feed-forward approach (FFN) and the other a nonlinear auto-regressive network (NARX), to predict ET and compared them to the commonly used empirical model Granger and Gray (GG). We trained our models on a nine-year eddy covariance (EC) dataset for Miscanthu s × giganteus ( M . × giganteus ) from Illinois (UIEF), then tested them using out-of-sample data from both UIEF and a different location in Iowa (SABR) to compare the accuracy of FFN, NARX, and GG models in estimating daily ET. A combination of air temperature (T a ) and solar radiation (R s ) was chosen as inputs due to the highest R 2 for FFN (R 2 = 0.79, 0.81, and 0.79 for training, testing, and validation, respectively) and only T a for NARX (R 2 = 0.70 for out-of-sample validation). The predictive power of the FFN model was superior to the NARX and GG models at the UIEF site (R 2 = 0.84, 0.70, and 0.83 for out-of-sample validation, respectively). Our analysis showed that ANN approaches are as accurate as empirical approaches for estimating ET but use fewer inputs.

54 ENVIRONMENTAL SCIENCES↗

Using Machine Learning to Predict Cloud Turbulent Entrainment–Mixing Processes

Different turbulent entrainment–mixing mechanisms between clouds and environment are essential to cloud–related processes; however, accurate representation of entrainment–mixing in weather/climate models still poses a challenge. This study exploits the use of machine learning (ML) to address this challenge. Four ML (Light Gradient Boosting Machine [LGB], eXtreme Gradient Boosting, Random Forest, and Support Vector Regression) are examined and compared. It is found that LGB performs best, and thus is selected to understand the impact of entrainment–mixing on microphysics using simulation data from Explicit Mixing Parcel Model. Compared with traditional parameterizations, the trained LGB provides more accurate microphysical properties (number concentration and cloud droplet spectral dispersion). The partial dependences of predicted microphysics on features exhibit a strong alignment with physical mechanisms and expectations, as determined by the interpreting method, thus overcoming the limitations of the “black box” scheme. The underlying mechanisms are that the smaller number concentration and larger spectral dispersion correspond to more inhomogeneous entrainment–mixing. Specifically, number concentration after entrainment–mixing is positively correlated with adiabatic number concentration and liquid water content affected by entrainment–mixing, and inversely correlated with adiabatic volume mean radius. Spectral dispersion after entrainment–mixing is negatively correlated with liquid water content affected by entrainment–mixing, turbulent dissipation rate and relative humidity of entrained air. Sensitivity analysis further suggests that number concentration is mainly determined by cloud microphysical properties whereas spectral dispersion is influenced by both cloud microphysical properties and environmental variables. The results indicate that the LGB scheme has the potential to enhance the representation of entrainment–mixing in weather/climate models.

54 ENVIRONMENTAL SCIENCES↗

Bayesian inference of structured latent spaces from neural population activity with the orthogonal stochastic linear mixing model

The brain produces diverse functions, from perceiving sounds to producing arm reaches, through the collective activity of populations of many neurons. Determining if and how the features of these exogenous variables (e.g., sound frequency, reach angle) are reflected in population neural activity is important for understanding how the brain operates. Often, high-dimensional neural population activity is confined to low-dimensional latent spaces. However, many current methods fail to extract latent spaces that are clearly structured by exogenous variables. This has contributed to a debate about whether or not brains should be thought of as dynamical systems or representational systems. Here, we developed a new latent process Bayesian regression framework, the orthogonal stochastic linear mixing model (OSLMM) which introduces an orthogonality constraint amongst time-varying mixture coefficients, and provide Markov chain Monte Carlo inference procedures. We demonstrate superior performance of OSLMM on latent trajectory recovery in synthetic experiments and show superior computational efficiency and prediction performance on several real-world benchmark data sets. We primarily focus on demonstrating the utility of OSLMM in two neural data sets: μ ECoG recordings from rat auditory cortex during presentation of pure tones and multi-single unit recordings form monkey motor cortex during complex arm reaching. We show that OSLMM achieves superior or comparable predictive accuracy of neural data and decoding of external variables (e.g., reach velocity). Most importantly, in both experimental contexts, we demonstrate that OSLMM latent trajectories directly reflect features of the sounds and reaches, demonstrating that neural dynamics are structured by neural representations. Together, these results demonstrate that OSLMM will be useful for the analysis of diverse, large-scale biological time-series datasets.

59 BASIC BIOLOGICAL SCIENCES↗

Unlocking hidden information in sparse small-angle neutron scattering measurements

Hypothesis Small-Angle Neutron Scattering (SANS) is a powerful technique for studying soft matter systems such as colloids, polymers, and lyotropic phases, providing nanoscale structural insights. However, its effectiveness is limited by low neutron flux, leading to long acquisition times and noisy data. Here, we hypothesize that Bayesian statistical inference using Gaussian Process Regression (GPR) can reconstruct high-fidelity scattering data from sparse measurements by leveraging intensity smoothness and continuity. Experiments and Simulations The method was benchmarked computationally and validated through SANS experiments on various soft matter systems, including wormlike micelles, colloidal suspensions, polymeric structures, and lyotropic phases. GPR-based inference was applied to both experimental and synthetic data to evaluate its effectiveness in noise reduction and intensity reconstruction. Findings GPR significantly enhances SANS data quality and therefore reducing measurement times by up to two orders of magnitude. This cost-effective approach maximizes experimental efficiency, enabling high-throughput studies and real-time monitoring of dynamic systems. It is particularly beneficial for weakly scattering and time-sensitive studies. Beyond SANS, this framework applies to other low-SNR techniques, including laboratory-based small-angle X-ray scattering and various dynamical scattering methods. Furthermore, it offers transformative potential for compact neutron sources, enhancing their viability for structural analysis in resource-limited settings.

Small angle neutron scattering↗

Geographical Insights into Suicide Mortality Through Spatial Machine Learning

Suicide mortality is a leading cause of death in the United States, with an upward trend that emphasizes its significance as a public health issue. Previous research has employed global models like ordinary least squares (OLS) regression and local models such as geographically weighted regression (GWR). While local models are useful for analyzing spatial variations in suicide mortality, they share limitations with traditional global models, particularly about their inability to handle multi-collinearity and non-linear relationships. Machine learning approaches, like random forests (RF), can address some of these limitations but often fail to account for spatial variability. This gap highlights the need for spatial ML models specifically designed to tackle suicide mortality. This research seeks to fill this void by using a geographically weighted random forest model (GWRF) to examine the associations between county-level suicide mortality in the U.S. from 2010 to 2020 and various social and environmental determinants of health. A key aspect of our methodology is disciplined feature selection, which reduces the pool of explanatory variables by about 90%. This refinement enhances the explanatory power of both global (R2 improved from 0.59 to 0.67) and local (R2 improved from 0.64 to 0.67) RF models while reducing their run times. An analysis of the importance scores for these selected features reveals that the drivers of suicide mortality vary by context. Thus, to effectively address regional disparities and inform targeted public health interventions, a holistic approach that incorporates multiple county-level characteristics is essential.

Lebakula, Viswadeep [ORNL] (ORCID:0000000152935914↗