Search NASA⌕ Search

SEARCH · Search NASA

Results for “regression models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

NASA's Functional Task Test: High Intensity Exercise Improves the Heart Rate Response to a Stand Test Following 70 Days of Bedrest

Cardiovascular adaptations due to spaceflight are modeled with 6deg head-down tilt bed rest (BR) and result in decreased orthostatic tolerance. We investigated if high-intensity resistive and aerobic exercise with and without testosterone supplementation would improve the heart rate (HR) response to a 3.5-min stand test and how quickly these changes recovered following BR. During 70 days of BR male subjects performed no exercise (Control, n=10), high intensity supine resistive and aerobic exercise (Exercise, n=9), or supine exercise plus supplemental testosterone (Exercise+T, n=8; 100 mg i.m., weekly in 2-week on/off cycles). We measured HR for 2 min while subjects were prone and for 3 min after standing twice before and 0, 1, 6, and 11 days after BR. Mixed-effects linear regression models were used to evaluate group, time, and interaction effects. Compared to pre-bed rest, prone HR was elevated on BR+0 and BR+1 in Control, but not Exercise or Exercise+T groups, and standing HR was greater in all 3 groups. The increase in prone and standing HR in Control subjects was greater than either Exercise or Exercise+T groups and all groups recovered by BR+6. The change in HR from prone to standing more than doubled on BR+0 in all groups, but was significantly less in the Exericse+T group compared to the Control, but not Exercise group. Exercise reduces, but does not prevent the increase in HR observed in response to standing. The significantly lower HR response in the Exercise+T group requires further investigation to determine physiologic significance.

Laurie, Steven S.↗

Three-Dimensional Aeroelastic and Aerothermoelastic Behavior in Hypersonic Flow

The aeroelastic and aerothermoelastic behavior of three-dimensional configurations in hypersonic flow regime are studied. The aeroelastic behavior of a low aspect ratio wing, representative of a fin or control surface on a generic hypersonic vehicle, is examined using third order piston theory, Euler and Navier-Stokes aerodynamics. The sensitivity of the aeroelastic behavior generated using Euler and Navier-Stokes aerodynamics to parameters governing temporal accuracy is also examined. Also, a refined aerothermoelastic model, which incorporates the heat transfer between the fluid and structure using CFD generated aerodynamic heating, is used to examine the aerothermoelastic behavior of the low aspect ratio wing in the hypersonic regime. Finally, the hypersonic aeroelastic behavior of a generic hypersonic vehicle with a lifting-body type fuselage and canted fins is studied using piston theory and Euler aerodynamics for the range of 2.5 less than or equal to M less than or equal to 28, at altitudes ranging from 10,000 feet to 80,000 feet. This analysis includes a study on optimal mesh selection for use with Euler aerodynamics. In addition to the aeroelastic and aerothermoelastic results presented, three time domain flutter identification techniques are compared, namely the moving block approach, the least squares curve fitting method, and a system identification technique using an Auto-Regressive model of the aeroelastic system. In general, the three methods agree well. The system identification technique, however, provided quick damping and frequency estimations with minimal response record length, and therefore o ers significant reductions in computational cost. In the present case, the computational cost was reduced by 75%. The aeroelastic and aerothermoelastic results presented illustrate the applicability of the CFL3D code for the hypersonic flight regime.

McNamara, Jack J.↗

Assessing the Transferability of Statistical Predictive Models for Leaf Area Index Between Two Airborne Discrete Return LiDAR Sensor Designs Within Multiple Intensely Managed Loblolly Pine Forest Locations in the South-Eastern USA

Leaf area is an important forest structural variable which serves as the primary means of mass and energy exchange within vegetated ecosystems. The objective of the current study was to determine if leaf area index (LAI) could be estimated accurately and consistently in five intensively managed pine plantation forests using two multiple-return airborne LiDAR datasets. Field measurements of LAI were made using the LiCOR LAI2000 and LAI2200 instruments within 116 plots were established of varying size and within a variety of stand conditions (i.e. stand age, nutrient regime and stem density) in North Carolina and Virginia in 2008 and 2013. A number of common LiDAR return height and intensity distribution metrics were calculated (e.g. average return height), in addition to ten indices, with two additional variants, utilized in the surrounding literature which have been used to estimate LAI and fractional cover, were calculated from return heights and intensity, for each plot extent. Each of the indices was assessed for correlation with each other, and was used as independent variables in linear regression analysis with field LAI as the dependent variable. All LiDAR derived metrics were also entered into a forward stepwise linear regression. The results from each of the indices varied from an R2 of 0.33 (S.E. 0.87) to 0.89 (S.E. 0.36). Those indices calculated using ratios of all returns produced the strongest correlations, such as the Above and Below Ratio Index (ABRI) and Laser Penetration Index 1 (LPI1). The regression model produced from a combination of three metrics did not improve correlations greatly (R2 0.90; S.E. 0.35). The results indicate that LAI can be predicted over a range of intensively managed pine plantation forest environments accurately when using different LiDAR sensor designs. Those indices which incorporated counts of specific return numbers (e.g. first returns) or return intensity correlated poorly with field measurements. There were disparities between the number of different types of returns and intensity values when comparing the results from two LiDAR sensors, indicating that predictive models developed using such metrics are not transferable between datasets with different acquisition parameters. Each of the indices were significantly correlated with one another, with one exception (LAI proxy), in particular those indices calculated from all returns, which indicates similarities in information content for those indices. It can then be argued that LiDAR indices have reached a similar stage in development to those calculated from optical-spectral sensors, but which offer a number of advantages, such as the reduction or removal of saturation issues in areas of high biomass.

Sumnall, Matthew↗

Using leaf and stomatal traits to predict biomass production and water use efficiency in Populus

Climate change is reshaping ecosystems, driving plants to adapt through leaf-trait plasticity that reflects strategies for growth and water use. Predicting biomass production and intrinsic water use efficiency (iWUE) remains challenging because of genetic, taxonomic, and environmental variability. Here, we used eastern cottonwood and Populus hybrids as a model system to test whether easily measurable leaf traits can serve as reliable predictors of performance, and whether adding stomatal and biochemical traits improves predictive power. Across two field sites in Mississippi, leaf mass per area (LMA), biomass production, iWUE, leaf area, and foliar nitrogen ( N %) differed significantly among taxa and sites, while other traits were conserved. Factorial analysis of mixed data (FAMD) revealed distinct clustering of taxa and sites, indicating coordinated variation among leaf and stomatal traits. Pairwise correlations highlighted fundamental trade-offs, with biomass positively related to LMA and petiole length but negatively associated with iWUE, N %, and carbon isotopic ratios (δ 13 C). Leaf temperature and leaf angle varied among taxa and were significantly correlated with LMA and petiole length, suggesting mechanisms of heat dissipation and leaf movability that link simple traits to gas exchange and productivity. Weighted multiple linear regression models explained 80%–91% of variation in biomass production and iWUE. Models using only LMA, petiole length, and stomatal metrics performed nearly as well as those incorporating N %, and δ 13 C, with complex traits adding approximately 10% explanatory power. These results demonstrate that simple morphological traits capture integrated functional trade-offs, while complex traits refine predictions. This tiered approach provides an efficient framework for selecting high-yielding, water-efficient genotypes of Populus and other hardwood species, offering practical pathways to enhance carbon uptake and iWUE under climate change.

biomass production↗

Response of Surface Shortwave Cloud Radiative Effect to Greenhouse Gases and Aerosols and Its Impact on Summer Maximum Temperature

Shortwave cloud radiative effects (SWCREs), defined as the difference of the shortwave radiative flux between all-sky and clear-sky conditions at the surface, have been reported to play an important role in influencing the Earth's energy budget and temperature extremes. In this study, we employed a set of global climate models to examine the SWCRE responses to CO2, black carbon (BC) aerosols, and sulfate aerosols in boreal summer over the Northern Hemisphere. We found that CO2 causes positive SWCRE changes over most of the NH, and BC causes similar positive responses over North America, Europe, and eastern China but negative SWCRE over India and tropical Africa. When normalized by effective radiative forcing, the SWCRE from BC is roughly 3–5 times larger than that from CO2. SWCRE change is mainly due to cloud cover changes resulting from changes in relative humidity (RH) and, to a lesser extent, changes in cloud liquid water, circulation, dynamics, and stability. The SWCRE response to sulfate aerosols, however, is negligible compared to that for CO2 and BC because part of the radiation scattered by clouds under all-sky conditions will also be scattered by aerosols under clear-sky conditions. Using a multilinear regression model, it is found that mean daily maximum temperature (Tmax) increases by 0.15 and 0.13 K per watt per square meter (W m−2) increase in local SWCRE under the CO2 and BC experiment, respectively. When domain-averaged, the contribution of SWCRE change to summer mean Tmax changes was 10 %–30 % under CO2 forcing and 30 %–50 % under BC forcing, varying by region, which can have important implications for extreme climatic events and socioeconomic activities.

surface shortwave cloud radiative effect↗

Machine-learning and first-principles investigation of lightweight medium-entropy alloys for hydrogen-storage applications

The transition to a low-carbon economy demands efficient and sustainable energy-storage solutions, with hydrogen emerging as a promising clean-energy carrier and with metal hydrides recognized for their hydrogen-storage capacity. Here, we leverage machine learning (ML) to predict hydrogen-to-metal (H/M) ratios and solution energy by incorporating thermodynamic parameters and local lattice distortion (LLD) as key features. Our best-performing ML model provides improvements to H/M ratios and solution energies over a broad class of medium-entripy alloys (easily extendable to multi-principal-element alloys), such as Ti–Nb-X (X = Mo, Cr, Hf, Ta, V, Zr) and Co–Ni-X (X = Al, Mg, V). Ti–Nb–Mo alloys reveal compositional effects in H-storage behavior, in particular Ti, Nb, and V enhance H-storage capacity, while Mo reduces H/M and hydrogen weight percent by 40–50 %. We attributed results in molybdenum-rich alloys to slow hydrogen kinetics, as validated by our pressure-composition-temperature (PCT) isotherm experiments on pure Ti and Ti 5 Mo 95 alloys. Density functional theory (DFT) and molecular dynamics (MD) simulations also confirm that Ti and Nb promote H diffusion, whereas Mo hinders it, highlighting the interplay between electronic structure, lattice distortions, and hydrogen uptake. Notably, our Gradient Boosting Regression model identifies LLD as a critical factor in H/M predictions. Here, to aid material selection, we present two periodic tables illustrating elemental effects on (a) H 2 wt% and (b) solution energy, derived from ML, and provide a reference for identifying alloying elements that enhance hydrogen solubility and storage.

08 HYDROGEN↗

Calibration and Data Analysis Recommendations for Three-Component Moment Balances

Fundamental characteristics of design, calibration, and application of three-component moment balances are investigated in great detail. These balances are typically used to determine loads on control surfaces, canards, or other parts that are attached to a wind tunnel model. First, three different descriptions of the load state of a moment balance are reviewed. Then, load transformations between different load formats and the combined load diagram for two of the three load components are discussed. An error analysis showed that it is critical to maximize the product of the distance between the bending moment gages and their sensitivities in order to minimize the overall error in the normal force prediction. In addition, it is important to apply a sufficient number of calibration loadings near the first bending moment gage. Then, unwanted near-linear dependencies between the two bending moment gage outputs can be avoided. The error in the bending moment prediction is also investigated that results from the elastic deformation of the metric part of the balance under load. Finally, the application of the Non-Iterative Method to three-component moment balance calibration data is described in order to obtain regression models that can be used to predict loads from measured outputs during a wind tunnel test.

Strain-Gage Balance↗

Continental shelf fish production estimation from CZCS chlorophyll data

A method for ocean fish production estimation was proposed for development. The method was to use data acquired with the Coastal Zone Color Scanner, and processed into chlorophyll concentrations by the GSFC ocean Sciences Division, in combination with fish production and primary production data acquired from different ocean areas. A linear relation exits between annual fish production and annual phytoplankton carbon production for a wide range of coastal ocean environments. The uses of several existing algorithms which relate primary production to CZCS chlorophyll data as input to the fish production regression model is proposed. A question relating phytoplankton production to CZCS chlorophyll was obtained by Eppley (1984) using chlorophyll data obtained from field samples, equivalent to chlorophyll data obtained from CZCS imagery, and primary production data obtained from ship-board observations on a wide variety of coastal and open ocean environments. This equation was modified with additional data and was successfully tested using CZCS data and field chlorophyll and phytoplankton production data obtained from northeastern North American continental shelf waters and Atlantic open ocean waters. The modified Eppley (1984) relation also estimated phytoplankton annual carbon production in the Sargasso Sea within the confidence limits of a mean value obtained from the Eppley (1984) equation for oceanic waters that provide about 90 percent of total ocean primary production. The modified Eppley production formula applied to CZCS chlorophyll data obtained from several northeastern North American coastal environments gave phytoplankton annual carbon production values similar to the values used in the fish production regression equation.

Iverson, Richard L.↗

Radiometric Modeling and Calibration of the Geostationary Imaging Fourier Transform Spectrometer (GIFTS)Ground Based Measurement Experiment

The ultimate remote sensing benefits of the high resolution Infrared radiance spectrometers will be realized with their geostationary satellite implementation in the form of imaging spectrometers. This will enable dynamic features of the atmosphere s thermodynamic fields and pollutant and greenhouse gas constituents to be observed for revolutionary improvements in weather forecasts and more accurate air quality and climate predictions. As an important step toward realizing this application objective, the Geostationary Imaging Fourier Transform Spectrometer (GIFTS) Engineering Demonstration Unit (EDU) was successfully developed under the NASA New Millennium Program, 2000-2006. The GIFTS-EDU instrument employs three focal plane arrays (FPAs), which gather measurements across the long-wave IR (LWIR), short/mid-wave IR (SMWIR), and visible spectral bands. The GIFTS calibration is achieved using internal blackbody calibration references at ambient (260 K) and hot (286 K) temperatures. In this paper, we introduce a refined calibration technique that utilizes Principle Component (PC) analysis to compensate for instrument distortions and artifacts, therefore, enhancing the absolute calibration accuracy. This method is applied to data collected during the GIFTS Ground Based Measurement (GBM) experiment, together with simultaneous observations by the accurately calibrated AERI (Atmospheric Emitted Radiance Interferometer), both simultaneously zenith viewing the sky through the same external scene mirror at ten-minute intervals throughout a cloudless day at Logan Utah on September 13, 2006. The accurately calibrated GIFTS radiances are produced using the first four PC scores in the GIFTS-AERI regression model. Temperature and moisture profiles retrieved from the PC-calibrated GIFTS radiances are verified against radiosonde measurements collected throughout the GIFTS sky measurement period. Using the GIFTS GBM calibration model, we compute the calibrated radiances from data collected during the moon tracking and viewing experiment events. From which, we derive the lunar surface temperature and emissivity associated with the moon viewing measurements.

Tian, Jialin↗

Patterns and Controls on Island-Wide Aboveground Biomass Accumulation in Second-Growth Forests of Puerto Rico

Understanding the heterogeneity of biomass accumulation in second-growth tropical forests following land use abandonment is important for informing ecosystem carbon models and forest restoration efforts. There is an urgent need for a broad sample of second-growth forests to enhance our knowledge of carbon accumulation in human-dominated landscapes, especially for older forests. Puerto Rico has predominantly second-growth forests, ranging in age from approximately 25 to more than 80 years. We used an island-wide sample of airborne lidar from the NASA Goddard Lidar, Hyperspectral, and Thermal (G-LiHT) Airborne Imager collected on March 2017, forest inventory data, and data on forest age, precipitation, soils, and land use to estimate aboveground biomass stocks in moist and wet, second-growth tropical forests. Biomass accumulation rates in Puerto Rico were lower, on average, than in other Neotropical forests. Median biomass across >16,700 ha of older second-growth forests was 105 Mg ha−1, and sampled biomass rarely surpassed 250 Mg ha−1. Differences in biomass by age were large and persistent across different substrates and land uses, with a plateau in the pattern of island-wide biomass accumulation after about 33 years. A spatial regression model showed that multiple factors were related to biomass accumulation, including time since abandonment, geologic substrate, past land use as coffee or pasture, precipitation, topographic wetness index, and slope. Our findings have important consequences for the total carbon storage and expected climate mitigation benefits of large-scale reforestation efforts, and highlight the value of airborne lidar for quantifying biomass variability in complex tropical landscapes.

Sebastián Martinuzzi↗

Exploring the environmental drivers of human blastomycosis cases in the Midwestern United States

Blastomycosis is a fungal infection endemic to the eastern United States (US) and Canada caused by the inhalation of the fungi Blastomyces spp. Currently, the environmental drivers of disease dynamics are poorly understood. The goal of our work was to explore what environmental conditions are associated with the annual presence of blastomycosis cases, and therefore are potentially explanatory of the ecological niche of Blastomyces. We examined the relationships between reported cases of blastomycosis in three Midwestern US states (Michigan, Minnesota, and Wisconsin) from 2007–2017 in relation to eleven hypothesized environmental conditions, including climate, stream and soil mineral content, and land cover variables. Then, we fit logistic regression models to explore the relationships between the environmental variables and yearly blastomycosis case occurrence. Mean soil moisture, stream sediment mercury content, percent of water within the county, and woody wetlands land cover were all positively associated with the presence of annual cases, with woody wetlands having the most consistent signal across the three states. We also found significant differences in the likelihood of case presence between US states that were not explained by the variables in our model, suggesting state-level differences in case reporting and disease awareness. Our results provide a perspective on potential biological hypotheses to further test regarding environmental controls on the life cycle and ecological niche of Blastomyces.

54 ENVIRONMENTAL SCIENCES↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗

What regulates decomposition in agroecosystems? Insights from reading the tea leaves

Litter decomposition is a critical Earth process, recycling nutrients and setting a portion of plant tissue on a path toward soil organic matter. Despite this importance, we still lack a good understanding of local factors that regulate decomposition, especially in agroecosystems where management plays an outsized role. Using a narrow range of climate and soils, we buried 1,308 pre-manufactured “litter bags” of differing residue quality (i.e., green and rooibos tea leaves) in 109 plots across several management practices to (1) explore the local controls on decomposition in agroecosystems and (2) test the robustness of the Tea Bag Index (TBI). We found that management practices intended to increase soil ecosystem services, that is, soil health, altered the decomposition of both teas. For example, adding nitrogen fertilizer and implementing perennial cropping decreased the extent of green tea decomposition (carbon-to-nitrogen ratio, or C:N = 12.8). No-tillage increased, but perennial cropping decreased, the rate of rooibos tea decomposition (C:N = 50.1). Cropped prairie accelerated green tea decomposition and increased the extent of red tea decomposition. A random forest regression model showed that soil temperature was the strongest predictor of green tea decomposition, but a soil health score also played a significant role in predicting the mass remaining. Soil texture and nutrient availability best predicted rooibos tea decomposition. Finer textured soils seemed to decelerate rooibos decomposition but increased the extent of decomposition. Furthermore, we demonstrated that the TBI metrics correlated somewhat well with empirically derived decomposition constants and were similarly sensitive to the effects of management. Still, the green tea stabilization factor had a substantial prediction bias. Our study increased our basic understanding of what regulates decomposition in agroecosystems. It also showed that the TBI can be a scientifically rigorous citizen science approach to monitoring changes in soil health.

60 APPLIED LIFE SCIENCES↗

Uncertainty-Aware Machine Learning for Small-Angle X-ray Scattering Analysis in Autonomous Experimentation

Small-angle X-ray scattering (SAXS) is a powerful high-throughput characterization tool for probing nanoscale structure in native sample environments, providing real-time morphological information such as nanoparticle size and shape during synthesis. However, automated SAXS data analysis for extracting meaningful structural parameters is non-trivial and remains a bottleneck in closed-loop experimentation towards autonomous materials discovery, which demands fast, reliable, and uncertainty-aware data analysis. Here, we develop a machine-learning approach for automated SAXS analysis tailored to closed-loop nanoparticle synthesis. A Random Forest (RF) regression model is trained on 100,000 synthetic SAXS curves generated from polydisperse spherical nanoparticles with realistic background contributions. Using normalized one-dimensional SAXS intensity profiles as input, the RF model directly predicts nanoparticle radius, size polydispersity, and background parameters, while the ensemble standard deviation across trees provides built-in uncertainty quantification (UQ). On synthetic data, we show that combining fit-quality metrics (R 2 , MAE) with thresholds on prediction uncertainty reliably identifies accurate parameter estimates without access to ground truth. We then apply the trained model to 365 experimental SAXS profiles of citrate-reduced gold nanoparticles synthesized using an automated droplet-flow microreactor with in situ SAXS at a synchrotron beamline, classifying the results into high- and low-confidence subsets based on UQ metrics. Finally, we integrate RF-based SAXS analysis into a simulated closed-loop optimization campaign using Gaussian process Bayesian optimization to minimize nanoparticle polydispersity, benchmarking against conventional automated Levenberg–Marquardt fitting. The RF-guided campaign exhibits substantially faster convergence and lower relative opportunity cost (∼0.07 vs ∼0.3), demonstrating that uncertainty-aware machine-learning SAXS analysis significantly enhances the efficiency and robustness of autonomous nanomaterials synthesis workflows.

Bayesian optimization↗

Leveraging hyperspectral phenotyping for accurate, non-destructive prediction of metabolite profiles in poplar under drought stress

Accurately predicting drought tolerance in woody perennial bioenergy crops is critical for sustainable biomass production under fluctuating precipitation. Hyperspectral imaging (HSI) in the visible-near-infrared (VNIR) and shortwave-infrared (SWIR) ranges offers a promising approach for predicting plant biochemical traits, yet its application in metabolite profiling remains underexplored. We integrated VNIR+SWIR HSI with untargeted metabolomics to investigate drought-induced metabolic shifts in Populus leaves from eight Populus genotypes. Metabolite profiling identified 127 compounds, with 73 showing significant drought responses spanning amino acids (AA), carbohydrates (CHO), phenolic glycosides (PG), organic acids (OA), fatty acids and alcohols (FA), terpenes (T), phenolic metabolites (P), and unclassified metabolites. Spectral analysis revealed consistently higher reflectance across VNIR and SWIR wavelengths in drought-stressed plants, corresponding with increased accumulation of AA and reduced CHO and PG levels. Least absolute shrinkage and selection operator (LASSO) regression modeling identified robust spectral predictors of metabolite concentrations, associating VNIR wavelengths (500–700 nm) predominantly with AA and P, whereas SWIR wavelengths (1680–1700 nm) reliably predicted CHO, OA, and T. Several stable spectral-metabolite associations persisted across the two watering regimes (drought vs. well-watered), highlighting their potential as spectral biomarkers for non-destructive stress monitoring. Minimal genotype-specific variation suggests that observed spectral and metabolic responses were driven primarily by environmental factors, likely reflecting limited genetic diversity among the commercial Populus genotypes examined. This work establishes VNIR+SWIR hyperspectral imaging as a powerful, non-destructive phenotyping tool for precision monitoring and targeted improvement of drought resilience in bioenergy crops.

Biochemical trait prediction↗

Evaluating Soil Moisture Retrievals from ESA's SMOS and NASA's SMAP Brightness Temperature Datasets

Two satellites are currently monitoring surface soil moisture (SM) using L-band observations: SMOS (Soil Moisture and Ocean Salinity), a joint ESA (European Space Agency), CNES (Centre national d'tudes spatiales), and CDTI (the Spanish government agency with responsibility for space) satellite launched on November 2, 2009 and SMAP (Soil Moisture Active Passive), a National Aeronautics and Space Administration (NASA) satellite successfully launched in January 2015. In this study, we used a multilinear regression approach to retrieve SM from SMAP data to create a global dataset of SM, which is consistent with SM data retrieved from SMOS. This was achieved by calibrating coefficients of the regression model using the CATDS (Centre Aval de Traitement des Donnes) SMOS Level 3 SM and the horizontally and vertically polarized brightness temperatures (TB) at 40 deg incidence angle, over the 2013 - 2014 period. Next, this model was applied to SMAP L3 TB data from Apr 2015 to Jul 2016. The retrieved SM from SMAP (referred to here as SMAP_Reg) was compared to: (i) the operational SMAP L3 SM (SMAP_SCA), retrieved using the baseline Single Channel retrieval Algorithm (SCA); and (ii) the operational SMOSL3 SM, derived from the multiangular inversion of the L-MEB model (L-MEB algorithm) (SMOSL3). This inter-comparison was made against in situ soil moisture measurements from more than 400 sites spread over the globe, which are used here as a reference soil moisture dataset. The in situ observations were obtained from the International Soil Moisture Network (ISMN; https:ismn.geo.tuwien.ac.at) in North of America (PBO_H2O, SCAN, SNOTEL, iRON, and USCRN), in Australia (Oznet), Africa (DAHRA), and in Europe (REMEDHUS, SMOSMANIA, FMI, and RSMN). The agreement was analyzed in terms of four classical statistical criteria: Root Mean Squared Error (RMSE),Bias, Unbiased RMSE (UnbRMSE), and correlation coefficient (R). Results of the comparison of these various products with in situ observations show that the performance of both SMAP products i.e. SMAP_SCA and SMAP_Reg is 48 similar and marginally better to that of the SMOSL3 product particularly over the PBO_H2O, SCAN, and USCRN sites. However, SMOSL3 SM was closer to the in situ observations over the DAHRA and Oznet sites. We found that the correlation between all three datasets and in situ measurements is best (R 0.80) over the Oznet sites and worst (R 0.58) over the SNOTEL sites for SMAP_SCA and over the DAHRA and SMOSMANIA sites (R 0.51 and R 0.45 for SMAP_Reg and SMOSL3, respectively). The Bias values showed that all products are generally dry, except over RSMN, DAHRA, and Oznet (and FMI for SMAP_SCA). Finally, our analysis provided interesting insights that can be useful to improve the consistency between SMAP and SMOS datasets.

SMOS↗

Impact of Grazing Duration and Environment on Soil Carbon in Reclaimed Uranium Mines Tailings: A Region Specific Study

ABSTRACT Grassland ecosystems, which cover over one‐third of the Earth's land area, store 10%–30% of global soil carbon (C). However, these ecosystems face substantial impacts from human activities, including mining. This study investigates the spatial distribution of soil C and related environmental factors in reclaimed grasslands on former uranium mine sites in Wyoming. We hypothesized that grazing duration and environmental factors would influence soil C levels. Interactions between topography, vegetation diversity, soil properties, and soil C in the context of grazing management in both natural and reclaimed grasslands from a wide range of periods from 1 year to 100 years were analyzed using geographically weighted regression models. Data collected from 2022 to 2023 showed that total carbon was consistently higher in natural grasslands (1.2%–4.9%) than in reclaimed grasslands (0.8%–1.3%). Additionally, soil C was significantly higher in natural grasslands grazed for 1 year compared to those grazed for 100 years. In contrast, reclaimed grasslands had lower soil C in areas grazed for 1 year compared to those grazed for 7 or 14 years. The absolute values of coefficients from environmental covariates indicated that areas grazed for a shorter duration (~1 year) were more influenced by biotic and abiotic factors than areas grazed for longer periods (> 7 years). Our findings show moderate grazing increases the resiliency of grassland ecosystems when grazed 7 years or longer and acknowledge the roles of topographic, soil, and vegetative factors in enhancing soil C concentration and developing sustainable land management practices in rangeland conditions.

Shilpakar, Chandan [Department of Plant Sciences U↗

Predictive links between microbial communities and biological oxygen utilization in the Arctic Ocean

Microbial metabolism influences rates of net community production (NCP), exerting a direct biological control on marine oxygen and carbon fluxes. In the Arctic, it is increasingly important to understand and quantify this process, as ecological and oceanographic conditions shift due to changing climate. Here, we describe potential ecological links between pelagic microbial diversity and an NCP precursor, biological oxygen utilization, using machine learning and paired observations of community structure and metabolic activity from a seasonally and spatially variable transect of the Arctic Ocean (2019–2020 MOSAiC Expedition). Community structure was determined using 16S (prokaryotic) and 18S (eukaryotic) rRNA gene amplicon sequencing, and metabolic activity was derived from ΔO 2 /Ar. Using self-organizing maps, we identified clear successional patterns in observed microbial community structure that were seasonally driven in the upper ocean and vertically stratified with depth. Metabolic activity was also stratified, with a primarily net heterotrophic water column (median −1.5% biological oxygen saturation), excepting periodic oxygen supersaturation (maximum: 13.6%) within the mixed layer. Using DNA sequences as predictor variables, we then constructed a random forest regression model that reliably reconstructed biological oxygen concentrations (root mean squared error = 4.14 μmol kg −1 ). Top predictors from this model were from heterotrophic (bacteria) or potentially mixotrophic (dinoflagellate) taxa. These analyses highlight biologically driven diagnostic tools that can be used to expand biogeochemical datasets and improve the microbial perspectives and metabolisms represented in ecological models of net productivity and carbon flux in a changing Arctic Ocean.

Chamberlain, Emelia J. [Univ. of San Diego, San Di↗