Search NASA⌕ Search

SEARCH · Search NASA

Results for “predictor”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Explaining drivers of housing prices with nonlinear hedonic regressions

Housing markets play a critical role in shaping the spatial and demographic evolution of urban areas. Simulating housing price dynamics can enhance projections of future urban development outcomes. However, traditional hedonic regressions for housing prices, which neglect nonlinear interactions among explanatory variables, often exhibit limited predictive performance. While machine learning (ML) methods can provide a more flexible representation of the relationships between predictors, they are often regarded as “black boxes” due to their complexity and lack of transparency. Interpretable ML techniques provide a promising route by combining the flexibility of ML methods with approaches to analyze the relationships between inputs and outputs. In this study, we employ interpretable ML to analyze the patterns driving the housing market in Baltimore, Maryland, USA. We train an Artificial Neural Network (ANN) to predict Baltimore housing prices based on structural characteristics (e.g., home size, number of stories) and locational attributes (e.g., distance to the city center). We then conduct sensitivity and Partial Dependence Plot (PDP) analyses to interpret the fitted ANN model. We find that the ML model achieves higher predictive accuracy and explains 16 % more of housing price variance than a traditional linear regression model. The interpretable ML model also reveals more nuanced and realistic nonlinear relationships between housing sales price and predictors as well as interactive effects underlying Baltimore home price dynamics. For instance, while the linear model indicates a steady housing price increase over time, our interpretable ML model detects a post-2008 decline, with smaller properties experiencing the sharpest drop.

97 MATHEMATICS AND COMPUTING↗

Machine Learning-Assisted Recovery of Delicate Kinetic Information from Transient Reactor Experiments

Identifying active sites and their roles in chemical reaction steps remains a vital challenge in heterogeneous catalysis. Transient experiments offer a unique way to probe active sites and distinguish subtle kinetic features. Although physics-based analysis methods may be well-developed, they can be highly susceptible to experimental noise, and smoothing methods may erase or even distort important features; a smooth curve is not always the best curve. We demonstrate a new workflow for the direct interpretation of intrinsic kinetic information from exit flux curves measured in transient reactor experiments. This workflow contains three artificial neural networks (ANNs), including a noise reducer, a concentration predictor, and a rate predictor to analyze experimental data, followed by the virtual TAP (VTAP) physics-based reactor model and density functional theory (DFT) calculations of adsorption energies on specific sites. We use this workflow to analyze the data from experiments titrating Pt/Al 2 O 3 and Pt/SiO 2 catalysts with carbon monoxide (CO) in the temporal analysis of products (TAP) reactor. Our workflow separates the time-evolving chemical reaction and mass transfer information contained in the TAP pulse response. The existence of strong- and weak-binding sites on the Pt/Al 2 O 3 catalyst is observed in the catalyst titration experiment in the transient reactor. The structures of the strong- and weak-binding sites are then identified by using DFT calculations. We find that the Pt/SiO 2 catalyst has only strong-binding sites, which aligns with the inactive support effect of SiO 2 . We demonstrate how machine learning methods provide unique insights with high-resolution data analysis that cannot be achieved by using state-of-the-art physics-based methods.

Adsorption↗

Large‐Scale Statistically Meaningful Patterns (LSMPs) Associated With Precipitation Extremes Over Northern California

Abstract We analyze large‐scale statistically meaningful patterns (LSMPs) that precede extreme precipitation (PEx) events over Northern California (NorCal). We find LSMPs by applying k‐means clustering to the two leading principal components of daily 500 hPa geopotential height anomalies two days before the onset, from October to March during 1948–2015. Statistical significance testing based on Monte Carlo simulations suggests a minimum of four statistically distinguished LSMP clusters. The four LSMP clusters are characterized as Northwest continental negative height anomaly, Eastward positive “Pacific‐North American Pattern (PNA),” Westward negative “PNA,” and Prominent Alaskan ridge. These four clusters, shown in multiple variables, evolve very differently and have differing links to the Arctic and tropical Pacific regions. Using binary forecast skill measures and a new copula‐based framework for predicting PEx events, we find LSMP indices that are useful predictors of NorCal PEx events, with moisture‐based variables being the best predictors of PEx events at least 6 days before the onset, and the lower atmospheric variables being better than their upper atmospheric counterparts any day in advance tested. To ensure statistical rigor, the LSMPs analyzed here (with the modified acronym) include local tests of both significance and consistency, which are not always featured in the literature on large‐scale meteorological patterns.

54 ENVIRONMENTAL SCIENCES↗

Updraft Width Modulates Ambient Atmospheric Controls on Convective Cloud Depth

Abstract The depth of convective clouds affects vertical transport of atmospheric constituents, influencing downstream weather and climate. Atmospheric controls on the maximum depth reached by moist convection are investigated with radar‐tracked convective cells tagged with sounding‐derived atmospheric parameters from a field campaign in central Argentina. Regression analyses show that narrow (<12‐km diameter) and wide (>16‐km diameter) cell depths respond to disparate factors, where cell areas are defined using composite reflectivity signatures. Undiluted lifted parcel indices including convective available potential energy (CAPE) and level of neutral buoyancy (LNB) are top predictors of wide cell maximum depth while mid‐tropospheric relative humidity is the top predictor of narrow cell maximum depth. Because narrow cells are more numerous than wide cells, the overall outcome of the full cell population does not strongly correlate with CAPE and LNB conditions. Tracked cells and atmospheric conditions in a simulation with 3‐km grid spacing covering the field campaign produce similar results to those observed. Narrow cells that are relatively deep have a cooler and moister mid‐troposphere with weaker free tropospheric subsidence, while relatively deep wide cells have much warmer and moister lower tropospheric conditions. These atmospheric differences are present 1 hr before cell initiation at both a fixed observing site and variable cell initiation locations. Simulated narrow cell maximum equivalent potential temperature decreases with height at a rate similar to the ambient vertical gradient, causing these cells to fall short of their LNB and supporting the view that entrainment‐driven dilution is a dominant control on their depth.

54 ENVIRONMENTAL SCIENCES↗

Multidecadal Fluctuations in the Observed ENSO‐Tropical Cyclone Teleconnection

Abstract El Niño‐Southern Oscillation (ENSO) is a skillful predictor for seasonal tropical cyclone (TC) activity in most TC basins. This study examines recent changes in the observed ENSO‐TC teleconnection strength, as measured by ENSO modulation of hurricane frequency. We find that the ENSO‐North Atlantic TC teleconnection fluctuated over time, with the strongest relationship occurring from the 1980s to the mid‐2000s. In the western and eastern North Pacific, the ENSO‐TC teleconnection has strengthened in recent decades. Periods with a strong ENSO‐TC teleconnection are associated with more favorable environmental conditions for TCs, with higher values of genesis potential indices. Positive phases of the Atlantic Multidecadal Oscillation coincided with periods of strong ENSO‐TC teleconnections in the Atlantic and North Pacific basins. A weaker Atlantic ENSO‐TC relationship was associated with negative phases of the Pacific Decadal Oscillation and the North Atlantic Oscillation. This research reveals climate conditions that modulate ENSO's utility for seasonal TC prediction. Plain Language Summary El Niño‐Southern Oscillation (ENSO) is a useful predictor for seasonal tropical cyclone (TC) activity in many basins. Here we found that the strength of the ENSO‐TC teleconnection, represented as the correlation between ENSO and the number of hurricanes and accumulated cyclone energy, has changed in the historical record. The ENSO‐TC teleconnection in the North Atlantic fluctuated over time, with a weak relationship during the 1960s and 1970s and a strong relationship during the 1980s to mid‐2000s. Meanwhile, the ENSO‐TC teleconnection strengthened in the North Pacific in recent decades, with strong teleconnections after the 1980s in the western North Pacific and after the 2000s in the eastern North Pacific. Periods of strong ENSO‐TC teleconnections are associated with more favorable environmental conditions for TCs, including higher values of genesis potential indices and higher mid‐tropospheric humidity, as well as positive phases of the Atlantic Multidecadal Oscillation. Additionally, the negative phase of the Pacific Decadal Oscillation leads to strong/weak ENSO‐TC teleconnections in the eastern North Pacific and North Atlantic, respectively. Furthermore, a negative North Atlantic Oscillation is associated with a weak ENSO‐North Atlantic TC teleconnection. This research highlights variations in ENSO's effectiveness for seasonal TC prediction. Key Points The observed impact of ENSO on tropical cyclone (TC) activity exhibits multidecadal fluctuations The ENSO‐TC teleconnection was strong in the Atlantic from the 1980s to mid‐2000s and strengthened over the North Pacific in recent decades The ENSO‐TC teleconnection is stronger in the Atlantic and North Pacific basins during a positive Atlantic Multidecadal Oscillation

ENSO↗

Environmental Factors Associated With Fall Phytoplankton Blooms in the Northern Bering and Chukchi Seas

This study investigates environmental drivers of fall phytoplankton blooms in the Arctic, focusing on the northern Bering and Chukchi seas. Random Forests models were used to analyze covariates of fall phytoplankton blooms from 2013 to 2018, incorporating shipboard, remote sensing, and modeled environmental properties. Four regional models and one comprehensive all-station model considered fall as well as midsummer conditions. Midsummer properties included suspended particulate matter, chlorophyll-a, and the proportion of degraded pheophytin to chlorophyll-a used as a proxy for bloom stage. Open water duration was one of the highest ranked factors in predicting fall blooms. Open water duration also influences the stage of midsummer (July) blooms as indicated by pheophytin proportions, which in turn were the highest-ranked factor for predicting fall bloom events in the Chirikov Basin (northern Bering Sea between St. Lawrence Island and the Bering Strait) and the Chukchi Sea. Wind direction, specifically easterly winds, was an important predictor in the northern Bering Sea. Maximum wind speed ranked highly at stations located within the nutrient-poor Alaska Coastal Current in the Chukchi Sea. However, stormy days, average and maximum wind speeds generally ranked low in importance as a predictor of fall bloom events. Other parameters, including photosynthetic active radiation, modeled nutrient concentrations, mixed layer depth, and time since sea ice breakup date showed strong but regionally varying relationships with fall blooms. Altogether, results from these Random Forests models suggest that high wind events and storms in the absence of sea ice provide an incomplete narrative for initiating fall bloom events.

Gaffey, C. B. [Clark University, Worcester, MA (Un↗

Environmental Controls on Water Vapor Deuterium Excess in the Coastal Boundary Layer: An Information Theory Perspective

We use information theory to quantify the environmental controls on water vapor deuterium excess (D-excess) in coastal Southern California from June 2023 through February 2024. Using Shannon entropy, mutual information (MI), and joint mutual information, metrics that capture both linear and nonlinear relationships, we identify the most informative variables and variable combinations governing D-excess across contrasting marine and continental regimes. Relative humidity with respect to sea surface temperature (RHS) is consistently the strongest individual predictor, explaining up to 27% of D-excess variability during marine conditions but only 10% in continental air masses. The Relative humidity(RHS) + sea surface temperature (SST) combination demonstrates synergistic effects, where their joint influence (explaining up to 36% of D-excess variability) exceeds what either variable achieves individually, confirming their coupled influence on deuterium excess. Wind direction complements RHS most effectively during continental conditions. The best three-variable combination (RHS + SST + Planetary Boundary Layer height) explains 38% of D-excess variability in marine air, while no combination exceeds 20% explanatory power during continental periods. Information theory shows that heteroscedasticity in D-excess relationships indicates regime shifts in controlling processes and quantifies fundamental constraints on predictor variables: some environmental factors like surface pressure or water vapor flux contain insufficient information content to explain D-excess variability regardless of their physical relevance. These results highlight the different predictability limits between marine and continental regimes, challenging the adequacy of linear models and providing a rigorous framework for quantifying the information content of isotope-climate relationships with implications for both modern and paleoclimate applications.

information theory↗

Dominant Controls on Preferential Flow and Their Implications for Future Soil Water Fluxes

Abstract Soil water flow, particularly preferential flow (PF), is a critical control on hydrological and biogeochemical processes, including groundwater recharge, contaminant transport, and carbon cycling. However, it remains challenging to predict PF occurrence across large environmental gradients. Here, we developed a deep learning (DL) model to estimate event‐scale soil water flow velocity and the probability of PF occurrence using high‐frequency soil moisture and precipitation data from 33 sites across the National Ecological Observatory Network. The model demonstrated high skill in predicting the binary occurrence of PF (91% F1‐score; 85% accuracy) but the performance was limited in predicting soil water velocity ( R 2 = 0.31). We found that precipitation characteristics (duration, volume, and intensity) were the most important predictors for soil water velocity. Among the non‐precipitation event variables, sand content showed relatively high predictive skill, though differences among non‐event climate variables were generally modest. Lower sand content was associated with increased predicted soil water velocity, a finding that highlights the role of soil structure in producing more non‐uniform flow, which contrasts with traditional uniform flow models. Projecting a reduced DL model under both moderate and high‐emissions future climate scenarios (2060–2099 Representative Concentration Pathways 4.5 and 8.5), we found ∼7.3% increase under RCP4.5 and ∼15% under RCP8.5 of soil water velocities compared to the historical simulation, while modeled likelihood of PF changed little. These findings suggest climate change is not making PF more frequent, but it is making existing PF pathways more efficient with important consequences for associated nutrient and contaminant transport under climate change. Plain Language Summary Water movement in soil is critical for water quality. While often modeled as a uniform flow process, in reality water moves rapidly through cracks and burrows in what is called “preferential flow” (PF), which limits natural filtration and can transport pollutants. We developed a deep learning model, trained on data from 33 U.S. sites, to predict when and how fast this PF occurs based on precipitation, soil, and climate data. The model showed that precipitation characteristics (duration, intensity, volume) were the most important predictors of PF. Lower soil sand content/higher clay content was associated with faster water flow, likely due to clay soils forming aggregates and cracks that water moves through rather than infiltrating uniformly. Further analyses based on climate projections suggest that the speed at which PF occurs will become more rapid under future climate scenarios compared to historical simulation. This highlights the need to represent PF in soil water models when assessing future water quality. Key Points The effect of precipitation peak intensity on soil water velocities declined with increasing precipitation intensity Antecedent soil moisture failed to predict preferential flow (PF), contrasting the high predictive power of sand content Climate predictions suggest that soil water velocities through PF paths will increase ∼15% by 2099

Li, Bonan↗

The Effect of Updraft Entrainment on Convective Cell Deepening in Realistic Large-Eddy Simulations

Entrainment of surrounding cooler and drier air into convective updrafts is one of the key processes that influence deep convection initiation and growth. Numerous studies have investigated the effect of entrainment on isolated convective cloud growth in idealized simulations, but the importance of this effect in realistic conditions with many interacting convective clouds remains uncertain. We examine the impact of entrainment on the depth reached by convective clouds in realistic large-eddy simulations (LES) over central Argentina during the Cloud, Aerosol, and Complex Terrain Interactions (CACTI) field campaign. Cloudy updrafts and their associated properties are assigned to convective cells tracked with radar reflectivity signatures. Several thousand convective cells are tracked over two high convective available potential energy (CAPE) and two low CAPE cases that support cells of varying depths and intensities. Entrainment is calculated explicitly as the fluxes of air into the outer surface of each cloudy updraft. Single-predictor logistic regression models are used to determine the relative importance of updraft, near-updraft, and preconvective initiation atmospheric conditions in predicting whether convective cells become deep. We then build a multiple-predictor regression model pairing important updraft and meteorological metrics with fractional entrainment rate. The probability of cells transitioning to deep convection is most sensitive to ambient 600-hPa relative humidity (42% of total metric contribution to cloud depth predictability), followed by low-level CAPE (28%), cloud-base updraft width (19%), and fractional entrainment (11%). Thus, the initial width of the updraft along with potential buoyancy and its dilution through the midtroposphere collectively determine whether deep convection will result from shallower clouds.

54 ENVIRONMENTAL SCIENCES↗

Dependence of Deep Convective Cell Properties on Meteorological and Aerosol Conditions during TRACER

Deep convective cells significantly influence Earth’s energy balance and water cycle. However, their accurate representation in numerical models remains challenging due to their small spatiotemporal scales and limited observational constraints. This study examines over ∼400 deep convective cells near Houston, observed by a dual-polarization C-band radar during the Tracking Aerosol Convection Interactions Experiment (TRACER) intensive observation period (June–September 2022). Cells are categorized by lifetime into short-lived (<40 min), intermediate-lived (40–80 min), and long-lived (80+ min) groups. Long-lived cells were broader (∼13.2 km at 2–4-km height) and deeper (∼11.4 km) than short-lived cells (∼6.4-km width, ∼7.31-km height). Using random forest (RF) modeling and correlation analyses, precipitable water vapor (PWV), 2–6-km lapse rate, 0–8-km bulk shear, and fine aerosol mass concentration (Mass_f) are identified as key predictors of cell lifetime. Higher PWV is associated with significantly longer convective cell lifetimes compared to the low-PWV group, particularly within low 2–6-km temperature lapse rate (LR_26km), moderate-to-higher 0–8-km bulk shear (BS_08km), and low-to-moderate Mass_f environments. RF analysis also identifies low-level (0–2 km) equivalent potential temperature, PWV, Mass_f, and surface latent heat flux as key predictors for cell width and height. Short-lived cells have higher aerosol number concentrations (500–1000-nm size range), linked to onshore wind conditions and marine aerosols; however, their low concentration suggests the sensitivity may reflect associated meteorological regimes rather than a direct aerosol effect. Long-lived cells have higher concentrations of organic and sulfate aerosols, while short-lived cells exhibit higher black carbon concentrations. These results highlight the intricate dependence of convective cell lifetimes and structure on environmental moisture, thermodynamics, wind shear, and aerosol characteristics.

54 ENVIRONMENTAL SCIENCES↗

Data for Propagation Method and Planting Density Influence Canopy Developmental Transition and Biomass Productivity in Miscanthus × giganteus

Understanding how establishment practices influence the mechanisms underlying Miscanthus × giganteus (miscanthus) productivity and canopy development is critical for optimizing management. Data was collected during the juvenile (2011–2013) and mature (2024) phases of a long-term field experiment established in Urbana, Illinois, to evaluate the effects of propagation method (plug propagation [PP] and rhizome propagation [RP]), planting density (1.0, 0.75, and 0.25 plants m⁻²), and nitrogen application (0 and 67 kg N ha⁻¹) on end-of-season biomass yield, tiller mass, tiller density, and tiller height. Linear regression models identified the dominant predictors of yield across stand ages and management regimes. Planting density, nitrogen (N) application, and propagation method significantly influenced early yield and canopy development. During the juvenile phase, biomass yield was driven by tiller density due to canopy expansion; in the mature phase, yield became driven by tiller mass. The PP plots produced higher tiller density than the RP plots, resulting in faster canopy closure and higher juvenile-phase yields. Rhizome-propagated (RP) plots produced lower tiller density, but individual tillers were 3.3–6.4 g tiller−1 heavier than PP tillers. After the canopy reached equilibrium, the PP and RP yields were similar because greater RP tiller mass compensated for its lower tiller density. Higher planting density resulted in greater yield and tiller density during the second year (2012), but this effect was absent from the third year (2013) onward. In the juvenile phase, N fertilization enhanced yield by 1.6–3.4 Mg ha−1. Initiating fertilization in 2013 on unfertilized plots produced biomass similar to that in fertilized plots, suggesting yield recovery in the mature phase. These findings revealed that establishment strategies, including propagation method and planting density, influence juvenile miscanthus canopy development and productivity, transitioning from tiller-density- to mass-dominated yields, but not mature phase productivity.

Miscanthus↗

Deep Learning Prediction of Protein Complex Structures

Proteins interact to form protein complex to carry out biological functions such as catalytic chemical reaction. Therefore, it is important to develop computational methods to predict protein-protein interaction and the structures of protein complexes to study and enhance protein function. In this project, we successfully developed several deep learning methods to predict inter-protein contacts and the reinforcement learning and optimization methods to reconstruct protein complex structures from predicted inter-chain contacts. The methods were integrated with the MULTICOM protein complex structure prediction system and applied to predict the complex structures of biomass production-related proteins of green algae. During the two and a half years of research and development, all the specific milestones of the project were achieved successfully. 16 publications/manuscripts were produced. 10 software tools were developed. A patent application was submitted. Our MULTICOM predictors leveraging some tools developed in this project were ranked among the top predictors in the 15th Critical Assessment of Techniques for Protein Structure Prediction (CASP15) in 2022.

59 BASIC BIOLOGICAL SCIENCES↗

Product Innovation to Increase Low-to-Moderate-Income Customers' Adoption of Community Solar PV

This study aims to comprehensively analyze community solar project preferences for consumers and suppliers by conducting three distinct analyses. First we analyze the predictors of community solar contract adoption to understand how individual priorities affect the probability of adoption. Using an original data set of survey responses from potential community solar customers, we analyzed the predictors of contract adoption by employing a weighted logit model. We find that individuals who were previously familiar with community solar projects were significantly more likely to adopt than those who were not familiar. Secondly, a survey of community solar developers and financiers identified industry perceived barriers to community solar access and inclusion. Thirdly, we gathered payment performance information from community solar initiatives to measure how financial risks are perceived and how they interact with customer demographics. Our study is beneficial to the public by providing insights into the drivers and barriers of community solar adoption and sheds light on the importance of understanding individual priorities in designing effective community solar policies. The community solar industry has changed significantly since the beginning of this project. The industry continues to grow at a rapid rate, with an additional 7 gigawatts expected to come online between 2022 and 2027. With federal pressure to meet climate goals, as the harms of climate change continue to impact everybody, legislators are looking to community solar as a method to achieving their states energy policy goals. These new policies that push for low-to-moderate inclusion, coupled with the increase in community solar capacity illustrate a new era for the community solar industry. A number of policies have arisen in the last few months that push for more inclusive practices including the groundbreaking Solar for All program run by the EPA. The research created a “best practice” contract that can then be used, in conjunction with the manuscript and validated study, to pitch the industry on a more inclusive community solar product.

14 SOLAR ENERGY↗

wa-hls4ml and lui-gnn: A benchmark and GNN-based surrogate model for hls4ml resource and latency estimation

As machine learning (ML) increasingly serves as a tool for addressing real-time challenges in scientific applications, the development of advanced tooling has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as model synthesis, are now becoming limiting factors in the rapid iteration of designs. To reduce these emerging constraints, multiple efforts are being launched toward designing an ML-based surrogate model that estimates resource usage of synthesized accelerator architectures. This model would reduce the design iteration time, especially when designing within a set of given hardware constraints. This approach shows considerable potential, but as it stands, the effort is early and would benefit from coordination and standardization to assist future work as it emerges. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of more than 100,000 fully connected neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. In addition to the resource utilization and latency data provided, the dataset includes generated artifacts and log files for many of the synthesized neural networks, in order to support future research in ML-based code generation. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, as well as the average performance across a subset of the dataset. We measure the performance of a given predictor model through multiple metrics, including $R^2$ score and SMAPE on regression tasks, as well as inference time to further characterize the estimator under test. Additionally, we introduce the latency/utilization inference graph neural network (lui-gnn), a surrogate model that uses a graph neural network to represent input architectures in the form of a directed graph. This graph representation allows for a diverse set of model architectures to all be effectively handled by a surrogate model. We present the architecture and performance of the model, as evaluated by the new proposed benchmark, including SMAPE, $R^2$ score, and inference times, and find that lui-gnn generally predicts latency and utilization for the 75\% quantile within several percent of the synthesized resources on the synthetic test dataset, indicating that this approach of estimating resource and latency via a surrogate models has promise and warrants further research.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Assessing Heterogeneity of Surface Water Temperature Following Stream Restoration and a High-Intensity Fire from Thermal Imagery

Thermal heterogeneity of rivers is essential to support freshwater biodiversity. Salmon behaviorally thermoregulate by moving from patches of warm water to cold water. When implementing river restoration projects, it is essential to monitor changes in temperature and thermal heterogeneity through time to assess the impacts to a river’s thermal regime. Lightweight sensors that record both thermal infrared (TIR) and multispectral data carried via unoccupied aircraft systems (UASs) present an opportunity to monitor temperature variations at high spatial (<0.5 m) and temporal resolution, facilitating the detection of the small patches of varying temperatures salmon require. Here, we present methods to classify and filter visible wetted area, including a novel procedure to measure canopy cover, and extract and correct radiant surface water temperature to evaluate changes in the variability of stream temperature pre- and post-restoration followed by a high-intensity fire in a section of the river corridor of the South Fork McKenzie River, Oregon. We used a simple linear model to correct the TIR data by imaging a water bath where the temperature increased from 9.5 to 33.4 °C. The resulting model reduced the mean absolute error from 1.62 to 0.35 °C. We applied this correction to TIR-measured temperatures of wetted cells classified using NDWI imagery acquired in the field. We found warmer conditions (+2.6 °C) after restoration (p < 0.001) and median absolute deviation for pre-restoration (0.30) to be less than both that of post-restoration (0.85) and post-fire (0.79) orthomosaics. In addition, there was statistically significant evidence to support the hypothesis of shifts in temperature distributions pre- and post-restoration (KS test 2009 vs. 2019, p < 0.001, D = 0.99; KS test 2019 vs. 2021, p < 0.001, D = 0.10). Moreover, we used a Generalized Additive Model (GAM) that included spatial and environmental predictors (i.e., canopy cover calculated from multispectral NDVI and photogrammetrically derived digital elevation model) to model TIR temperature from a transect along the main river channel. This model explained 89% of the deviance, and the predictor variables showed statistical significance. Collectively, our study underscored the potential of a multispectral/TIR sensor to assess thermal heterogeneity in large and complex river systems.

Barker, Matthew I. (ORCID:0000000252864930)↗

Bayesian chain graph models to characterize microbe-environment dynamics

Microbiome data require statistical models that can simultaneously decode microbes' reaction to the environment and interactions among microbes. While a multiresponse linear regression model seems like a straight-forward solution, we argue that treating it as a graphical model is problematic given that the regression coefficient matrix does not encode the conditional dependence structure between response and predictor nodes. This observation is especially important in biological settings when we have prior knowledge on the edges from specific experimental interventions that can only be properly encoded under a conditional dependence model. Here, we propose a chain graph model with two sets of nodes (predictors and responses) whose solution yields a graph with edges that indeed represent conditional dependence, thus agreeing with the experimenter's intuition on the average behavior of nodes under treatment. The solution to our model is sparse via the Bayesian linear regression (LASSO). In addition, we propose an adaptive extension so that different shrinkages can be applied to different edges to incorporate edge-specific prior knowledge. Our model is computationally inexpensive through an efficient Gibbs sampling algorithm and can account for binary, counting, and compositional responses via an appropriate hierarchical structure. We test the performance of our model in a variety of simulated datasets, thereby showing superior performance to state-of-the-art approaches. We further apply our model to human gut and soil microbial compositional datasets, and we highlight that CG-LASSO can estimate biologically meaningful network structures in the data.

compositional data↗

Observational benchmarks inform representation of soil organic carbon dynamics in land surface models

Abstract. Representing soil organic carbon (SOC) dynamics in Earth system models (ESMs) is a key source of uncertainty in predicting carbon–climate feedbacks. Machine learning models can help identify dominant environmental controllers and establish their functional relationships with SOC stocks. The resulting knowledge can be integrated into ESMs to reduce uncertainty and improve predictions of SOC dynamics over space and time. In this study, we used a large number of SOC field observations (n=54 000), geospatial datasets of environmental factors (n=46), and two machine learning approaches (namely random forest, RF, and generalized additive modeling, GAM) to (1) identify dominant environmental controllers of global and biome-specific SOC stocks, (2) derive functional relationships between environmental controllers and SOC stocks, and (3) compare the identified environmental controllers and predictive relationships with those in models used in Phase 6 of the Coupled Model Intercomparison Project (CMIP6). Our results showed that the diurnal temperature, drought index, cation exchange capacity, and precipitation were important observed environmental predictors of global SOC stocks. While the RF model identified 14 environmental factors that describe climatic, vegetation, and edaphic conditions as important predictors of global SOC stocks (R2=0.61, RMSE = 0.46 kg m−2), current ESMs oversimplify the relationships between environmental factors and SOC, with precipitation, temperature, and net primary productivity explaining > 96 % of the variability in ESM-modeled SOC stocks. Further, our study revealed notable disparities among the functional relationships between environmental factors and SOC stocks simulated by ESMs compared with observed relationships. To improve SOC representations in ESMs, it is imperative to incorporate additional environmental controls, such as the cation exchange capacity, and refine the functional relationships to align more closely with observations.

54 ENVIRONMENTAL SCIENCES↗

Xanthos-Lake Dataset

The Xanthos-Lake v1.0 dataset provides the input data, trained machine-learning models, and simulation outputs needed to characterize lake water balance, snow and ice conditions, and mixing-layer temperature within the Xanthos global hydrological modeling framework. The dataset supports lake representation across a wide range of lake sizes and hydroclimatic conditions by combining xLSIM, a basin-specific machine-learning emulator of lake snow, ice, ice-cover fraction, and mixing-layer temperature, with the Xanthos-Lake water-balance model. The archive contains NetCDF datasets used to train and evaluate xLSIM, trained model weights, processed meteorological and lake-property inputs, and basin- and lake-category-specific simulation outputs. These materials are organized into four primary data groups, described below. Snowice_model_inputs: Contains the NetCDF input data used to train xLSIM. The xLSIM machine-learning framework uses three lake-based datasets. The meteorological forcing dataset provides monthly relative humidity, specific humidity, surface wind speed, maximum and minimum air temperature, downward longwave and shortwave radiation, snowfall, surface air pressure, and total precipitation. Lake surface area is included as an additional static predictor. The target-state dataset provides lake ice thickness, snow depth, snow cover, and lake mixing-layer temperature, while a companion lake-surface dataset provides the lake ice-cover fraction. Before training, ice thickness and snow depth are converted from meters to centimeters, mixing-layer temperature is converted from kelvin to degrees Celsius and constrained to nonnegative values, and ice-cover fraction is converted from a fraction to a percentage. The predictor variables are normalized using statistics calculated across the selected lakes and time steps. Snowice_model_outputs: Contains the NetCDF outputs generated by xLSIM. For each basin, xLSIM produces a file containing observed and predicted lake-state variables for the training, validation, and testing periods. The modeled variables include lake ice thickness, snow depth, snow cover, mixing-layer temperature, and lake ice-cover fraction. For basins without a sufficiently persistent snow-and-ice signal, the emulator predicts only mixing-layer temperature. The outputs also include training and validation loss histories, the selected model configuration, identifiers of the lakes used in training, and SHAP-based feature-importance information at the global, lake, and seasonal-regime levels. The trained machine-learning model weights are provided separately within the dataset archive. Together, these files support model evaluation and subsequent coupling with the Xanthos-Lake water-balance framework. XanthosLAKES: Contains the NetCDF input data used by the Xanthos-Lake framework. Monthly meteorological inputs include relative and specific humidity, downward shortwave and longwave radiation, mean, maximum, and minimum air temperature, wind speed, precipitation, snowfall, and surface air pressure. Static lake-property datasets provide lake identifiers, geographic locations, surface area, volume, mean depth, elevation, drainage area, fetch, outlet-routing information, and associated Xanthos grid-cell attributes. Separate bathymetric datasets provide the coefficients of the area–depth and volume–depth relationships for each aggregated lake unit. GLEV-based records provide observed lake surface area and evaporation data used to initialize lake states, define reference conditions, and calibrate and evaluate the model. Xanthos-Lake Outputs: Contains the basin- and lake-category-specific NetCDF outputs generated by Xanthos-Lake. Monthly variables include lake surface area, storage volume, outlet discharge, evaporation rate, evaporation volume, lake–groundwater exchange, lake inflow, ice thickness, snow depth, snow-cover fraction, ice-cover fraction, and mixing-layer temperature. The files also contain lake-specific calibration and validation statistics, including normalized root-mean-square error, mean absolute error, Nash–Sutcliffe efficiency, Kling–Gupta efficiency, and percent bias. Stored calibrated and derived parameters include the weir discharge coefficient, fractional freeboard, groundwater exchange coefficient, reference water level, corresponding reference surface area and storage volume, weir-width adjustment factor, and the fraction of routed inflow entering the lake. Basin identifiers, lake category, simulation period, calibration and validation periods, and parameter-schema information are retained as NetCDF metadata.

Abeshu, Guta [Pacific Northwest National Laborator↗