Search NASA⌕ Search

SEARCH · Search NASA

Results for “Streams”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Human activities shape global patterns of decomposition rates in rivers

Rivers and streams contribute to global carbon cycling by decomposing immense quantities of terrestrial plant matter. However, decomposition rates are highly variable and large-scale patterns and drivers of this process remain poorly understood. Using a cellulose-based assay to reflect the primary constituent of plant detritus, we generated a predictive model (81% variance explained) for cellulose decomposition rates across 514 globally distributed streams. A large number of variables were important for predicting decomposition, highlighting the complexity of this process at the global scale. Predicted cellulose decomposition rates, when combined with genus-level litter quality attributes, explain published leaf litter decomposition rates with high accuracy (70% variance explained). Finally, our global map provides estimates of rates across vast understudied areas of Earth and reveals rapid decomposition across continental-scale areas dominated by human activities.

54 ENVIRONMENTAL SCIENCES↗

Leptothrix ochracea genomes reveal potential for mixotrophic growth on Fe(II) and organic carbon

ABSTRACT Leptothrix ochracea creates distinctive iron-mineralized mats that carpet streams and wetlands. Easily recognized by its iron-mineralized sheaths, L. ochracea was one of the first microorganisms described in the 1800s. Yet it has never been isolated and does not have a complete genome sequence available, so key questions about its physiology remain unresolved. It is debated whether iron oxidation can be used for energy or growth and if L. ochracea is an autotroph, heterotroph, or mixotroph. To address these issues, we sampled L. ochracea -rich mats from three of its typical environments (a stream, wetlands, and a drainage channel) and reconstructed nine high-quality genomes of L. ochracea from metagenomes. These genomes contain iron oxidase genes cyc2 and mtoA, showing that L. ochracea has the potential to conserve energy from iron oxidation. Sox genes confer potential to oxidize sulfur for energy. There are genes for both carbon fixation (RuBisCO) and utilization of sugars and organic acids (acetate, lactate, and formate). In silico stoichiometric metabolic models further demonstrated the potential for growth using sugars and organic acids. Metatranscriptomes showed a high expression of genes for iron oxidation; aerobic respiration; and utilization of lactate, acetate, and sugars, as well as RuBisCO, supporting mixotrophic growth in the environment. In summary, our results suggest that L. ochracea has substantial metabolic flexibility. It is adapted to iron-rich, organic carbon-containing wetland niches, where it can thrive as a mixotrophic iron oxidizer by utilizing both iron oxidation and organics for energy generation and both inorganic and organic carbon for cell and sheath production. IMPORTANCE Winogradsky's observations of L. ochracea led him to propose autotrophic iron oxidation as a new microbial metabolism, following his work on autotrophic sulfur-oxidizers. While much culture-based research has ensued, isolation proved elusive, so most work on L. ochracea has been based in the environment and in microcosms. Meanwhile, the autotrophic Gallionella became the model for freshwater microbial iron oxidation, while heterotrophic and mixotrophic iron oxidation is not well-studied. Ecological studies have shown that Leptothrix overtakes Gallionella when dissolved organic carbon content increases, demonstrating distinct niches. This study presents the first near-complete genomes of L. ochracea , which share some features with autotrophic iron oxidizers, while also incorporating heterotrophic metabolisms. These genome, metabolic modeling, and transcriptome results give us a detailed metabolic picture of how the organism may combine lithoautotrophy with organoheterotrophy to promote Fe oxidation and C cycling and drive many biogeochemical processes resulting from microbial growth and iron oxyhydroxide formation in wetlands.

59 BASIC BIOLOGICAL SCIENCES↗

eCounter: Inline Per-IP Network Monitoring at Millisecond Resolution via eBPF

Scientific data acquisition (SciDAQ) systems are shifting from archive-based workflows to streaming paradigms, where real-time, fine-grained network monitoring becomes essential. While P4-enabled devices offer per-packet in-band observability, they require specialized switches and routers. Host-side tools like Prometheus exporters lack sufficient temporal granularity. To bridge this gap, we present eCounter, a lightweight, hardware-agnostic, inline telemetry agent built on extended Berkeley Packet Filter (eBPF). eCounter captures per-interface ingress and egress traffic, categorized by IP address and protocol, at millisecond to sub-millisecond resolution. In a 100 Gbps environment, it continuously exports up to 3,257 time-series bins per second with only 4% CPU utilization at a 35¿KiB/s data rate. We evaluate eCounter across diverse NIC MTU settings, hook types, CPU architectures and operating systems, and observed negligible impact on concurrent high-throughput streaming applications. Complexity analysis confirms that it can be readily scaled to distributed SciDAQ deployments.

Mei, Xinxin [Computational Sciences and Technology↗

Accelerating Advanced Light Source Science Through Multi-Facility HPC Workflows

Synchrotron light sources support a wide array of techniques to investigate materials, often producing complex, high-volume data that challenge traditional workflows. At the Advanced Light Source (ALS), we developed infrastructure to move microtomography data over ESnet to ALCF and NERSC, where CPU- and GPU-based algorithms generate 3D reconstructed volumes of experimental samples. We employ two data movement and reconstruction models: real-time processing as data streams directly to NERSC compute nodes, and automated file transfer to NERSC and ALCF file systems. The streaming pipeline provides users with feedback in under ten seconds, while the file-based workflow produces high-quality reconstructions suitable for deeper analysis in 20-30 minutes. This infrastructure enables users to utilize HPC resources without direct access to backend systems. We plan to extend this architecture to more endstations, supporting our beamline scientists and users.

Abramov, David↗

Data Assimilation for Robust UQ Within Agent-Based Simulation on HPC Systems

Agent-based simulation provides a powerful tool for in silico system modeling. However, these simulations do not provide built-in methods for uncertainty quantification (UQ). Within these types of models a typical approach to UQ is to run multiple realizations of the model then compute aggregate statistics. This approach is limited due to the compute time required for a solution. When faced with an emerging biothreat, public health decisions need to be made quickly and solutions for integrating near real-time data with analytic tools are needed. We propose an integrated Bayesian UQ framework for agent-based models based on sequential Monte Carlo sampling. Given streaming or static data about the evolution of an emerging pathogen this Bayesian framework provides a distribution over the parameters governing the spread of a disease through a population. These estimates of the spread of a disease may be provided to public health agencies seeking to abate the spread. By coupling agent-based simulations with Bayesian modeling in a data assimilation, our proposed framework provides a powerful tool for modeling dynamical systems in silico. We propose a method which reduces model error and provides a range of realistic possible outcomes. Moreover, our method addresses two primary limitations of ABMs: the lack of UQ and an inability to assimilate data. Our proposed framework combines the flexibility of an agent-based model with UQ provided by the Bayesian paradigm in a workflow which scales well to HPC systems. We provide algorithmic details and results on a simulated outbreak with both static and streaming data.

Spannaus, Adam [ORNL] (ORCID:0000000225213657)↗

The LCLStream Ecosystem for Multi-Institutional Dataset Exploration

We describe a new end-to-end experimental data streaming framework designed from the ground up to support new types of applications – AI training, extremely high-rate X-ray time-of-flight analysis, crystal structure determination with distributed processing, and custom data science applications and visualizers yet to be created. Throughout, we use design choices merging cloud microservices with traditional HPC batch execution models for security and flexibility. This project makes a unique contribution to the DOE Integrated Research Infrastructure (IRI) landscape. By creating a flexible, API-driven data request service, we address a significant need for high-speed data streaming sources for the X-ray science data analysis community. With the combination of data request API, mutual authentication web security framework, job queue system, high-rate data buffer, and complementary nature to facility infrastructure, the LCLStreamer framework has prototyped and implemented several new paradigms critical for future generation experiments.

Rogers, David [ORNL] (ORCID:0000000251871768)↗

Integrated structural dynamics uncover new modes of B12 photoreceptor activation

**SACLA** A crystallographic pump-power titration was first carried-out with pump laser fluences of 12, 30, 60 and 120 uJ.cm-2 and a time delay of 3 us. Then, a time-series was performed with a pump-laser fluence of 30 uJ.cm-2 (~2.4 absorbed photon per chromophore) and time-delays of 10 ns, 300 ns, 3 µs, 100 us and 3 ms. Finally, two time-delays (10 ns and 3 us) were collected with a pump laser fluence of 12 uJ.cm-2. Raw images, crystFEL streams and merged mtzs are available for all collected dat **SwissFEL** A time series with a pump laser fluence of 30 mJ.cm-2 was performed at time delys of 3 µs and 10 ms. Raw images, crystFEL streams and merged mtzs are available for all collected datasets. Refined detector geometry for each experimental campaign is also provided (crystFEL format)

RIOS-SANTACRUZ, Ronald↗

Integrated structural dynamics uncover new modes of B12 photoreceptor activation

**SACLA** A crystallographic pump-power titration was first carried-out with pump laser fluences of 12, 30, 60 and 120 ??J.cm-2 and a time delay of 3 ??s. Then, a time-series was performed with a pump-laser fluence of 30 ??J.cm-2 (~2.4 absorbed photon per chromophore) and time-delays of 10 ns, 300 ns, 3 ??s, 100 ??s and 3 ms. Finally, two time-delays (10 ns and 3 ??s) were collected with a pump laser fluence of 12 ??J.cm-2. Raw images, crystFEL streams and merged mtzs are available for all collected dat **SwissFEL** A time series with a pump laser fluence of 30 mJ.cm-2 was performed at time delys of 3 ??s and 10 ms. Raw images, crystFEL streams and merged mtzs are available for all collected datasets. Refined detector geometry for each experimental campaign is also provided (crystFEL format)

RIOS-SANTACRUZ, Ronald↗

Distinctive Pattern of Global Warming in Ocean Heat Content

Abstract Huge heat anomalies in the atmosphere and ocean in recent years are not yet explained. Strong characteristic patterns in temperatures for upper layers of the ocean occurred from 2000 to 2023 in the presence of global warming from increasing atmospheric greenhouse gases. Here, we show that the deep tropics are warming, although sharply modulated by El Niño–Southern Oscillation events, with strong heating in the extratropics near 40°N and 40°–45°S but little heating near 20°N and 25°–30°S. The heating is most clearly manifested in zonal-mean ocean heat content and is evident in sea surface temperatures. The strongest heating is in the Southern Hemisphere, where aerosol effects are small. Estimates are made of the contributions to heating of top-of-atmosphere (TOA) radiation, atmospheric energy transports, surface fluxes of energy, and redistribution of energy by surface winds and ocean currents. The patterns of change are not directly related to TOA radiation but are evident in net surface energy fluxes and inferred ocean heat transports, underscoring their coupled origin. Changes in the atmospheric circulation through a poleward shift in ocean jet streams and storm tracks are reflected in surface wind-driven ocean Ekman transports. As well as human-induced climate change, internal natural variability is likely in play. Hence, the atmosphere and ocean currents are systematically redistributing heat from global warming, profoundly affecting local climates. Significance Statement As the climate changes, it has been difficult to discern meaningful patterns. Distinctive patterns of change have occurred in the ocean when examined as zonal averages around latitude bands. Most excess heat from global warming resides in the ocean and, since 2005, has become focused into bands near 40°N and 40°S, with little net warming in the subtropics. The strongest warming is in the Southern Hemisphere, although sea surface temperatures have increased more in the Northern Hemisphere. Changes in the atmospheric circulation through a poleward shift in the jet stream and storm tracks are primarily responsible along with corresponding changes in ocean currents. These changes are linked through surface exchanges of energy via heat, moisture, and wind stress.

Trenberth, Kevin E. [National Center for Atmospher↗

Rapid measurement of soluble xylo-oligomers using near-infrared spectroscopy (NIRS) and multivariate statistics: calibration model development and practical approaches to model optimization

Rapid monitoring of biomass conversion processes using techniques such as near-infrared (NIR) spectroscopy can be substantially quicker and less labor-, resource-, and energy-intensive than conventional measurement techniques such as gas or liquid chromatography (GC or LC) due to the lack of solvents and preparation methods, as well as removing the need to transfer samples to an external lab for analytical evaluation. The purpose of this study was to determine the feasibility of rapid monitoring of a biomass conversion process using NIR spectroscopy combined with multivariate statistical modeling, and to examine the impact of (1) subsetting the samples in the original dataset by process location and (2) reducing the spectral range used in the calibration model on model performance. We develop multivariate calibration models for the concentrations of soluble xylo-oligosaccharides (XOS), monomeric xylose, and total solids at multiple points in a biomass conversion process which produces and then purifies XOS compounds from sugar cane bagasse. A single model using samples from multiple locations in the process stream showed acceptable performance as measured by standard statistical measures. However, compared to the single model, we show that separate models built by segregating the calibration samples according to process location show improved performance. We also show that combining an understanding of the sample spectra with simple multivariate analysis tools can result in a calibration model with a substantially smaller spectral range that provides essentially equal performance to the full-range model. We demonstrate that real-time monitoring of soluble xylo-oligosaccharides (XOS), monomeric xylose, and total solids concentration at multiple points in a process stream using NIR spectroscopy coupled with multivariate statistics is feasible. Segregation of sample populations by process location improves model performance. Models using a reduced spectral range containing the most relevant spectral signatures show very similar performance to the full-range model, reinforcing the importance of performing robust exploratory data analysis before beginning multivariate modeling.

09 BIOMASS FUELS↗

Odor exposure during imprinting periods increases odorant-specific sensitivity and receptor gene expression in coho salmon ( Oncorhynchus kisutch )

ABSTRACT Pacific salmon are well known for their homing migrations; juvenile salmon learn odors associated with their natal streams prior to seaward migration, and then use these retained odor memories to guide them back from oceanic feeding grounds to their river of origin to spawn several years later. This memory formation, termed olfactory imprinting, involves (at least in part) sensitization of the peripheral olfactory epithelium to specific odorants. We hypothesized that this change in peripheral sensitivity is due to exposure-dependent increases in the expression of odorant receptor (OR) proteins that are activated by specific odorants experienced during imprinting. To test this hypothesis, we exposed juvenile coho salmon, Oncorhynchus kisutch, to the basic amino acid odorant l-arginine during the parr–smolt transformation (PST), when imprinting occurs, and assessed sensitivity of the olfactory epithelium to this and other odorants. We then identified the coho salmon ortholog of a basic amino acid odorant receptor (BAAR) and determined the mRNA expression levels of this receptor and other transcripts representing different classes of OR families. Exposure to l-arginine during the PST resulted in increased sensitivity to that odorant and a specific increase in BAAR mRNA expression in the olfactory epithelium relative to other ORs. These results suggest that specific increases in ORs activated during imprinting may be an important component of home stream memory formation and this phenomenon may ultimately be useful as a marker of successful imprinting to assess management strategies and hatchery practices that may influence straying in salmon.

Dittman, Andrew H. (ORCID:000000016482359X)↗

Dayflow-PR: High-Resolution Streamflow Reanalysis for Puerto Rico, Version 1.0

This dataset presents a high-resolution historical streamflow reanalysis for NHDPlusV2 stream reaches across Puerto Rico (PR) spanning 1950 - 2019. The reanalysis is generated using the calibrated VIC-RAPID hydrologic modeling framework at the Hydrologic Unit Code Sub-basin (HUC08) scale, forced with sub-daily and daily meteorological forcings from Daymet. Runoff is simulated on 1- and 6-km grids, and the resulting total runoff is routed through the NHDPlusV2 river network using the RAPID routing model to produce Naturalized Streamflow Reanalysis. Where complete observational records are available over 1980 - 2019, streamflows are assimilated (substituted) and subsequently routed downstream through the river network to produce Assimilated Streamflow Reanalysis. The dataset includes streamflow outputs from eight distinct hydrologic modeling configurations along with key performanc evaluation metrics at daily and monthly scales, supporting a wide range of water resource applications. This dataset is derived to support the Non-Powered Dam Assessment, as well as 9505 Secure Water Assessment projects for the US Department of Energy (DOE) Water Power Technologies Office (WPTO). For further details, refer to Ghimire et al. (2023), Kao et al. (2024), and Ghimire et al. (2025).

13 HYDRO ENERGY↗

Stable Water Isotope Data for the East River Watershed, Colorado (2014-2025)

The stable water isotope data for the East River Watershed, Colorado, consists of delta2H (hydrogen) and delta18O (oxygen) values from samples collected at multiple, long-term monitoring sites including streams, groundwater wells, springs, and a precipitation collector used to establish a local meteoric water line (LMWL) for the watershed. These locations represent important and/or unique end-member locations for which stable isotope values can be diagnostic of the connection between precipitation inputs as snow and rain and riverine export. Such locations include drainages underline entirely or largely by shale bedrock, land covered dominated by conifers, aspens, or meadows, and drainages impacted by historic mining activity and the presence of naturally mineralized rock. Developing a long-term record of water isotope values from a diversity of environments is a critical component of quantifying the impacts of both climate change and discrete climate perturbations, such as drought, forest mortality, and wildfire, on water export. Such data may be combined with stream gaging stations co-located at each surface water monitoring site to relate seasonal variations in water export to their stable isotopic signature. Data for liquid water delta2H and delta18O values are reported in units of parts per thousand (per-mil; ‰). This data package contains (1) a zip file (isotope_data_2014-2025.zip) containing a total of 95 files: 96 data files of isotope data from across the Lawrence Berkeley National Laboratory (LBNL) Watershed Function Scientific Focus Area (SFA) which is reported in .csv files per location and a locations.csv (1 file) with latitude and longitude for each location; (2) a file-level metadata (v6_20260901_flmd.csv) file that lists each file contained in the dataset with associated metadata; and (3) a data dictionary (v6_20260901_dd.csv) file that contains terms/column_headers used throughout the files along with a definition, units, and data type. Missing values within the anion data files are noted as either "-9999" or "0.0" for not detectable (N.D.) data. There are a total of 43 locations containing isotope data. Update on 2022-06-10: versioned updates to this dataset was made along with these changes: (1) updated isotope data for all locations up to 2021-12-31 and (2) the addition of the file-level metadata (flmd.csv) and data dictionary (dd.csv) were added to comply with the File-Level Metadata Reporting Format. Update on 2022-09-09: Updates were made to reporting format specific files (file-level metadata and data dictionary) to correct swapped file names, add additional details on metadata descriptions on both files, add a header_row column to enable parsing, and add version number and date to file names (v2_20220909_flmd.csv and v2_20220909_dd.csv). Update on 2023-08-08: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-03-13. The file level metadata and data dictionary files were updated to reflect the additional data added. Update on 2024-03-11: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2024-02-19. Further, revisions to the data files were made to remove incorrect data points (from 1970 and 2001). The reporting format specific files were updated to reflect the additional data added. Update on 2025-05-15: Updates were made to both the data files and reporting format specific files. New available isotope data was added, up until the end of WY2024 (September 30, 2024). International Generic Sample Numbers (IGSNs), when registered, were added to the data files. The reporting format specific files were updated to reflect the additional data added. Update on 2026-09-01: Updates were made to both the data files and reporting format specific files. New available isotope data was added, up until the end of WY2025 (September 30, 2025).

54 ENVIRONMENTAL SCIENCES↗

Anion Data for the East River Watershed, Colorado (2014-2025)

The anion data for the East River Watershed, Colorado, consist of fluoride, chloride, sulfate, nitrate, and phosphate concentrations collected at multiple, long-term monitoring sites that include stream, groundwater, and spring sampling locations. These locations represent important and/or unique end-member locations for which solute concentrations can be diagnostic of the connection between terrestrial and aquatic systems. Such locations include drainages underlined entirely or largely by shale bedrock, land covered dominated by conifers, aspens, or meadows, and drainages impacted by historic mining activity and the presence of naturally mineralized rock. Developing a long-term record of solute concentrations from a diversity of environments is a critical component of quantifying the impacts of both climate change and discrete climate perturbations, such as drought, forest mortality, and wildfire, on the riverine export of multiple anionic species. Such data may be combined with stream gauging stations co-located at each monitoring site to directly quantify the seasonal and annual mass flux of these anionic species out of the watershed. This data package contains (1) a zip file (anion_data_2014_2025.zip) containing a total of 386 files: 387 data files of anion data from across the Lawrence Berkeley National Laboratory (LBNL) Watershed Function Scientific Focus Area (SFA) which is reported in .csv files per location and a locations.csv (1 file) with latitude and longitude for each location; (2) a file-level metadata (v7_20260901_flmd.csv) file that lists each file contained in the dataset with associated metadata; (3) a data dictionary (v7_20260901_dd.csv) file that contains terms/column_headers used throughout the files along with a definition, units, and data type; and (4) a anion MDL fact sheet (anion_MDLs_202608 in PDF and docx formats). Missing values within the anion data files are noted as either "-9999" or "0.0" for not detectable (N.D.) data. There are a total of 47 locations containing anion data. Update on 2022-06-10: versioned updates to this dataset was made along with these changes: (1) updated anion data for all locations up to 2021-12-31, (2) removal of units from column headers in datafiles, (3) added row underneath headers to contain units of variables, (4) restructure of units to comply with CSV reporting format requirements, and (5) the addition of the file-level metadata (flmd.csv) and data dictionary (dd.csv) were added to comply with the File-Level Metadata Reporting Format. Update on 2022-09-09: Updates were made to reporting format specific files (file-level metadata and data dictionary) to correct swapped file names, add additional details on metadata descriptions on both files, add a header_row column to enable parsing, and add version number and date to file names (v2_20220909_flmd.csv and v2_20220909_dd.csv). Update on 2022-12-20: Updates were made to both the data files and reporting format specific files. Conversion issues affecting ER-PLM locations for anion data was resolved for the data files. Additionally, the flmd and dd files were updated to reflect the updated versions of these files. Available data was added up until 2022-03-14. Update on 2023-08-08: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-05-19. The file level metadata and data dictionary files were updated to reflect the additional data added. Update on 2024-03-11: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-09-11. Further, revisions to the data files were made to remove incorrect data points (from 1970 and 2001). The reporting format specific files were updated to reflect the additional data added. Update on 2025-05-15: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until the end of WY2024 (September 30, 2024). International Generic Sample Numbers (IGSNs), when registered, were added to the data files. The reporting format specific files were updated to reflect the additional data added. Update on 2026-09-01: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until the end of WY2025 (September 30, 2025). An anion MDL document was included in this update.

54 ENVIRONMENTAL SCIENCES↗

Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification" Willard et al. (2025).

This data release provides all data and code used in the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025)" to model stream temperature, evaluate, and assess results. The associated manuscript explores the effect of different ensemble construction techniques across different common machine learning (ML) architectures for predictions in unmonitored basins. Modeling was done using long short-term memory (LSTM), gated recurrent unit (GRU), temporal convolution network (TCN), and extreme gradient boosting (XGBoost) models, and stream site coverage spans 1362 locations across the conterminous United States. The ensemble construction techniques investigated include ensemble by random weight initialization, differing hyperparameters, different random subsets of training data, different subselections of input features, different architectures, and Monte Carlo Dropout. The data is organized into these items items:Code repository and data for the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025).Code: stream_temp_ml_regionalization.zip contains the code repositoryData to run the code:- data_dir.zip -- contains all files that should be moved to the "DATA_DIR" variable defined in the "set_env_vars.sh" script in the code repository- metadata_dir.zip -- contains all files that should be moved to the "METADATA_DIR" variable defined in the "set_env_vars.sh" script in the code repositoryData produced by the code and used in the paper:- outputs_dir.zip - contains model output and results (outputs_dir/results), model weights (outputs_dir/models), and all other outputs used for the paper including feature importances.To cite this code, please use the following BibTeX or MLA entries:bibtex:@misc{willard2025streamensembles,author = {Jared Willard and Charuleka Varadharajan},title = {Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification"},year = {2024},doi = {10.15485/2527393},publisher = {ESS-DIVE Repository},url = {https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2527393}}MLA: Willard, Jared, et al. Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification". 2025. ESS-DIVE Repository, doi:10.15485/2448016.

54 ENVIRONMENTAL SCIENCES↗

Mountain Basin Controls on the Snow-to-Streamflow Signal: An AIC-Weighted Multiple Linear Regression Framework

A regression-based analysis quantifies how basin characteristics modulate the snow-to-streamflow signal. First, we use the ERA5-Land reanalysis gridded product (European Centre for Medium Range Weather Forecasts reanalysis 5 -Land component) for 4,655 hydrologic unit code - 10 (HUC10) mountain basins across the western United States (US) for water years 1987–2024. Linear regressions are performed for peak snow water equivalent (SWE) and annual streamflow for each mountain basin. Models use ordinary least squares in Python’s statsmodels package. After which, an Akaike Information Criterion (AIC)–weighted ensemble multiple linear regression (MLR) framework with 47 watershed traits is used to predict the linear regression coefficient of determination (r-squared) defining the ability of peak SWE to predict annual streamflow across all mountain basin. Predictor sets are constrained to avoid multicollinearity by excluding models with variance inflation factors (VIF) greater than 5. Mountain basin traits included in the MLR include seasonal climate, topography, vegetation type and structure, and bedrock geology. Accepted models are considered if their AIC is within 2.0 of the model with the minimum AIC, or best model. To compare predictor influence across acceptable models, we computed standardized regression coefficients. To evaluate structural redundancy among models, we constructed binary inclusion vectors for each acceptable model, denoting whether a predictor was present (1) or absent (0). Core predictor variables are defined as occurring in at least 67% of the acceptable models. For this regional analysis, only one model was found acceptable, with higher snow-to-streamflow translation (higher r-squared) occurring in colder mountain basins with higher relative winter precipitation, more snow accumulation and a lower fraction of annual precipitation that falls in the spring and summer. The second component of the data package uses previously published, high-resolution output from an integrated hydrological model of the East River watershed using the U.S. Geological Survey Groundwater and Surface water Flow model (GSFLOW, doi:10.15485/1998576). East River MLR expands upon the approach described above to explore the response of five streamflow metrics—annual streamflow, runoff efficiency, 7-day minimum flow, low-flow duration, and non-perennial stream fraction to snow system indicators including peak SWE, snow-covered area, snow disappearance date, and the fraction of basin area characterized by low-to-no snow, as well as seasonal precipitation and temperature, and annual hydrologic variables representing soil moisture, evapotranspiration (ET), the partitioning of incoming precipitation to evapotranspiration (ET/P), groundwater storage, and groundwater inflow to streams. MLR was done on all water years (P0: 1987-2024) and for each period as determined in the split analysis using pooled regression techniques (P1: 1987-2011 and P2: 2012-2024) to evaluate shifting predictor variable emphasis on streamflow generation. Results indicate that since 2012, peak SWE has lost statistical strength in its prediction of annual streamflow and runoff efficiency, and the indirect influence of spring temperature has emerged as critically important. Low-flow metrics remain largely influenced by soil moisture, vegetation water use and groundwater inflows with summer precipitation becoming a direct influence on minimum summer flow. Together, these data and Python-based analysis tools provide a framework for identifying the key watershed characteristics that control how streamflow responds to snow from year to year. The package also helps quantify uncertainty in statistical models and assess how snow–streamflow relationships vary across regions and over time. This dataset contains comma-separated values files (.csv), text files (.txt), python code files (.py), figure files (.png), and shapefiles (.cpg, .dbf, .prj, .sbn, .sbx, .shp, .xml). Further details on file contents and MLR execution can be found in the readme file and the FLMD files. Work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Historic climate, cosmogenic 10Be, denudation-rate, and geospatial datasets from the Pikes Peak region, Colorado, USA

This data package contains geographic information system (GIS) layers and tabular datasets associated with the study of elevation-dependent denudation rates on Pikes Peak in the Front Range of the Rocky Mountains, Colorado, USA. The package includes GIS layers used to produce the study-area map, including sample locations, sample watershed boundaries, the Pikes Peak batholith, Pleistocene glacier extent, weather station locations, and elevation and hillshade rasters, together with comma-separated value (CSV) tables and matching CSV data dictionaries. These mapped layers provide the geographic framework for interpreting denudation patterns across the Pikes Peak region and for relating sample locations to watershed geometry, bedrock setting, glacial history, and nearby climate stations. The first group of tables reports climate and geospatial context for the study area. These files include station-based temperature and precipitation data used to characterize elevational gradients in mean annual climate and monthly climate seasonality, sample locations, denudation-rate and topographic metrics, fixed frost-cracking model parameters, frost-cracking intensity and precipitation-frequency metrics, and stream-power inversion results. Together, these data provide the basis for evaluating how denudation varies with elevation, climate, and landscape form across sampled catchments on Pikes Peak. The second group of tables reports cosmogenic nuclide and erosion-model results used in the denudation analysis. Included files contain accelerator mass spectrometry (AMS) measurements for in situ-produced cosmogenic beryllium-10 (10Be), including sample identifiers, measured 10Be:9Be ratios, analytical uncertainties, carrier mass, quartz mass, blank corrections, blank-group statistics, and calculated 10Be concentrations and uncertainties. Additional tables summarize stream-power-law inversion results for sampled catchments, including optimized model parameters, predicted erosion rates, residual metrics, channel-pixel counts, and convergence status, as well as regression equations and summary statistics used to evaluate relationships among elevation, climate, frost cracking, precipitation forcing, and denudation rate. The package contains GIS files, comma-separated value files (.csv), Microsoft Excel files (.xlsx), CSV data dictionaries, a file-level metadata table, and a readme text file.

10Be cosmogenic nuclides↗

pH-Driven Restructuring of Hydration Layers and Cation Ad-sorption at the Alumina-Water Interface

Oxide-water interfaces underpin ion separation, catalysis, and electrochemical energy technologies, where the electrical double layer (EDL) controls adsorption, transport, and reactivity. Yet, the molecular-scale link between pH-dependent surface protonation, hydration-layer structure, and counter-ion adsorption remains poorly defined. Here, we combine in situ crystal truncation rod (CTR) and resonant anomalous X-ray reflectivity (RAXR) with streaming potential measurements and ab initio molecular dynamics (AIMD) simulations to resolve the chemical and structural evolution of the EDL at the single-crystal alumina (012)-water interface in 10 mM Rb+ over pH 3-12. CTR measurements reveal two distinct adsorbed water layers at ~2.2 and ~3.5 Å above the surface that each shift toward the substrate at transition pHs near 6.5 and 10.6, respectively, directly reflecting changes in primary hydration layer structure in response to the deprotonation of bridging and terminal aluminol groups. RAXR shows a 10-fold increase in Rb+ coverage and a decrease in mean adsorption height from ~3.5 to ~2.7 Å with increasing pH, indicating enhanced counter-ion binding accompanied by Stern layer contraction. Streaming potential measurements demonstrate that the zeta potential, i.e., potential at the hydrodynamic shear plane, is positive at pH 3 and becomes negative at pH ≥3.5. This negative charge magnitude increases with increasing pH, consistent with progressive surface deprotonation at higher pH. AIMD identifies inner- and outer-sphere Rb+ complexes whose adsorption heights and coordination geometries depend sensitively on the protonation state of surface oxygens, providing atomistic support for the experimentally inferred trends. These measurements establish two discrete, site-specific pH transitions in hydration-layer structure that track aluminol (de)protonation and quantitatively link them to a pH-driven contraction of the Stern layer (increasing Rb+ coverage and decreasing adsorption height). This provides a direct structural basis for connecting surface acid-base chemistry to ion binding distances at an oxide-water interface.

Electrical double layer (EDL), Surface protonation↗