Search NASA⌕ Search

Engineering topics

Stegen, James C.

Publications and source records attributed to Stegen, James C..

At least 19 records

A functional microbiome catalogue crowdsourced from North American rivers

Predicting elemental cycles and maintaining water quality under increasing anthropogenic influence requires knowledge of the spatial drivers of river microbiomes. However, understanding of the core microbial processes governing river biogeochemistry is hindered by a lack of genome-resolved functional insights and sampling across multiple rivers. Here we used a community science effort to accelerate the sampling, sequencing and genome-resolved analyses of river microbiomes to create the Genome Resolved Open Watersheds database (GROWdb). GROWdb profiles the identity, distribution, function and expression of microbial genomes across river surface waters covering 90% of United States watersheds. Specifically, GROWdb encompasses microbial lineages from 27 phyla, including novel members from 10 families and 128 genera, and defines the core river microbiome at the genome level. GROWdb analyses coupled to extensive geospatial information reveals local and regional drivers of microbial community structuring, while also presenting foundational hypotheses about ecosystem function. Building on the previously conceived River Continuum Concept, we layer on microbial functional trait expression, which suggests that the structure and function of river microbiomes is predictable. We make GROWdb available through various collaborative cyberinfrastructures, so that it can be widely accessed across disciplines for watershed predictive modelling and microbiome-based management practices.

59 BASIC BIOLOGICAL SCIENCES↗

Laboratory time series moisture manipulative experiment from sediment across the contiguous US: time series aerobic respiration and geochemistry (v2)

This dataset supports a broader study examining the effects of wetting and drying on hyporheic zone respiration across the contiguous United States (CONUS). The dataset provides data generated from a laboratory moisture manipulation experiment. The contents include time series aerobic respiration and moisture; dissolved oxygen; sediment geochemistry data; and field metadata (including qualitative information on instream and river corridor characteristics). Samples were collected as part of the WHONDRS CONUS-Scale Model-Sample Study (CM). This study was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the CONUS. The data package associated with the CM study is available at https://data.ess-dive.lbl.gov/view/doi:10.15485/1923689. CM sampling began in April 2022 and ended in October 2023. This study uses subsamples from a subset of CM samples collected between June 2022 and June 2023. The original field samples were labeled as CM_###. Subsequent subsamples for this study were labeled as EC_###. The labels from the field samples and the EC subsamples can be mapped directly based on the digits following the prefix and underscore (i.e., EC_001 is a subsample from CM_001). See the critical details section below for more details on sample naming. This data package was originally published in August 2024. It was updated in February 2026 (v2; new and modified files). See the change history section in the readme for more details. For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) field protocol; and a (6) a subfolder with sediment sample data from the incubation experiment. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC); (2) total nitrogen (TN); (3) adenosine triphosphate (ATP); (4) percent carbon and nitrogen; (5) effect size; (6) iron (II); (7) gravimetric moisture; (8) respiration rates and raw dissolved oxygen values; (9) specific conductance; (10) pH; (11) temperature; (12) a summary containing median values of each data type for each treatment (wet and dry); (13) methods codes; (14) FTICR-MS methods; and (15) a subfolder of 9.4 Tesla FTICR-MS data. This folder contains three subfolders, one containing the sediment .xml data files, one containing the sediment CoreMS output files, the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .ref, or .xml.

54 ENVIRONMENTAL SCIENCES↗

Prediction of Distributed River Sediment Respiration Rates Using Community-Generated Data and Machine Learning

River sediment microbial respiration is a key indicator of ecosystem functioning and the biogeochemical fluxes across this critical zone link surface and subsurface waters. As such, there is tremendous interest in measuring and mapping these respiration rates. Respiration observations are expensive and labor intensive; there is limited data available to the community. An open science, collaborative initiative is collecting samples for respiration rate analysis and multi-scale metadata; this evolving data set is being used for making machine learning (ML) predictions at unsampled sites to help inform continued community engagement. However, it is a challenge to find an optimum configuration for ML models to work with this feature-rich (i.e., 100+ possible input variables) data set. Here, we present results from a two-tiered approach to managing the analysis of this complex data set: (a) a stacked ensemble of models that automatically optimizes hyperparameters and manages the training of many models and (b) feature permutation importance to detect the most important features in the models. The major elements of this workflow are modular, portable, open, and cloud-based thus making this implementation a potential template for other applications. The models developed here predict that sediment organic matter chemistry is one of the most important features for predicting sediment respiration rate. Other larger-scale, important features fall into the categories of climatic, ecological, geological, and fluvial settings. Leveraging these larger-scale features to generate data-driven estimates of river sediment respiration rates reveals spatially consistent but heterogeneous patterns across the river network of the Columbia River Basin.

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with the manuscript "Organic Molecules are Deterministically Assembled in River Sediments"

This data package is associated with the publication "Organic Molecules are Deterministically Assembled in River Sediments" submitted to Scientific Reports (Stegen et al., 2024). The study applies community ecology methods to dissolved organic matter (DOM) chemistry from variably inundated riverbed sediments to uncover principles governing DOM composition at a reach-scale. This data package documents the workflow used to process and generate the main findings in the manuscript. The R scripts reference the raw, unprocessed Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data from another data package, available on ESS-DIVE at https://data.ess-dive.lbl.gov/view/doi:10.15485/1834208. The scripts then process the raw FTICR-MS data and generate the findings and figures presented in the associated manuscript. In brief, this study demonstrates that DOM assemblages in variably inundated sediments are primarily governed by deterministic variable selection, including sediment moisture effecting the degree of deterministic assembly. See the manuscript for more details pertaining to interpretation and implications of the findings. This data package is associated with the GitHub repository found at https://github.com/WHONDRS-Hub/ECA_2020_Sed.This data package is comprised of 6 scripts and 7 folders. The file-level metadata file (file ending in "flmd.csv") lists all files contained in this data package and descriptions for each. The data dictionary (file ending in "dd.csv) describes all tabular data columns and their respective definitions and units. The FTICR_Processing_Scripts produce the outputs found in the "Processed_Data" folder. The remaining scripts (located in the parent directory) produce the outputs found in the following four folders: (1) "MCD_Dendrograms", "MCD_Randomizations", "MCD_bNTI_Outcomes", and "OM_Null_Modeling". The fifth script additionally takes the three comma-separated values (CSV) files found in the parent directory as input ("VGC_texture.csv", "merged_weights.csv", and "ECA2_FTICR_BetaDisp.csv"). The outputs of each of the five scripts serve as the input to the following script, with the final outputs stored in the folder "OM_Null_Modeling".

54 ENVIRONMENTAL SCIENCES↗

Linkages Between Mineral Element Composition of Soils and Sediments With Hyporheic Zone Dissolved Organic Matter Chemistry Across the Contiguous United States

The hyporheic zone is a hotspot for biogeochemical cycling where interactions with mineral metals preserve the release and biodegradation of organic matter (OM). A small fraction of OM can still be exchanged between localized sediments and the overlying water column, and recent evidence suggests there exists a longitudinal structuring in sediment dissolved OM (DOM) chemistry across the continental United States (CONUS). In this study, we tested a hypothesis that water extractable sediment DOM chemistry could be explained by sediment metal contents and integrative watershed scale features at the CONUS scale. Crowdsourced samples were characterized for high resolution mass spectrometry and coupled with sediment metals determined via x-ray fluorescence as well as with land cover and soil elemental information obtained from national databases. Our results highlight weak relationships between DOM chemistry and elemental composition at the CONUS scale indicating limited transferability of organo-metal linkages into multi-scale hydrobiogeochemical models.

58 GEOSCIENCES↗

Timeseries Unlabeled and Labeled Photos, Modeled Stream Elevation, and (Meta)Data of Variably Inundated Streams Across The Yakima River Basin, Washington, United States (v2)

This dataset is associated with the “River Monitoring Photos” (RMP) study and subsequent manuscript (Bao et al. 2025. Monitoring river flow status using low-cost wildlife camera and image segmentation artificial intelligence doi: 10.1016/j.envsoft.2025.106715). Game camera timeseries photos were collected to evaluate stream variable inundation via changes in width. A subset of photos was labeled for training the YOLOv8 and Mask2Former models and used to segment water surface fractions from all the game camera photos.This data package was originally published in March 2024. It was updated in October 2025 (v2) to add additional photos and files associated with the manuscript (i.e., processed data, labeled photos, and trained models). For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.In addition to a readme, this data package also includes two file-level metadata (FLMD) files that describes each file and two data dictionaries (DD) that describe all column/row headers and variable definitions. This dataset is comprised of (1) file-level metadata; (2) data dictionary; (3) readme; (4) field metadata; and (5) folders containing game camera photos and manuscript-associated files. Each Yakima River Basin site has a folder that contains subfolders for each month photos were collected. There is also a folder for files associated with the manuscript which has subfolders for labeled data, trained models, Yakima River Basin site water surface fractions, and USGS site water surface fractions. All files are .csv, .json, .txt, .yaml, .pth, .pt, or .pdf. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Models, data, and scripts associated with “Prediction of Distributed River Sediment Respiration Rates using Community-Generated Data and Machine Learning”

This data package is associated with the publication “Prediction of Distributed River Sediment Respiration Rates using Community-Generated Data and Machine Learning’’ submitted to the Journal of Geophysical Research: Machine Learning and Computation (Scheibe et al. 2024). River sediment respiration observations are expensive and labor intensive to obtain and there is no physical model for predicting this quantity. The Worldwide Hydrobiogeochemisty Observation Network for Dynamic River Systems (WHONDRS) observational data set (Goldman et al.; 2020) is used to train machine learning (ML) models to predict respiration rates at unsampled sites. This repository archives training data, ML models, predictions, and model evaluation results for the purposes of reproducibility of the results in the associated manuscript and community reuse of the ML models trained in this project. One of the key challenges in this work was to find an optimum configuration for machine learning models to work with this feature-rich (i.e. 100+ possible input variables) data set. Here, we used a two-tiered approach to managing the analysis of this complex data set: 1) a stacked ensemble of ML models that can automatically optimize hyperparameters to accelerate the process of model selection and tuning and 2) feature permutation importance to iteratively select the most important features (i.e. inputs) to the ML models. The major elements of this ML workflow are modular, portable, open, and cloud-based, thus making this implementation a potential template for other applications. This data package is associated with the GitHub repository found at Please see the file level metadata (flmd; “sl-archive-whondrs_flmd.csv”) for a list of all files contained in this data package and descriptions for each. Please see the data dictionary (dd; “sl-archive-whondrs_dd.csv”) for a list of all column headers contained within comma separated value (csv) files in this data package and descriptions for each. The GitHub repository is organized into five top-level directories: (1) “input_data” holds the training data for the ML models; (2) “ml_models” holds machine learning models trained on the data in “input_data”; (3) “scripts” contains data preprocessing and postprocessing scripts and intermediate results specific to this data set that bookend the ML workflow; (4) “examples” contains the visualization of the results in this repository including plotting scripts for the manuscript (e.g., model evaluation, FPI results) and scripts for running predictions with the ML models (i.e., reusing the trained ML models); (5) “output_data” holds the overall results of the ML model on that branch. Each trained ML model resides on its own branch in the repository; this means that inputs and outputs can be different branch-to-branch. Furthermore, depending on the number of features used to train the ML models, the preprocessing and postprocessing scripts, and their intermediate results, can also be different branch-to-branch. The “main-*” branches are meant to be starting points (i.e. trunks) for each model branch (i.e. sprouts). Please see the Branch Navigation section in the top-level README.md in the GitHub repository for more details. There is also one hidden directory “.github/workflows”. This hidden directory contains information for how to run the ML workflow as an end-to-end automated GitHub Action but it is not needed for reusing the ML models archived here. Please the top-level README.md in the GitHub repository for more details on the automation.

13C↗

Data and scripts associated with a manuscript investigating dissolved organic matter and microbial community linkages across seven globally distributed rivers

This data package is associated with the publication “Meta-metabolome ecology reveals that geochemistry and microbial functional potential are linked to organic matter development across seven rivers” submitted to Science of the Total Environment. This data package includes the data necessary to replicate the analyses presented within the manuscript to investigate dissolved organic matter (DOM) development across broad spatial distances and within divergent biomes. Specifically, we included the Fourier transform ion cyclotron mass spectrometry (FTICR-MS) data, geochemistry data, annotated metagenomic data, and results from ecological null modeling analyses in this data package. Additionally, we included the scripts necessary to generate the figures from the manuscript. Complete metagenomic data associated with this data package can be found at the National Center for Biotechnology (NCBI) under Bioproject PRJNA946291. This dataset consists of (1) four folders; (2) a file-level metadata (flmd) file; (3) a data dictionary (dd) file; (4) a factor sheet describing samples; and (5) a readme. The FTICR Data folder contains (1) the processed Fourier transform ion cyclotron mass spectrometry (FTICR-MS) data; (2) a transformation-weighted characteristics dendrogram generated from the FTICR-MS data; and (3) the script used to generate all FTICR-MS related figures. The Geochemical Data folder contains (1) the single geochemistry data file and (2) the R script responsible for generating associated figures. The Metagenomic Data folder contains (1) annotation information across different levels; (2) carbohydrate active enzyme (CAZyme) information from the dbCAN database (Yin et al., 2012); (3) phylogenetic tree data (FASTAs, alignments, and tree file); and (4) the scripts necessary to analyze all of these data and generate figures. The Null Modeling Data folder contains (1) data generated during null modeling for each river and all rivers combined and (2) the R scripts necessary to process the data. All files are .csv, .pdf, .tsv, .tre, .faa, .afa, .tree, or .R.

54 ENVIRONMENTAL SCIENCES↗

Data and Scripts Associated with the Manuscript “Water Column Respiration in the Yakima River Basin is Explained by Temperature, Nutrients and Suspended Solids”

This data package is associated with the publication “Water Column Respiration in the Yakima River Basin is Explained by Temperature, Nutrients and Suspended Solids” published in EGU Biogeochemistry (Laan et al. 2025). In this research, water column respiration (ERwc) data, surface water chemistry data, organic matter (OM) chemistry data, and publicly available geospatial data were used in analysis to evaluate the variability in ERwc at 47 sites across the Yakima River basin in Washington, USA. In addition to this readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. The data package includes the data inputs, and outputs, and R scripts to reproduce all the analyses performed in the manuscript and create manuscript figures. The data package is comprised of three main folders (Code, Data, and Figures). The Code folder is comprised of four scripts and three analysis-specific subfolders that contain the R scripts to perform the analyses described in the publication and create publication figures. The Data folder is comprised of two “.csv” files and four subfolders that contain data input and output files. The Published_Data folder contains a readme that directs the user to download the appropriate files and add to this folder when using scripts. The Figures folder includes figures from the manuscript in “.pdf” and “.png” formats and a folder with intermediate figure files. This data package is associated with a GitHub repository which can be found at https://github.com/river-corridors-sfa/rcsfa-RC2-SPS-ERwc. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected some of these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Ultrahigh-resolution mass spectrometry data associated with the manuscript “A functional microbiome catalog crowdsourced from North American rivers"

This data package is associated with the publication “A functional microbiome catalog crowdsourced from North American rivers” submitted to Nature (Borton et al., 2024); (https://www.biorxiv.org/content/10.1101/2023.07.22.550117v1). Predicting elemental cycles and maintaining water quality under increasing anthropogenic influence requires understanding the spatial drivers of river microbiomes. However, the unifying microbial determinants governing river biogeochemistry are hindered by a lack of genome-resolved functional insights and sampling across multiple rivers. Here we employed a community science effort to accelerate the sampling of river microbiomes to create the Genome Resolved Open Watersheds database (GROWdb). GROWdb is a publicly available resource that paves the way for watershed predictive modeling and microbiome-based management practices. This resource profiled the identity, distribution, function, and expression of thousands of microbial genomes across rivers covering 90% of United States watersheds. We identified the most cosmopolitan microbiome members, while also revealing local drivers of strain endemism across ecological dimensions. We provide the first evidence that microbial functional trait expression followed the tenets of the River Continuum Concept, suggesting the structure and function of river microbiomes is predictable. The Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) data were one of many different data types used in establishing the ecological dimensions along which different microbes were detected .This data package only contains the processed FTICR-MS data associated with this manuscript; all other data is accessible via Zenodo (https://zenodo.org/records/8173287), GitHub (https://github.com/jmikayla1991/Genome-Resolved-Open-Watersheds-database-GROWdb), KBase (https://doi.org/10.25982/109073.30/1895615), and NCBI via Bioproject PRJNA946291.This dataset consists of (1) a file-level metadata (flmd) file; (2) a data dictionary (dd) file; (3) a readme; (4) three Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) processed data files (a ‘data’ file containing peak-by-sample observations, a ‘mol’ file containing peak metadata, and a transformation profile containing transformation-by-sample observations). All files are .csv or .pdf.

54 ENVIRONMENTAL SCIENCES↗

Chemodiversity of riverine dissolved organic matter: Effects of local environments and watershed characteristics

Riverine dissolved organic matter (DOM) is crucial to global carbon cycling and aquatic ecosystems. However, the geographical patterns and environmental correlates of DOM chemodiversity remain elusive especially in surface waters and sediments of global rivers. Here, we systematically analyzed DOM molecular diversity and composition in surface waters and sediments across 97 globally distributed rivers using data from the Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS) consortium. Here, we further examined the associations of molecular richness and composition with geographical, climatic, physicochemical variables, as well as the watershed characteristics. We found that molecular richness significantly decreased toward higher latitudes, but only in sediments (r = -0.24, P < 0.001). The environmental variables like precipitation and non-purgeable organic carbon showed strong associations with DOM molecular richness and composition. Interestingly, we identified that less-documented factors like watershed characteristics were also related to DOM molecular richness and composition, and were stronger in waters than sediments. Such associations were further supported by the fact that 11 out of 13 watershed characteristics (e.g., land cover variables) showed more positive than negative correlations with molecular abundance especially in waters. The molecules positively and negatively correlated with specific watershed characteristics had distinct traits, such as stoichiometric ratios and aromaticity. Land covers shifted from natural (e.g., forests) to human-modified (e.g., impervious area) were associated with systematic changes in DOM molecular chemistry. Our findings imply that it may be possible to use a small set of broadly available data types to predict DOM molecular richness and composition across diverse river systems. Elucidation of mechanisms underlying these relationships will provide further enhancements to such predictions, especially when extrapolating to unsampled systems.

54 ENVIRONMENTAL SCIENCES↗

Systems and methods for determining ground water-surface water interactions

Systems for determining GW/SW interaction are provided. The systems can include: a sensing assembly comprising sensors for pressure, fluid conductivity, temperature, and transfer resistance; processing circuitry operatively coupled to the sensing assembly and configured to receive data from the sensing assembly and process the data to provide a GW/SW interaction, wherein the data includes pressure, fluid conductivity, temperature, transfer resistance data. Methods for determining GW/SW interaction are provided. The methods can include: receiving real time data including pressure, fluid conductivity, temperature, and transfer resistance; from at least some of the data received simulating the SW/GW interaction; and fitting the real time data with the simulated data to provide actual SW/GW interaction.

Johnson, Timothy C.↗

Microdiverse bacterial clades prevail across Antarctic wetlands

Antarctica's extreme environmental conditions impose selection pressures on microbial communities. Indeed, a previous study revealed that bacterial assemblages at the Cierva Point Wetland Complex (CPWC) are shaped by strong homogeneous selection. Yet which bacterial phylogenetic clades are shaped by selection processes and their ecological strategies to thrive in such extreme conditions remain unknown. Here, we applied the phyloscore and feature-level βNTI indexes coupled with phylofactorization to successfully detect bacterial monophyletic clades subjected to homogeneous (HoS) and heterogenous (HeS) selection. Remarkably, only the HoS clades showed high relative abundance across all samples and signs of putative microdiversity. The majority of the amplicon sequence variants (ASVs) within each HoS clade clustered into a unique 97% sequence similarity operational taxonomic unit (OTU) and inhabited a specific environment (lotic, lentic or terrestrial). Our findings suggest the existence of microdiversification leading to sub-taxa niche differentiation, with putative distinct ecotypes (consisting of groups of ASVs) adapted to a specific environment. We hypothesize that HoS clades thriving in the CPWC have phylogenetically conserved traits that accelerate their rate of evolution, enabling them to adapt to strong spatio-temporally variable selection pressures. Variable selection appears to operate within clades to cause very rapid microdiversification without losing key traits that lead to high abundance. Variable and homogeneous selection, therefore, operate simultaneously but on different aspects of organismal ecology. The result is an overall signal of homogeneous selection due to rapid within-clade microdiversification caused by variable selection. It is unknown whether other systems experience this dynamic, and we encourage future work evaluating the transferability of our results.

59 BASIC BIOLOGICAL SCIENCES↗

Schneider Springs Fire Study 2023 for Ecosystem Respiration Rates: Surface Water Chemistry and Hydrologic Sensor Data across the Yakima River Basin, Washington, USA (v2)

This dataset supports a broader study examining the drivers of spatial variability in wildfire impacts across the Yakima River Basin. Data provided within this dataset were generated from sample collection across 17 total sites (8 sites affected by a recent wildfire, 9 sites unaffected by a recent wildfire) within multiple rivers throughout the Yakima River Basin in Washington, USA from May-July 2023. Fire affected sites are defined as those affected by the 2021 Schneider Springs Fire, based on the drainage area of the streams being within the 2021 Schneider Springs Fire burn perimeter or not (Figure 1, below). The contents include surface water geochemistry data (dissolved organic carbon; total dissolved nitrogen; total suspended solids); short-term sonde data (specific conductivity; turbidity; pH; chlorophyll A; temperature); stream depth data; stream velocity; manual chamber open channel respiration data; sensor time-series data (oxygen; water pressure; barometric pressure); field metadata (including qualitative information on in stream and river corridor characteristics); and environmental context photos taken in the field. The dataset also includes a summary file of the sensor data and plots of the sensor data. Sensors were only recovered at 15 out of the 17 sites, and not all sensors were recovered at all 15 sites (see Methods section for more details), therefore all data does not exist at all sites. Data from a 2022 study at the same sites, as well as additional sites, can be found at https://data.ess-dive.lbl.gov/view/doi:10.15485/1969566. The data package was originally published in November 2023. It was updated in June 2025 (v2; modified files). See the change history section in the readme for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of one folder with field photos and one main data folder with two subfolders. The main data folder consists of (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) field protocol; (5) readme; (6) international generic sample number (IGSN) mapping file; and (7) stream depth and averages. The sensor data subfolder consists of (1) sensor installation methods summary; (2) stream velocity; and (3) six subfolders. The BarotrollAtm (barometric pressure; temperature), DepthHOBO (water pressure; temperature), MantaRiver (specific conductivity; turbidity; pH; chlorophyll A; temperature), EXO (specific conductivity; pH; temperature), miniDOT (dissolved oxygen; temperature), and miniDOTManualChamber (dissolved oxygen; temperature) contain time-series data, plots, and summary files. The sample data subfolder consists of (1) total suspended solids (TSS) data; (2) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (3) total dissolved nitrogen (TN) data and averages; and (4) methods codes. All files are .csv, .pdf, .jpg, .jpeg, or .mov.

54 ENVIRONMENTAL SCIENCES↗

Geology and elevation shape bacterial assembly in Antarctic endolithic communities

Ice free areas of continental Antarctica are among the coldest and driest environments on Earth, and yet, they support surprisingly diverse and highly adapted microbial communities. Endolithic growth is one of the key adaptations to such extreme environments and often represents the dominant life-form. Despite growing scientific interest, little is known of the mechanisms that influence the assembly of endolithic microbiomes across these harsh environments. Here, we used metagenomics to examine the diversity and assembly of endolithic bacterial communities across Antarctica within different rock types and over a large elevation range. While granite supported richer and more heterogeneous communities than sandstone, elevation had no apparent effect on taxonomic richness, regardless of rock type. Conversely, elevation was clearly associated with turnover in community composition, with the deterministic process of variable selection driving microbial assembly along the elevation gradient. The turnover associated with elevation was modulated by geology, whereby for a given elevation difference, turnover was consistently larger between communities inhabiting different rock types. Overall, selection imposed by elevation and geology appeared stronger than turnover related to other spatially-structured environmental drivers. Our findings indicate that at the cold-arid limit of life on Earth, geology and elevation are key determinants of endolithic bacterial heterogeneity. This also suggests that warming temperatures may threaten the persistence of such extreme-adapted organisms.

54 ENVIRONMENTAL SCIENCES↗

Coupled primary production and respiration in a large river contrasts with smaller rivers and streams

Abstract Although time series in ecosystem metabolism are well characterized in small and medium rivers, patterns in the world's largest rivers are almost unknown. Large rivers present technical difficulties, including depth measurements, gas exchange (, ) estimates, and the presence of large dams, which can supersaturate gases. We estimated reach‐scale metabolism for the Hanford Reach of the Columbia River (Washington state, USA), a free‐flowing stretch with an average discharge of 3173 . We calculated from semi‐empirical models and directly estimated it from tracer measurements. We fixed at the median value from these calculations (0.5 ), and used maximum likelihood to estimate reach‐scale, open‐channel metabolism. Both gross primary production (GPP) and ecosystem respiration (ER) were high (GPP range: 0.3–30.8 g , ER range: 0.8–30.6 g ), with peak GPP and ER occurring in the late summer or early fall. GPP increased exponentially with temperature, consistent with metabolic theory, while light was seasonally saturating. Annual average GPP, estimated at 1500 g carbon , was in the top 2% of estimates for other rivers. GPP and ER were tightly coupled and 90% of GPP was immediately respired, resulting in net ecosystem production near 0. Patterns in the Hanford Reach contrast with those in small‐medium rivers, suggesting that metabolism magnitudes and patterns in large rivers may not be simply scaled from knowledge of smaller rivers.

54 ENVIRONMENTAL SCIENCES↗