Search NASA⌕ Search

Engineering topics

Son, Kyongho

Publications and source records attributed to Son, Kyongho.

Water Column Respiration in the Yakima River Basin is Explained by Temperature, Nutrients and Suspended Solids

Understanding aquatic ecosystem metabolism involves the study of two key processes: carbon fixation via primary production and organic C mineralization as total ecosystem respiration (ERtot). In streams and rivers, ERtot includes respiration in the water column (ERwc) and in the sediments (ERsed). While literature surveys suggest that ERsed is often a dominant contributor to ERtot, recent studies indicate that the relative influence of sediment-associated processes versus water column processes can fluctuate along the river continuum. Still, a comprehensive understanding of the factors contributing to these shifts within basins and across stream orders is needed. Here we contribute to this need by measuring ERwc and collecting water samples across 47 sites in the Yakima River basin, Washington, USA. We found that ERwc rates varied throughout the basin during baseflow conditions, ranging from –7.38 to 0.36 g O2 m?3 d?1, and encompassed the range of ERwc literature values. Additionally, by comparing to ERtot estimates for rivers across the contiguous United States, we suggest that the contribution of ERwc rates to reach-scale ERtot rates across the Yakima River was likely highly variable, but we did not test this directly. We observed that temperature, nutrient concentrations (dissolved organic carbon, total dissolved nitrogen), and total suspended solids explained 41% of ERwc variability across the basin. Our findings highlight the potential relevance of water column processes in aquatic ecosystem metabolism, with the Yakima River basin serving as an environmentally diverse river network representative of the larger Columbia River basin that spans much of the Pacific Northwest region of the United States. Our results are generally congruent with previous work, suggesting that the observed variability and suite of associated environmental factors influencing ERwc are potentially transferable across basins.

Laan, Maggi M.↗

Modeling the Effects of Artificial Drainage on Agriculture-Dominated Watersheds Using a Fully Distributed Integrated Hydrology Model

In agriculture-dominated watersheds where natural drainage is poor, agricultural ditches (narrow engineered channels) and tile drains (perforated pipes) are widely employed to enhance surface and subsurface drainage, respectively. Despite their relatively small scale, these features exert substantial control over the hydro-biogeochemical function of watersheds and their effects need to be represented in the models. We introduce a novel strategy to incorporate the effects of artificial agricultural drainage into a fully distributed basin-scale integrated surface-subsurface hydrology models. In our approach, narrow agriculture ditches for surface drainage are resolved efficiently using ditch-aligned computational meshes that are hydrologically conditioned to ensure connectivity in the stream/ditch network. For tile drainage in the subsurface, we use the physically based Hooghoudt's drainage equation as a subgrid model and route the water drained through tiles to the nearest ditch. Without site-specific calibration, this model reproduced observed streamflow in the Portage River Watershed (>1,000 km 2 ) as recorded by a USGS gauge with good accuracy (normalized KGE = 0.81) and outperformed a calibrated SWAT model (normalized KGE = 0.68). Numerical experiments confirm that artificial drainage reduces surface inundations and effectively controls the water table. At the watershed scale, artificial drainage increases baseflow but has little effect on watershed discharges above the 90th percentile. The strong physical underpinnings and reduced need for calibration allow us to study the impacts of artificial drainage on distributed hydrological response in terms of fluxes and states and provide a platform for investigating watershed-scale nutrient transport.

54 ENVIRONMENTAL SCIENCES↗

Evaluating SWAT + model uncertainties for human and natural outcomes: Application in a Great Lakes agricultural watershed

Nutrient exports from agricultural lands in the Great Lakes Region pose significant threats to water quality and ecological health through eutrophication, hypoxia, and harmful algal blooms. Climate change and agricultural adaptation practices complicate future nutrient loading due to intensified hydrologic cycles and land use decisions. Our research focuses on evaluating the Soil and Water Assessment Tool (SWAT) plus model parametric uncertainties for human and natural outcomes across different scales. These factors are integral to ensuring a balance between productive agricultural practices and maintaining the health of watershed hydrology. However, uncertainties in modeling such complex interactions pose significant challenges, limiting our ability to precisely determine critical factors that influence crop yield and soil moisture. Our analysis employs Sobol global sensitivity analysis to evaluate first-order, second order, and total-order indices for SWAT crop growth parameter, ensuring comprehensive assessment of individual and interactive effects on model outputs. The objective is to identify the parameters that significantly affect model outputs for crop yield and soil moisture and improve our understanding of their interactions at the basin and hydrological response unit (HRU) scale. Our case study, the Portage River Watershed, which drains into Lake Erie, is chosen to better capture finer scale interactions crucial for predicting nutrient loading under future climate scenarios. This foundational work is aimed at setting the stage for the future development of an agent-based model (ABM). The ABM model would incorporate SWAT outputs to dynamically simulate decision-making processes.

Bunyon, Enock↗

Modeling the Effects of Artificial Drainage on Agriculture-dominated Watersheds using a Fully Distributed Integrated Hydrology Model: Datasets, scripts, model files

This model-data archive supports the research paper that demonstrates the integration of agricultural drainage features—specifically, narrow engineered ditches and tile drains—into a fully distributed, basin-scale integrated surface-subsurface hydrology model (ISSHM), Amanzi-ATS. The model employs innovative computational meshes aligned with agricultural ditches and incorporates the physically based Hooghoudt's drainage equation to simulate tile drainage, offering a novel strategy that enhances the accuracy of hydrological simulations.The archived dataset includes input parameters, model configurations, and select simulation outputs for the Amanzi-ATS model that successfully captured the streamflow patterns in the Portage River Watershed as validated by USGS gauge readings. Jupyter notebook for the preparation of model inputs and post-processing of outputs are also included. The model's predictive performance achieved a normalized Kling-Gupta Efficiency (KGE) of 0.81, surpassing SWAT without the necessity for site-specific calibration.The Amanzi-ATS model presented in this modeL-data archive allows for numerical experiments to explore the shifts in the flow structure under different drainage scenarios. As a tool for advancing the understanding of distributed hydrological responses and nutrient cycling, this archived model provides valuable insights for researchers, modelers, and decision-makers involved in watershed management and environmental modeling.The Watershed Workflow package is implemented in Python3. The Jupyter notebooks can be executed through multiple open-source tools, for example, Anaconda Jupyter Lab, VS Studio Code, etc. Other data files include CSV and HDF5 files, which can be read through Python scripts. The input files for the ATS model, open-source integrated hydrology, and transport model, are in XML format and can be edited in any commonly used text editors.

54 ENVIRONMENTAL SCIENCES↗

Modeling the Effects of Artificial Drainage on Agriculture-dominated Watersheds using a Fully Distributed Integrated Hydrology Model: Datasets, code, models files

This model-data archive is for a modeling study focused on the effects of artificial drainage on agriculture-dominated watersheds using a fully distributed integrated hydrology model. This study introduced a novel strategy to represent the effects of tile drains and agricultural ditches in a fully distributed basin-scale integrated hydrology model. The numerical experiments in this study reveal the effects of surface and subsurface drainage on various hydrological states and fluxes. The details about the datasets can be found in the "readme" document attached with the dataset.

Agricultural Watersheds↗

Data and Scripts Associated with the Manuscript “Water Column Respiration in the Yakima River Basin is Explained by Temperature, Nutrients and Suspended Solids”

This data package is associated with the publication “Water Column Respiration in the Yakima River Basin is Explained by Temperature, Nutrients and Suspended Solids” published in EGU Biogeochemistry (Laan et al. 2025). In this research, water column respiration (ERwc) data, surface water chemistry data, organic matter (OM) chemistry data, and publicly available geospatial data were used in analysis to evaluate the variability in ERwc at 47 sites across the Yakima River basin in Washington, USA. In addition to this readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. The data package includes the data inputs, and outputs, and R scripts to reproduce all the analyses performed in the manuscript and create manuscript figures. The data package is comprised of three main folders (Code, Data, and Figures). The Code folder is comprised of four scripts and three analysis-specific subfolders that contain the R scripts to perform the analyses described in the publication and create publication figures. The Data folder is comprised of two “.csv” files and four subfolders that contain data input and output files. The Published_Data folder contains a readme that directs the user to download the appropriate files and add to this folder when using scripts. The Figures folder includes figures from the manuscript in “.pdf” and “.png” formats and a folder with intermediate figure files. This data package is associated with a GitHub repository which can be found at https://github.com/river-corridors-sfa/rcsfa-RC2-SPS-ERwc. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected some of these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Uncertainty characterization in a coupled human-natural system: Modeling agricultural adaptation in the Great Lakes Region

The Great Lakes Region's water quality and ecological health are threatened by the export of nutrients from agricultural lands, which causes eutrophication, hypoxia, and destructive algal blooms. The intensification of hydrologic cycles brought about by climate change is expected to exacerbate nutrient loading in the region, and, at the same time, agricultural adaptation to changing conditions is also expected to affect loading through shifting amounts and timing of fertilization. Quantifying these future effects and their interactions necessitates modeling both the human and natural processes as a coupled system, by pairing land use and agricultural management with hydrologic modeling. At the same time, compounding uncertainties arising from the complex interactions in both systems significantly limit our predictive understanding of the region's impacts. This study utilizes the Soil and Water Assessment Tool (SWAT), developed for simulating the impact of various farmer decisions on watershed functions in Western Lake Erie watersheds, and an under-development agent-based model (ABM) for agricultural management decisions. The aim of this study is to use global sensitivity analysis on the coupled ABM and SWAT models to quantify how uncertainty in both models interactively affects nutrient loading. To do so, we will conduct Sobol sensitivity analysis experiments at different levels of coupling assumptions to quantify how various uncertain factors (e.g., soil moisture and crop choice) and their interactions affect our estimates of nutrient loading. The results of this analysis will allow us to quantify how complex interactions and dependencies between both systems amplify the effect of uncertainties. Insights gained from this study will have broader implications for modeling the adaptive co-evolution of human and natural systems under climate change and can inform effective management of nutrient loading in the Great Lakes Region.

Climate Change↗

Riverine organic matter functional diversity increases with catchment size

A large amount of dissolved organic matter (DOM) is transported to the ocean from terrestrial inputs each year (~0.95 Pg C per year) and undergoes a series of abiotic and biotic reactions, causing a significant release of CO 2 . Combined, these reactions result in variable DOM characteristics (e.g., nominal oxidation state of carbon, double-bond equivalents, chemodiversity) which have demonstrated impacts on biogeochemistry and ecosystem function. Despite this importance, however, comparatively few studies focus on the drivers for DOM chemodiversity along a riverine continuum. Here, we characterized DOM within samples collected from a stream network in the Yakima River Basin using ultrahigh-resolution mass spectrometry (i.e., FTICR-MS). To link DOM chemistry to potential function, we identified putative biochemical transformations within each sample. We also used various molecular characteristics (e.g., thermodynamic favorability, degradability) to calculate a series of functional diversity metrics. We observed that the diversity of biochemical transformations increased with increasing upstream catchment area and landcover. This increase was also connected to expanding functional diversity of the molecular formula. This pattern suggests that as molecular formulas become more diverse in thermodynamics or degradability, there is increased opportunity for biochemical transformations, potentially creating a self-reinforcing cycle where transformations in turn increase diversity and diversity increase transformations. We also observed that these patterns are, in part, connected to landcover whereby the occurrence of many landcover types (e.g., agriculture, urban, forest, shrub) could expand DOM functional diversity. For example, we observed that a novel functional diversity metric measuring similarity to common freshwater molecular formulas (i.e., carboxyl-rich alicyclic molecules) was significantly related to urban coverage. These results show that DOM diversity does not decrease along stream networks, as predicted by a common conceptual model known as the River Continuum Concept, but rather are influenced by the thermodynamic and degradation potential of molecular formula within the DOM, as well as landcover patterns.

54 ENVIRONMENTAL SCIENCES↗

Spatial Study 2022: Water Column, Sediment, and Total Ecosystem Respiration Rates across the Yakima River Basin, Washington, USA (v2)

This dataset supports a broader study examining the drivers of spatial variability in sediment respiration rates in the Yakima River Basin and is associated with the manuscript “Sediment-associated processes account for most of the spatial variation in ecosystem respiration in the Yakima River basin” submitted to Nature Communications Earth & Environment (Garayburu-Caruso et al., in review). The dataset provides ecosystem metabolism estimates generated from streamMetabolizer (Appling et al.; 2018) using data collected during the same five-week period at 48 sites within multiple rivers throughout the Yakima River Basin in Washington, USA. Additionally, it includes the scripts used for the analysis and producing the figures in the manuscript. The contents include streamMetabolizer inputs and outputs and additional relevant data needed to generate the main manuscript results. The data included are: total ecosystem respiration, water respiration, calculated sediment-associated respiration, gross primary production outputs from the river corridor model for the Yakima River Basin, median grain size (d50), depth, dissolved oxygen, water temperature, pressure, and annual oxygen consumption. The associated GitHub repository can be found at https://github.com/river-corridors-sfa/SSS_metabolism. Samples collected during this study were labeled as “Second Spatial Study” or “SSS.” Raw time series sensor data, total suspended solids, and depth data from SSS were published at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1969566. A subset of data from the SSS samples were published in the contiguous United States (CONUS)-Scale Model-Sample (CM) study data package available at https://data.ess-dive.lbl.gov/view/doi:10.15485/1923689 that presents data from across the CONUS. They include dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC), total nitrogen (TN), grain size, aerobic sediment respiration, dissolved oxygen (DO), and temperature. Parent IDs and Site IDs are consistent between the SSS and CM data packages, and they can be mapped directly so data across packages can be used together. Field metadata for the samples in this da This dataset is comprised of one main data folder with four subfolders. The main data folder contains of (1) file-level metadata; (2) data dictionary; (3) total/water column/sediment respiration; (4) gross primary production (GPP); (5) median grain size (d50); and (6) annual oxygen consumption. The “Figures” subfolder contains the figures used in the paper and all intermediate files (including geospatial files). The “Published_Data” contains a readme directing the user to download the public data to reproduce analyses and figures. The “Scripts” folder contains all scripts used in the analyses that were not part of running StreamMetabolizer. Lastly, the “Stream_Metabolizer” folder contains all files associated with running StreamMetabolizer including (1) model input files, (2) model output files, (3) processing scripts, (4) histogram plots of the outputs, and (5) an R project. All files are .csv, .pdf, .R, .Rmd, .Rproj, .html, .png, .txt, .qgz, .cpg, .dbf, .prj, .shp, .shp.ea.iso.xml, .shp.iso.xml, .shx, .sbn. ta package can be found at either link. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Data associated with “Different methods of estimating riverbed sediment grain size diverge at the basin scale ” (v2)

This data package is associated with the publication “Different methods of estimating riverbed sediment grain size diverge at the basin scale” published in Frontiers in Earth Science (Regier et al., 2025). The distribution of sediment grain size in streams and rivers is often quantified by the median grain size (d50), a key metric for understanding and predicting hydrologic and biogeochemical function of streams and rivers. Manual methods to measure d50 are time-consuming and ignore larger grains, while model-based methods to estimate d50 often over-generalize basin characteristics, and therefore cannot accurately represent site-scale heterogeneity. Here, we apply a machine learning-enabled photogrammetry methodology (You Only Look Once, or YOLO) for estimating d50 for grains > 2 mm based on images collected from streams and rivers throughout the Yakima River Basin (YRB). To understand how such methods may help bridge the gaps in resolution and accuracy between manual and catchment characteristics model-based d50 estimates, we compared YOLO d50 values to manual and model-based estimates across the YRB. We found distinct differences among methods for d50 averages and variability, and relationships between d50 estimates and basin characteristics. Source images can be found at https://data.ess-dive.lbl.gov/view/doi:10.15485/1892052. This data package was originally published in May 2023. It was updated August 2025 (v2; new and modified files). File and folder names were not revised to indicate changes. See the change history section in the readme for more details. In addition to the readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of one main data folder containing (1) file-level metadata; (2) data dictionary; (3) readme; (4) and subfolders containing data, figures, and scripts. The data folder contains datasets used for the analyses in the manuscript in image, text-delimited or geospatially-referenced formats. The figures folder contains the figures from the manuscript in different formats. The scripts folder contains all of the scripts used to complete the analyses in the manuscript. All files are .csv, .rds, .dbf, .prj, .shp, .shx, .jpg, .png, .R, .Rproj, or .pdf. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected some of these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Spatial Study 2022: Surface Water Samples, Cotton Strip Degradation, and Hydrologic Sensor Data across the Yakima River Basin, Washington, USA (v3)

This dataset supports a broader study examining the drivers of spatial variability in sediment respiration rates in the Yakima River Basin. The dataset provides data and photos generated from sample collection during the same one-week period at 48 sites within multiple rivers throughout the Yakima River Basin in Washington, USA. The contents include surface water geochemistry data; river substrate grain size photos; stream depth data; manual chamber open channel respiration data; and field metadata (including qualitative information on instream and river corridor characteristics). Grain size photos can be used to improve estimates of channel substrate D50 data. The dataset also includes tensile strength and photos from cotton strip field degradation experiments; five-week sensor time series temperature, dissolved oxygen, pressure, pH, specific conductance, chlorophyll A, and turbidity data; plots of the sensor data; and R scripts used to generate the plots. Samples collected during this study were labeled as “Second Spatial Study” or “SSS.” A subset of data from the SSS samples were published in the contiguous United States (CONUS)-Scale Model-Sample (CM) study data package available at https://data.ess-dive.lbl.gov/view/doi:10.15485/1923689 that presents data from across the CONUS. SSS data published in the CM data package were not included in this data package. They include dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC), total nitrogen (TN), grain size, aerobic sediment respiration, dissolved oxygen (DO), and temperature. Parent IDs and Site IDs are consistent between the SSS and CM data packages, and they can be mapped directly so data across packages can be used together. Additionally, sensor data from a similar 2021 spatial study can be found at https://data.ess-dive.lbl.gov/view/doi:10.15485/1892052 and 2021 sample data can be found at https://data.ess-dive.lbl.gov/view/doi:10.15485/1898914. The 2021 spatial study had some sites in common with this 2022 spatial study. This dataset is comprised of three photo folders and one main data folder with six subfolders. The photo folders contain photographs and videos of cotton strip retrieval and sediment quadrats. The main data folder consists of (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) total suspended solids (TSS) data and cotton strip tensile strength data and averages; (5) field protocol; (6) readme; (7) methods codes; (8) international generic sample number (IGSN) mapping file; (9) sensor installation methods summary; (10) stream depth and averages; and (11) Ultrameter data and averages. The Sonar subfolder consists of Sonar time-series depth data and a processing script. The BarotrollAtm, DepthHOBO, MantaRiver, miniDOT, and miniDOTManualChamber subfolders contain time-series data, plots, and summary files. All files are .csv, .pdf, .txt, .R, .Rmd, .jpg, .jpeg, .AVI, .mp4, or .mov. The data package was originally published in April 2023. It was updated in August 2023 (v2; modified files) and September 2024 (v3; modified files). See the change history section in the readme for details. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

WHONDRS River Corridor Sediment and Water Geochemistry and In Situ Sensor Data from Machine-Learning-Informed Sites across the Contiguous United States (v6)

This dataset supports a broader study examining hyporheic zone respiration rates to improve predictive models at a contiguous United States (CONUS) scale. The CONUS-Scale Model-Sample Study (CM) was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the CONUS. New machine learning models were created every month to guide sampling locations. Data from the resulting samples were used to test and rebuild the machine learning models for the next round of sampling guidance. Sampling began in April 2022 and ended in October 2023. In addition to the widely distributed CONUS sites, a more spatially focused sampling occurred in the Yakima River Basin, WA in summer 2022. Data from this more spatially intensive sampling occurred under the label “Second Spatial Study (SSS)” and were also included in the machine learning models. Other data types collected from SSS that were not part of CM were published in a separate data package (https://data.ess-dive.lbl.gov/view/doi:10.15485/1969566). This data package was originally published in February 2023. It was updated in June 2023 (v2; new and modified files); December 2023 (v3; new and modified files); June 2024 (v4; new and modified files); April 2024 (v5; new and modified files); and September 2025 (v6; modified files). See the change history section in the readme for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of two folders of field photos and videos, one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) international generic sample number (IGSN) mapping file; (6) field protocols; (7) a subfolder with sample data; and (8) a subfolder with sensor data. The sample data subfolder contains (1) surface water and sediment dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) surface water and sediment total nitrogen data and averages; (3) surface water major cations and anions and averages; (4) sediment grain size data; (5) sediment iron (II) data and averages; (6) wet sediment mass, dry sediment mass, water mass, and wet sediment volume in incubation and sediment ICR vials; (7) sediment incubation respiration rate data and averages; (8) normalized respiration rate data and averages; (9) methods codes; (10) sediment specific surface area; (11) sediment percent carbon and nitrogen; (12) sediment gravimetric moisture and averages; (15) sediment X-ray diffraction (XRD) data; (16) sediment adenosine triphosphate (ATP) and averages; (17) a subfolder with sediment incubation respiration data, scripts, and plots; (18) surface water and sediment FTICR methods; and (19) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains five subfolders, one containing the sediment .xml data files, one containing the water .xml files, one containing the sediment CoreMS output files, one containing the water CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS).The sensor data subfolder contains (1) a subfolder with miniDOT dissolved oxygen and temperature data and plots; (2) miniDOT dissolved oxygen and temperature summary data; and (3) miniDOT installation methods. All files are .csv, .pdf, .R, .xml, .d, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, .png, .mov, or .mp4. CORRECTION: Carbon and nitrogen content are reported as percentages. The current column headers "01395_C_percent_per_mg" and "01397_N_percent_per_mg" are incorrect. These should read "01395_C_percent" and "01397_N_percent" and will be corrected in the next version of this data package. We thank the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, Cowiche Canyon Conservatory, Washington State Parks and Recreation Commission (Scientific Research Permit #210901), and the Confederated Tribes and Bands of the Yakama Nation for access to field locations where the samples labeled “SSS” were collected. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview. WHONDRS consortium members were asked to provide any acknowledgments for the collection of samples labeled “CM” and the following is a list of acknowledgments that were submitted with their corresponding Site IDs: (MART) Research activities were conducted in part on the Wind River Experimental Forest within the Gifford Pinchot National Forest; (MP- 100379) Philadelphia is part of Lenapehoking, the ancestral homelands of the Lenape peoples; (MP-102398) Land surveyed is the ancestral homelands of the Nookhose'iinenno (Arapaho), Tsis tsis'tas (Cheyenne), and Nuuchu (Ute); (MP-100749 and MP- 100747) Georgia Coastal Ecosystem LTER, OCE-1832178; (SP-70 and SP-72) Eastern Shoshone, Shoshone-Bannock; (MP- 102944) Funded by Oregon Watershed Enhancement Board. On the traditional lands of the Confederated Tribes of the Siletz, Confederated Tribes of the Grand Rhonde, and the Clatsop-Nehalem Confederated Tribe; (MP- 100607) Holiday Creek is located on the traditional territory of the Monacan Indian Nation; (SP-45) Lafayette Blue Springs State Park; (MP-102420) NSF DEB-2016749; (MP-100019) New Hampshire Agriculture Experiment Station; (SP-35) Rayonier (land owner; https://www.rayonier.com/); (MP- 101276) US Department of Energy, Office of Science, Biological and Environmental Research, Subsurface Biogeochemical Research, Watershed Dynamics and Evolution SFA at ORNL; (MP- 103224) Watershed Dynamics and Evolution SFA at ORNL; (MP- 101584) Traditional lands of the Oceti Sakowin (Dakota, Lakota, Nakoda) and Anishinaabe Peoples.

54 ENVIRONMENTAL SCIENCES↗

Model Inputs, Outputs, and Scripts associated with: “Spatial microbial respiration variations in the hyporheic zones within the Columbia River Basin”

This data package is associated with the publication “Spatial microbial respiration variations in the hyporheic zones within the Columbia River Basin” published in the Journal of Geophysical Research: Biogeosciences (Son et al. 2022) available at doi: 10.1029/2021JG006654. This data package includes the key model inputs/outputs of the river corridor model for the Columbia River Basin (CRB) and the model source codes, which were used in the manuscript. The model is a carbon-nitrogen-coupled river corridor model (RCM), and the model is used to quantify hyporheic zone (HZ) aerobic and anaerobic respiration at the NHDPLUS stream reach scales. The RCM used in this study combines empirical substrate models derived from observations and three microbially driven reactions to compute respiration of the HZ for each National Hydrography Dataset (NHD) reach within the CRB. The reactions in HZs of each NHD reach include anaerobic respiration and two-step anaerobic respiration via denitrification. Our HZ respiration estimates are limited to the lotic (or flowing) stream/river systems, and do not account for the respiration process in water column. Note that the RCM only simulates the HZ’s contribution to the dissolved carbon dioxide (CO2) concentrations in the streams, and the CO2 emissions to the atmosphere are not modelled. The model computes at hourly timesteps because of the fast reaction rates. The key input data of the model are exchange flux, residence time, and stream solute (dissolved organic carbon (DOC), dissolved oxygen (DO), and nitrate concentrations). These inputs are constant over time and represent long-term averaged values.This modeling framework successfully quantified HZ respiration components over multiple scales. It revealed key mechanisms driving the spatial variation of HZ aerobic and anaerobic respiration in reaches with varying hydrologic and substrate conditions. Thus, this modeling study offers a testing hypothesis in different river system (e.g., climate and biomes) for the HZ respiration processes, and can be used as a sampling design tool for large-scale HZ experimental studies.This dataset contains five folders: (1) model_inputs, (2) model_outputs, (3) Rscripts, (4) figures, and (5) model_codes. It also contains a readme, file level metadata (FLMD), and data dictionary (dd). Please see the FLMD for a list of all the files contained in this data package and descriptions for each. The model_inputs folder contains the model inputs used to drive the model simulations. The model_outputs folder contains key model output files from the river corridor model. The Rscripts folder contains the Rscripts for pre- and post- processing model results. The figures folder contains the raw figures associated with the manuscript. The model_codes folder includes key model source codes/input files. All files are .jpg, .jpeg, .out, .e, .od, .dat, .sub, .F90, .0, .R, .sbx, .cpg, .sbn, .shx, .shp, .dbf, .prj, .tfw, .tif, .xml, .pdf, or .csv.

54 ENVIRONMENTAL SCIENCES↗

Model Inputs, Outputs, and Scripts associated with: “Combined effects of stream hydrology and land use on basin-scale hyporheic zone denitrification in the Columbia River Basin”

This data package is associated with the publication “Combined effects of stream hydrology and land use on basin‐scale hyporheic zone denitrification in the Columbia River Basin”, published in Water Resource Research (Son et al.2022) available at https://doi.org/10.1029/2021WR031131. This data package includes the key model inputs/outputs of the river corridor model for the Columbia River Basin (CRB) and the model source codes used in the manuscript. The model is a carbon-nitrogen-coupled river corridor model (RCM), and the model is used to quantify hyporheic zone (HZ) denitrification at the NHDPLUS stream reach scales. The RCM used in this study combines empirical substrate models derived from observations and three microbially driven reactions, including two-step denitrification and aerobic respiration, are considered within the HZ. The key input data of the model are exchange flux, residence time, and stream solute (dissolved organic carbon (DOC), dissolved oxygen (DO), and nitrate concentrations). These inputs are constant over time and represent long-term averaged values. This study uses the RCM to explore the spatial patterns of HZ denitrification across reaches with different sizes and land use in the CRB. Our main objective is to use the RCM as a virtual reality model, and the machine-learning models as surrogates that encapsulate the complexities of the physics-based model while identifying the importance of different variables that are not evident in the model conceptualization. We do not include a direct comparison of the modeled HZ denitrification and measurements; however, the RCM can capture the overall spatial patterns of the HZ denitrification because the model inputs and its reaction networks are based on well-established theory and a physical-based model. The combination of the model-based predictions and a machine-learning approach (e.g., random forest) is used to improve our understanding of what variables of the model are associated with spatial patterns of the modeled denitrification across reaches with different sizes and land uses, and to develop a proxy model using measurable variables to reproduce the simulated patterns.This dataset contains five folders: (1) model_inputs, (2) model_outputs, (3) Rscripts, (4) figures, and (5) model_codes. It also contains a readme, file level metadata (FLMD), and data dictionary (dd). Please see the FLMD for a list of all the files contained in this data package and descriptions for each. The model_inputs folder contains the model inputs used to drive the model simulations. The model_outputs folder contains key model output files from the river corridor model. The Rscripts folder contains the Rscripts for pre- and post- processing model results. The figures folder contains the raw figures associated with the manuscript. The model_codes folder includes key model source codes/input files. All files are .jpg, .jpeg, .out, .e, .od, .dat, .sub, .F90, .0, .R, .sbx, .cpg, .sbn, .shx, .shp, .dbf, .prj, .tfw, .tif, .xml, .pdf, or .csv.

54 ENVIRONMENTAL SCIENCES↗

Combined Effects of Stream Hydrology and Land Use on Basin‐Scale Hyporheic Zone Denitrification in the Columbia River Basin

Abstract Denitrification in the hyporheic zone (HZ) of river corridors is crucial to removing excess nitrogen in rivers from anthropogenic activities. However, previous modeling studies of the effectiveness of river corridors in removing excess nitrogen via denitrification were often limited to the reach‐scale and low‐order stream watersheds. We developed a basin‐scale river corridor model for the Columbia River Basin with random forest models to identify the dominant factors associated with the spatial variation of HZ denitrification. Our modeling results suggest that the combined effects of hydrologic variability in reaches and substrate availability influenced by land use are associated with the spatial variability of modeled HZ denitrification at the basin scale. Hyporheic exchange flux can explain most of spatial variation of denitrification amounts in reaches of different sizes, while among the reaches affected by different land uses, the combination of hyporheic exchange flux and stream dissolved organic carbon (DOC) concentration can explain the denitrification differences. Also, we can generalize that the most influential watershed and channel variables controlling denitrification variation are channel morphology parameters (median grain size (D50), stream slope), climate (annual precipitation and evapotranspiration), and stream DOC‐related parameters (percent of shrub area). The modeling framework in our study can serve as a valuable tool to identify the limiting factors in removing excess nitrogen pollution in large river basins where direct measurement is often infeasible.

54 ENVIRONMENTAL SCIENCES↗