Search NASA⌕ Search

Engineering topics

Lin, Xinming

Publications and source records attributed to Lin, Xinming.

Water Column Respiration in the Yakima River Basin is Explained by Temperature, Nutrients and Suspended Solids

Understanding aquatic ecosystem metabolism involves the study of two key processes: carbon fixation via primary production and organic C mineralization as total ecosystem respiration (ERtot). In streams and rivers, ERtot includes respiration in the water column (ERwc) and in the sediments (ERsed). While literature surveys suggest that ERsed is often a dominant contributor to ERtot, recent studies indicate that the relative influence of sediment-associated processes versus water column processes can fluctuate along the river continuum. Still, a comprehensive understanding of the factors contributing to these shifts within basins and across stream orders is needed. Here we contribute to this need by measuring ERwc and collecting water samples across 47 sites in the Yakima River basin, Washington, USA. We found that ERwc rates varied throughout the basin during baseflow conditions, ranging from –7.38 to 0.36 g O2 m?3 d?1, and encompassed the range of ERwc literature values. Additionally, by comparing to ERtot estimates for rivers across the contiguous United States, we suggest that the contribution of ERwc rates to reach-scale ERtot rates across the Yakima River was likely highly variable, but we did not test this directly. We observed that temperature, nutrient concentrations (dissolved organic carbon, total dissolved nitrogen), and total suspended solids explained 41% of ERwc variability across the basin. Our findings highlight the potential relevance of water column processes in aquatic ecosystem metabolism, with the Yakima River basin serving as an environmentally diverse river network representative of the larger Columbia River basin that spans much of the Pacific Northwest region of the United States. Our results are generally congruent with previous work, suggesting that the observed variability and suite of associated environmental factors influencing ERwc are potentially transferable across basins.

Laan, Maggi M.↗

Integrating Analytical Solutions and U-Net Model for Predicting Groundwater Contaminant Plumes in Pump-and-Treat Systems

Pump-and-treat (P&T) is a common technique for groundwater remediation involving the extraction and treatment of contaminated water above ground. Optimizing the design and operation of the P&T well network is essential for maximizing the system’s effectiveness and efficiency. However, this optimization often necessitates many model evaluations, leading to computationally demanding tasks. This study introduces a novel approach that integrates analytical solutions for groundwater dynamics with the U-Net (Ronneberger et al., 2015) deep learning framework to predict groundwater contaminant plume migration under dynamic pumping conditions. By incorporating the Thiem equation (Thiem, 1906) into the input preprocessing, the U-Net model transforms sparse well data into a continuous spatial field that captures the hydraulic impacts of pumping activities. This integration enables the model to leverage both deep learning capabilities and classical physics-based groundwater theories, enhancing prediction accuracy and computational efficiency. These advancements can facilitate rapid, large-scale evaluations of P&T optimization simulations, allowing for timely and effective decision-making in well placement and system management. We demonstrate the model's robust performance across both simplified transient 2D models and a more complex 3D heterogeneous site model at the 200 West P&T facility at the Hanford Site. The U-Net-based model offers substantial computational advantages, reducing simulation times significantly compared to full physics-based models and providing a powerful tool for rapid site evaluation and P&T system optimization, such as evaluating alternative P&T well network designs. Our findings highlight the potential of advanced machine learning models to significantly enhance the efficiency and sustainability of groundwater remediation efforts, offering a novel application of U-Net architecture in environmental science.

Pump-and-treat↗

A unified ensemble soil moisture dataset across the continental United States

Abstract A unified ensemble soil moisture (SM) package has been developed over the Continental United States (CONUS). The data package includes 19 products from land surface models, remote sensing, reanalysis, and machine learning models. All datasets are unified to a 0.25-degree and monthly spatiotemporal resolution, providing a comprehensive view of surface SM dynamics. The statistical analysis of the datasets leverages the Koppen-Geiger Climate Classification to explore surface SM’s spatiotemporal variabilities. The extracted SM characteristics highlight distinct patterns, with the western CONUS showing larger coefficient of variation values and the eastern CONUS exhibiting higher SM values. Remote sensing datasets tend to be drier, while reanalysis products present wetter conditions. In-situ SM observations serve as the basis for wavelet power spectrum analyses to explain discrepancies in temporal scales across datasets facilitating daily SM records. This study provides a comprehensive soil moisture data package and an analysis framework that can be used for Earth system model evaluations and uncertainty quantification, quantifying drought impacts and land–atmosphere interactions and making recommendations for drought response planning.

54 ENVIRONMENTAL SCIENCES↗

An ML-based terrestrial data fusion and augmentation framework to enable advanced understanding of the terrestrial carbon and water interactions

Soil moisture is essential to the terrestrial carbon and water cycles and land–atmosphere interactions. There are various types of soil moisture data, and each type has the distinct spatiotemporal strengths and limitations, depending on the diverse applications and retrieval methodologies of different data types (Li et al., in review; The PNNL-82151 FY23 Report). However, the limitations of different soil moisture data in terms of accuracy and spatiotemporal coverage hinder our ability to further understand the soil moisture dynamics across scales. To have a gap free soil moisture data product with a fine spatiotemporal coverage and vertical profiles, we train extreme gradient boosting (XGBoost) models by using (1) in-situ soil moisture measurements from the International Soil Moisture Network (ISMN), (2) soil moisture from the ECMWF reanalysis (ERA) at the 9 km and sub-daily spatiotemporal resolution, (3) the Daymet meteorological fields, and (4) data products that characterize surface conditions, including soil texture, organic content, topography, vegetation type, and rooting depth. We use the trained XGBoost models that have consistent performance across seven soil layers, i.e., 0–5 cm, 5–10 cm, 10–20 cm, 20–40 cm, 40–60 cm, 60–100 cm, and 100–200 cm, and the gridded model predictors to generate a soil moisture data at the 1 km and daily spatiotemporal resolution for the Continental United States (CONUS) from 2001–2020. This dataset can be broadly used for Earth system model benchmark, monitoring extreme weathers, making informed decisions regarding agriculture, water resource management, climate change mitigation, and ecosystem preservation.

58 GEOSCIENCES↗

On the transferability of residence time distributions in two 10-km long river sections with similar hydromorphic units

Quantifying hydrologic exchange fluxes (HEFs) at the stream-groundwater interface and their residence time distributions (RTDs) in the subsurface are important for managing the water quality and ecosystem health in dynamic river corridors. However, direct simulating high-spatial resolution HEFs and RTDs can be time-consuming, especially for watershed-scale modeling. Efficient surrogate models linking RTDs to hydromorphic units (HUs) can be alternatives for simulating RTDs in large-scale models. A common concern of these surrogate models, though, is the transferability of the relationship between the RTDs and HUs from one river corridor to another. To address this issue, this work evaluates the HEFs and resulting RTD-HU relationships for two 10-km long river corridors along the Columbia River leveraging a one-way coupled three-dimensional transient surface-subsurface water transport modeling framework we previously developed. Applying such a framework at the two river corridors with similar HUs allows for quantitative comparisons of HEFs and RTDs using both statistical tests and machine learning classification models. Finally, our comparison shows that the similarity and transferability of the RTD-HU relationship is very low for the two investigated river sections, which suggests that devising a general algorithm to estimate RTDs based solely on surface water hydrodynamics and short-distance river channel topography data, as well as HU classification, might be nearly impossible.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning Analysis of Impact of Western US Fires on Central US Hailstorms

Fires, including wildfires, harm air quality and essential public services like transportation, communication, and utilities. These fires can also influence atmospheric conditions, including temperature and aerosols, potentially affecting severe convective storms. Here, we investigate the remote impacts of fires in the western United States (WUS) on the occurrence of large hail (size: $\geqslant$ 2.54 cm) in the central US (CUS) over the 20-year period of 2001–20 using the machine learning (ML), Random Forest (RF), and Extreme Gradient Boosting (XGB) methods. The developed RF and XGB models demonstrate high accuracy (> 90%) and F1 scores of up to 0.78 in predicting large hail occurrences when WUS fires and CUS hailstorms coincide, particularly in four states (Wyoming, South Dakota, Nebraska, and Kansas). The key contributing variables identified from both ML models include the meteorological variables in the fire region (temperature and moisture), the westerly wind over the plume transport path, and the fire features (i.e., the maximum fire power and burned area). Importantly, the results confirm a linkage between WUS fires and severe weather in the CUS, corroborating the findings of our previous modeling study conducted on case simulations with a detailed physics model.

54 ENVIRONMENTAL SCIENCES↗

Data and Scripts Associated with the Manuscript “Water Column Respiration in the Yakima River Basin is Explained by Temperature, Nutrients and Suspended Solids”

This data package is associated with the publication “Water Column Respiration in the Yakima River Basin is Explained by Temperature, Nutrients and Suspended Solids” published in EGU Biogeochemistry (Laan et al. 2025). In this research, water column respiration (ERwc) data, surface water chemistry data, organic matter (OM) chemistry data, and publicly available geospatial data were used in analysis to evaluate the variability in ERwc at 47 sites across the Yakima River basin in Washington, USA. In addition to this readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. The data package includes the data inputs, and outputs, and R scripts to reproduce all the analyses performed in the manuscript and create manuscript figures. The data package is comprised of three main folders (Code, Data, and Figures). The Code folder is comprised of four scripts and three analysis-specific subfolders that contain the R scripts to perform the analyses described in the publication and create publication figures. The Data folder is comprised of two “.csv” files and four subfolders that contain data input and output files. The Published_Data folder contains a readme that directs the user to download the appropriate files and add to this folder when using scripts. The Figures folder includes figures from the manuscript in “.pdf” and “.png” formats and a folder with intermediate figure files. This data package is associated with a GitHub repository which can be found at https://github.com/river-corridors-sfa/rcsfa-RC2-SPS-ERwc. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected some of these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Quantifying Drivers of Methane Hydrobiogeochemistry in a Tidal River Floodplain System

The influence of coastal ecosystems on global greenhouse gas (GHG) budgets and their response to increasing inundation and salinization remains poorly constrained. In this study, we have integrated an uncertainty quantification (UQ) and ensemble machine learning (ML) framework to identify and rank the most influential processes, properties, and conditions controlling methane behavior in a freshwater floodplain responding to recently restored seawater inundation. Our unique multivariate, multiyear, and multi-site dataset comprises tidal creek and floodplain porewater observations encompassing water level, salinity, pH, temperature, dissolved oxygen (DO), dissolved organic carbon (DOC), total dissolved nitrogen (TDN), partial pressure of carbon dioxide (pCO 2 ), nitrous oxide (pN 2 O), methane (pCH 4 ), and the stable isotopic composition of methane (δ 13 CH 4 ). Additionally, we incorporated topographical data, soil porosity, hydraulic conductivity, and water retention parameters for UQ analysis using a previously developed 3D variably saturated flow and transport floodplain model for a physical mechanistic understanding of factors influencing groundwater levels and salinity and, therefore, CH 4 . Principal component analysis revealed that groundwater level and salinity are the most significant predictors of overall biogeochemical variability. The ensemble ML models and UQ analyses identified DO, water level, salinity, and temperature as the most influential factors for porewater methane levels and indicated that approximately 80% of the total variability in hourly water levels and around 60% of the total variability in hourly salinity can be explained by permeability, creek water level, and two van Genuchten water retention function parameters: the air-entry suction parameter α and the pore size distribution parameter m. These findings provide insights on the physicochemical factors in methane behavior in coastal ecosystems and their representation in local- to global-scale Earth system models.

54 ENVIRONMENTAL SCIENCES↗

Schneider Springs Fire Study 2023 for Ecosystem Respiration Rates: Surface Water Chemistry and Hydrologic Sensor Data across the Yakima River Basin, Washington, USA (v2)

This dataset supports a broader study examining the drivers of spatial variability in wildfire impacts across the Yakima River Basin. Data provided within this dataset were generated from sample collection across 17 total sites (8 sites affected by a recent wildfire, 9 sites unaffected by a recent wildfire) within multiple rivers throughout the Yakima River Basin in Washington, USA from May-July 2023. Fire affected sites are defined as those affected by the 2021 Schneider Springs Fire, based on the drainage area of the streams being within the 2021 Schneider Springs Fire burn perimeter or not (Figure 1, below). The contents include surface water geochemistry data (dissolved organic carbon; total dissolved nitrogen; total suspended solids); short-term sonde data (specific conductivity; turbidity; pH; chlorophyll A; temperature); stream depth data; stream velocity; manual chamber open channel respiration data; sensor time-series data (oxygen; water pressure; barometric pressure); field metadata (including qualitative information on in stream and river corridor characteristics); and environmental context photos taken in the field. The dataset also includes a summary file of the sensor data and plots of the sensor data. Sensors were only recovered at 15 out of the 17 sites, and not all sensors were recovered at all 15 sites (see Methods section for more details), therefore all data does not exist at all sites. Data from a 2022 study at the same sites, as well as additional sites, can be found at https://data.ess-dive.lbl.gov/view/doi:10.15485/1969566. The data package was originally published in November 2023. It was updated in June 2025 (v2; modified files). See the change history section in the readme for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of one folder with field photos and one main data folder with two subfolders. The main data folder consists of (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) field protocol; (5) readme; (6) international generic sample number (IGSN) mapping file; and (7) stream depth and averages. The sensor data subfolder consists of (1) sensor installation methods summary; (2) stream velocity; and (3) six subfolders. The BarotrollAtm (barometric pressure; temperature), DepthHOBO (water pressure; temperature), MantaRiver (specific conductivity; turbidity; pH; chlorophyll A; temperature), EXO (specific conductivity; pH; temperature), miniDOT (dissolved oxygen; temperature), and miniDOTManualChamber (dissolved oxygen; temperature) contain time-series data, plots, and summary files. The sample data subfolder consists of (1) total suspended solids (TSS) data; (2) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (3) total dissolved nitrogen (TN) data and averages; and (4) methods codes. All files are .csv, .pdf, .jpg, .jpeg, or .mov.

54 ENVIRONMENTAL SCIENCES↗

An Online Prototype Toolset for Predicting and Optimizing P&T Performance (FY23 Status Report)

A new web-based toolset is being developed to support ongoing remediation optimization efforts and implementation of an adaptive site management strategy for the 200 West Area Pump-and-Treat (P&T) system at the Hanford Site. This toolset, comprising the well performance index tool and the well optimization pre-screening tool, will offer a user-friendly interface to predict and optimize the P&T well network’s performance at a preliminary level. Efforts in fiscal year (FY) 2023 focused on three main components: updating the existing deep learning model for predicting P&T performance, designing and developing a prototype of a web-based performance index tool, and initiating the conceptual design of the well optimization pre-screening tool. The well performance index tool is based on a pre-trained deep learning model that allows users to select a target contaminant and well screen length, then visualize the predicted performance of potential new wells across the site. The well optimization pre-screening tool includes two separate modules: the pre-computed scenario viewer, which organizes and visualizes offline optimization simulation results, and the quick analysis module, which provides real-time model prediction using user-specified well locations. In FY24, the plan is to add web-based applications to SOCRATES for both the well performance prediction tool and the optimization prescreening tool, with accompanying user and theory guides. These tools are intended to enable an accessible, easily applied, and transparent approach to remedy planning and decision-making.

97 MATHEMATICS AND COMPUTING↗

An Inventory of AI-ready Benchmark Data for US Fires, Heatwaves, and Droughts

Extreme weather events, including fires, heatwaves, and droughts, have significant impacts on earth, environmental, and energy systems. Mechanistic and predictive understanding, as well as probabilistic risk assessment of these extreme weather events, are crucial for detecting, planning for, and responding to these extremes. Records of extreme weather events provide an important data source for understanding present and future extremes, but the existing data needs preprocessing before it can be used for analysis. Moreover, there are many nonstandard metrics defining the levels of severity or impacts of extremes. In this study, we compile a comprehensive benchmark data inventory of extreme weather events, including fires, heatwaves, and droughts. The dataset covers the period from 2001 to 2020 with a daily temporal resolution and a spatial resolution of 0.5°×0.5° (~55km×55km) over the continental United States (CONUS), and a spatial resolution of 1km × 1km over the Pacific Northwest (PNW) region, together with the co-located and relevant meteorological variables. By exploring and summarizing the spatial and temporal patterns of these extremes in various forms of marginal, conditional, and joint probability distributions, we gain a better understanding of the characteristics of climate extremes. The resulting AI/ML-ready data products can be readily applied to ML-based research, fostering and encouraging AI/ML research in the field of extreme weather. This study can contribute significantly to the advancement of extreme weather research, aiding researchers, policymakers, and practitioners in developing improved preparedness and response strategies to protect communities and ecosystems from the adverse impacts of extreme weather events. Usage Notes We presented a long term (2001-2020) and comprehensive data inventory of historical extreme events with daily temporal resolution covering the separate spatial extents of CONUS (0.5°×0.5°) and PNW(1km×1km) for various applications and studies. The dataset with 0.5°×0.5° resolution for CONUS can be used to help build more accurate climate models for the entire CONUS, which can help in understanding long-term climate trends, including changes in the frequency and intensity of extreme events, predicting future extreme events as well as understanding the implications of extreme events on society and the environment. The data can also be applied for risk accessment of the extremes. For example, ML/AI models can be developed to predict wildfire risk or forecast HWs by analyzing historical weather data, and past fires or heateave , allowing for early warnings and risk mitigation strategies. Using this dataset, AI-driven risk assessment models can also be built to identify vulnerable energy and utilities infrastructure, imrpove grid resilience and suggest adaptations to withstand extreme weather events. The high-resolution 1km×1km dataset ove PNW are advantageous for real-time, localized and detailed applications. It can enhance the accuracy of early warning systems for extreme weather events, helping authorities and communities prepare for and respond to disasters more effectively. For example, ML models can be developed to provide localized HW predictions for specific neighborhoods or cities, enabling residents and local emergency services to take targeted actions; the assessment of drought severity in specific communities or watersheds within the PNW can help local authorities manage water resources more effectively.

Lin, Xinming↗

Evaluation of Multi-Fidelity Soil Moisture Products Across the Continental United States

We have aggregated the most recent soil moisture datasets from a diverse range of sources, encompassing the Continental United States (CONUS). These sources encompass gridded data from remote sensing products, reanalysis products, machine learning-based projects, and land surface modeling products. Additionally, we have obtained and processed in-situ soil moisture observations from the International Soil Moisture Network. The collected datasets exhibit variations in both temporal and spatial resolutions. Among the 20 datasets, six are available at a spatial resolution of 0.25 degrees, while three are at a coarser spatial resolution of 25 km. To minimize spatial interpolation, we conducted data uncertainty evaluations at the 0.25-degree spatial resolution. For our data evaluations, we maintained a monthly temporal resolution, which effectively captures soil moisture seasonality and interannual variability. Our data processing strategy preserves the raw data and interpolated data at their original temporal resolutions. Datasets with higher temporal resolutions, including daily, three-hourly, and hourly datasets, are set aside for subsequent analyses. These analyses will delve into topics such as soil moisture changes and recovery during extreme weather events. Furthermore, we have processed auxiliary data to enhance our evaluation, leveraging tools such as Google Earth Engine. This includes incorporating topography data, land use land cover data, Köppen-Geiger climate classification, and more to provide a comprehensive assessment from multiple sources.

Li, Lingcheng↗

An Inventory of AI-ready Benchmark Data for US Fires, Heatwaves, and Droughts

Extreme weather events, including fires, heatwaves, and droughts, have significant impacts on earth, environmental, and energy systems. Mechanistic and predictive understanding, as well as probabilistic risk assessment of these extreme weather events, are crucial for detecting, planning for, and responding to these extremes. Records of extreme weather events provide an important data source for understanding present and future extremes, but the existing data needs preprocessing before it can be used for analysis. Moreover, there are many nonstandard metrics defining the levels of severity or impacts of extremes. In this study, we compile a comprehensive benchmark data inventory of extreme weather events, including fires, heatwaves, and droughts. The dataset covers the period from 2001 to 2020 with a daily temporal resolution and a spatial resolution of 0.5°×0.5° (~55km×55km) over the continental United States (CONUS), and a spatial resolution of 1km × 1km over the Pacific Northwest (PNW) region, together with the co-located and relevant meteorological variables. By exploring and summarizing the spatial and temporal patterns of these extremes in various forms of marginal, conditional, and joint probability distributions, we gain a better understanding of the characteristics of climate extremes. The resulting AI/ML-ready data products can be readily applied to ML-based research, fostering and encouraging AI/ML research in the field of extreme weather. This study can contribute significantly to the advancement of extreme weather research, aiding researchers, policymakers, and practitioners in developing improved preparedness and response strategies to protect communities and ecosystems from the adverse impacts of extreme weather events.

54 ENVIRONMENTAL SCIENCES↗

Spatial Study 2022: Water Column, Sediment, and Total Ecosystem Respiration Rates across the Yakima River Basin, Washington, USA (v2)

This dataset supports a broader study examining the drivers of spatial variability in sediment respiration rates in the Yakima River Basin and is associated with the manuscript “Sediment-associated processes account for most of the spatial variation in ecosystem respiration in the Yakima River basin” submitted to Nature Communications Earth & Environment (Garayburu-Caruso et al., in review). The dataset provides ecosystem metabolism estimates generated from streamMetabolizer (Appling et al.; 2018) using data collected during the same five-week period at 48 sites within multiple rivers throughout the Yakima River Basin in Washington, USA. Additionally, it includes the scripts used for the analysis and producing the figures in the manuscript. The contents include streamMetabolizer inputs and outputs and additional relevant data needed to generate the main manuscript results. The data included are: total ecosystem respiration, water respiration, calculated sediment-associated respiration, gross primary production outputs from the river corridor model for the Yakima River Basin, median grain size (d50), depth, dissolved oxygen, water temperature, pressure, and annual oxygen consumption. The associated GitHub repository can be found at https://github.com/river-corridors-sfa/SSS_metabolism. Samples collected during this study were labeled as “Second Spatial Study” or “SSS.” Raw time series sensor data, total suspended solids, and depth data from SSS were published at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1969566. A subset of data from the SSS samples were published in the contiguous United States (CONUS)-Scale Model-Sample (CM) study data package available at https://data.ess-dive.lbl.gov/view/doi:10.15485/1923689 that presents data from across the CONUS. They include dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC), total nitrogen (TN), grain size, aerobic sediment respiration, dissolved oxygen (DO), and temperature. Parent IDs and Site IDs are consistent between the SSS and CM data packages, and they can be mapped directly so data across packages can be used together. Field metadata for the samples in this da This dataset is comprised of one main data folder with four subfolders. The main data folder contains of (1) file-level metadata; (2) data dictionary; (3) total/water column/sediment respiration; (4) gross primary production (GPP); (5) median grain size (d50); and (6) annual oxygen consumption. The “Figures” subfolder contains the figures used in the paper and all intermediate files (including geospatial files). The “Published_Data” contains a readme directing the user to download the public data to reproduce analyses and figures. The “Scripts” folder contains all scripts used in the analyses that were not part of running StreamMetabolizer. Lastly, the “Stream_Metabolizer” folder contains all files associated with running StreamMetabolizer including (1) model input files, (2) model output files, (3) processing scripts, (4) histogram plots of the outputs, and (5) an R project. All files are .csv, .pdf, .R, .Rmd, .Rproj, .html, .png, .txt, .qgz, .cpg, .dbf, .prj, .shp, .shp.ea.iso.xml, .shp.iso.xml, .shx, .sbn. ta package can be found at either link. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Spatial Study 2022: Surface Water Samples, Cotton Strip Degradation, and Hydrologic Sensor Data across the Yakima River Basin, Washington, USA (v3)

This dataset supports a broader study examining the drivers of spatial variability in sediment respiration rates in the Yakima River Basin. The dataset provides data and photos generated from sample collection during the same one-week period at 48 sites within multiple rivers throughout the Yakima River Basin in Washington, USA. The contents include surface water geochemistry data; river substrate grain size photos; stream depth data; manual chamber open channel respiration data; and field metadata (including qualitative information on instream and river corridor characteristics). Grain size photos can be used to improve estimates of channel substrate D50 data. The dataset also includes tensile strength and photos from cotton strip field degradation experiments; five-week sensor time series temperature, dissolved oxygen, pressure, pH, specific conductance, chlorophyll A, and turbidity data; plots of the sensor data; and R scripts used to generate the plots. Samples collected during this study were labeled as “Second Spatial Study” or “SSS.” A subset of data from the SSS samples were published in the contiguous United States (CONUS)-Scale Model-Sample (CM) study data package available at https://data.ess-dive.lbl.gov/view/doi:10.15485/1923689 that presents data from across the CONUS. SSS data published in the CM data package were not included in this data package. They include dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC), total nitrogen (TN), grain size, aerobic sediment respiration, dissolved oxygen (DO), and temperature. Parent IDs and Site IDs are consistent between the SSS and CM data packages, and they can be mapped directly so data across packages can be used together. Additionally, sensor data from a similar 2021 spatial study can be found at https://data.ess-dive.lbl.gov/view/doi:10.15485/1892052 and 2021 sample data can be found at https://data.ess-dive.lbl.gov/view/doi:10.15485/1898914. The 2021 spatial study had some sites in common with this 2022 spatial study. This dataset is comprised of three photo folders and one main data folder with six subfolders. The photo folders contain photographs and videos of cotton strip retrieval and sediment quadrats. The main data folder consists of (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) total suspended solids (TSS) data and cotton strip tensile strength data and averages; (5) field protocol; (6) readme; (7) methods codes; (8) international generic sample number (IGSN) mapping file; (9) sensor installation methods summary; (10) stream depth and averages; and (11) Ultrameter data and averages. The Sonar subfolder consists of Sonar time-series depth data and a processing script. The BarotrollAtm, DepthHOBO, MantaRiver, miniDOT, and miniDOTManualChamber subfolders contain time-series data, plots, and summary files. All files are .csv, .pdf, .txt, .R, .Rmd, .jpg, .jpeg, .AVI, .mp4, or .mov. The data package was originally published in April 2023. It was updated in August 2023 (v2; modified files) and September 2024 (v3; modified files). See the change history section in the readme for details. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning of Key Variables Impacting Extreme Precipitation in Various Regions of the Contiguous United States

Abstract Amplification in extreme precipitation intensity and frequency can cause severe flooding and impose significant social and economic consequences. Variations in extreme precipitation intensity, frequencies, and return periods can be attributed to many physical variables across spatial and temporal scales. Here we employ ensemble machine learning (ML) methods, namely random forest (RF), eXtreme Gradient Boosting (XGB), and artificial neural networks (ANN), to explore key contributing variables to monthly extreme precipitation intensity and frequency in six regions over the United States. We further establish emulators for return periods. Results show that the ML models for intensity perform better in regions with obvious seasonality (i.e., Northern Great Plains, Southern Great Plains, and West Coast) than the other three regions (Northeast, Southwest, and Rocky Mountains), while for frequency the models perform well for most regions. The Shapley additive explanation is used to help explain the relationships between extreme precipitation characteristics and identify top variables for RF and XGB. We find that latent heat flux, relative humidity, soil moisture, and large‐scale subsidence are key common variables across the regions for both monthly intensity and frequency, and their compound effects are non‐negligible. The developed ML models capture the probability and return period of extreme precipitation well for all regions and may be used for decision making (e.g., infrastructure planning and design).

54 ENVIRONMENTAL SCIENCES↗

Predicting future well performance for environmental remediation design using deep learning

Here in this study, we developed a deep learning (DL) framework with a multi-channel three-dimensional convolutional neural network (MC3D-CNN) to predict well performance and thereby assist future environmental remediation design. Such prediction of extraction well performance at designated locations is critical for configuring pump-and-treat (P&T) well network design and operation, setting reasonable target closure dates for overall remedying, and estimating remedy costs. The framework is developed with operational and monitoring data routinely collected during P&T remedy operations, including well extraction and injection rates as well as in situ contaminant concentrations. Traditionally, the collected data were rarely used for purposes other than assessing past well performance and the accuracy of the conceptual site model. However, recent advances in data-driven computational approaches enable better use of the large datasets to inform future well performance, enhance site characterization, and improve remediation planning. In this study, we established a DL framework to integrate transient three-dimensional contaminant plumes and multiple aquifer properties (e.g., hydraulic conductivity and hydrostratigraphic maps) to identify characteristic patterns controlling and representing extraction well mass recovery, aiming at providing future mass recovery estimates for existing wells and candidate wells at any proposed locations. We evaluated our framework by using a realistic synthetic dataset generated from a well-calibrated flow and transport model used in the 200 West Area of the U.S. Department of Energy’s Hanford Site in southeastern Washington state. The multi-channel feature in our framework allows integration of various types and temporal densities of training datasets for DL model development. Overall, we found that the trained DL model achieved an accuracy of over 90% in ranking extraction well performance in validation datasets, and over 80% in predicting high-performance-ranking well locations. This data-informed approach provides a flexible tool to support adaptive site management, streamline decision-making, and potentially reduce remediation time and costs. Our DL framework can be used as a filtering tool to improve the current P&T network optimization design by reducing the number of candidate well locations.

54 ENVIRONMENTAL SCIENCES↗