Search NASA⌕ Search

SEARCH · Search NASA

Results for “Daily Groundwater”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Data and scripts associated with “Sequential Precipitation Input Tagging (SPIT) to Estimate Water Transit Times and Hydrologic Tracer Dynamics within Water-Tagging Enabled Hydrologic Models” (v3)

This data package is associated with the publication “Sequential Precipitation Input Tagging (SPIT) to Estimate Water Transit Times and Hydrologic Tracer Dynamics within Water-Tagging Enabled Hydrologic Models” submitted to Journal of Advances in Modeling Earth Systems (Butler et al. 2025). This study developed the Sequential Precipitation Input Tagging (SPIT) framework to tag input precipitation and estimate water transit times and hydrologic tracers. SPIT tags all precipitation events at regular intervals over an extended period (monthly tags over seven years) in a hydrologic model from 2016-2022. SPIT is applied at six National Ecological Observatory Network (NEON) sites across the continental United States to calculate transit time distributions (TTD) and derive from these mean transit times (MTT), fractions of young water (Fyw), and hydrologic tracer concentrations in stream water (δ18O) within a water-tagging enabled version of the Weather Research and Forecast (WT-WRF-Hydro) model with national water model (NWM) configurations. We go on to validate WT-WRF-Hydro estimates against Butler et al. (2023), who analyzed the same NEON sites using stable water isotope data to estimate water transit times. This new tracking method provides a detailed picture of water movement and helps improve predictions about water availability in the future. This data package was originally published in January 2025. It was updated May 2025 (v2; new and modified files) and October 2025 (v3; new and modified files). File and folder names were not revised to indicate changes. See the change history section in the readme for more details. This data package contains the data and scripts used to develop the SPIT framework WT-WRF-Hydro (Water Tagging Weather Research and Forecasting Hydrologic) model and is associated with the following GitHub repository: https://github.com/zbutler33/SPIT-Framework. This data package contains five parent folders: (1) “Manipulated_outputs”, (2) “Metadata”, (3) “Observed”, (4) “Outputs”, and (5) “Scripts”. Each of these parent folders contains additional subfolders and files. Please see the FLMD (“v*_Butler_2024_WT_WRF_Hydro_flmd.csv”) for a list of all the files contained in this data package and descriptions for each. See the data dictionary (“v*_Butler_2024_WT_WRF_Hydro_dd.csv”) for definitions and units of all of the tabular (files ending in “.csv” and ".tsv") column headers.

54 ENVIRONMENTAL SCIENCES↗

California Trees Seasonally Use Augmented Water Sources: Water Isotope Tracking in a Groundwater‐Dependent Ecosystem

Sustainable groundwater management must account for the needs of groundwater dependent ecosystems. To understand the relationship of ecosystems and seasonal water use, we studied the stable isotope composition (δ 18 O and δ 2 H) of water in streamside trees in a semi-arid streamside environment (Livermore, California, USA). We sampled seven trees at two sites every other month from April 2024 through April 2025 for tree xylem stable water isotope signatures. These data were compared to potential source waters: precipitation, imported surface water, soil water and regional groundwaters. Large daily precipitation events were found to be isotopically similar to regional groundwater and were thus treated as one water source. A Bayesian mixing model using stable water isotopes was used to determine the ratios of these three potential source waters (small daily precipitation events, groundwater/large daily precipitation events and imported water) present in tree xylem water. On average, tree water sources include 32% imported water (SD = 9%), 35% small daily precipitation events (SD = 10%) and 33% groundwater (SD = 3%), with significant seasonal variation (t summer-winter = 30.8, p < 0.01), particularly drawing more imported water (more than 55%) in the summer. While small daily precipitation events contribute only 10% of the total precipitation in our dataset, it represents a third of water used by trees. In addition, while these ecosystems are designated as groundwater dependent ecosystems, the trees use approximately one third imported water and even more during dry summer months. This approach provides water managers with a practical tool for quantifying ecosystem water needs, supporting data-driven decisions and regulatory compliance. While California's Sustainable Groundwater Management Act emphasizes supporting groundwater dependent ecosystems, it allows flexibility in demonstrating benefits. In conclusion, our methodology provides a way to document that management actions (such as managed aquifer recharge with imported water) deliver measurable co-benefits to GDEs.

Geosciences↗

Groundwater and Surface Water Flow (GSFLOW) model files to explore bedrock circulation depth and porosity in Copper Creek, Colorado

This data package contains integrated hydrological model input and output files for Copper Creek, Colorado (24 km2), a tributary of the East River located in the headwaters of the Upper Colorado River Basin. The model code is the U.S. Geological Survey (USGS) Groundwater and Surface Water Flow (GSFLOW) model. The model contains a 100-m grid resolution and a daily timestep. The land surface model is dynamically linked to a three-dimensional groundwater flow model that allows for streamflow gaining and losing conditions. The groundwater model contains 12 model layers and extends 400 m below land surface. The original Copper Creek model was modified to contain geologic layers representing saprolite, shallow bedrock, and deep bedrock. Endmember depth versus hydraulic conductivity relationships and porosity values for fractured crystalline rock are simulated. For the shallow case, median flow depths occur in the shallow saprolite at depths <8 m, while the deep case promotes a median groundwater flow depth of 100 m. With this modeling framework we compare streamflow response to a plausible worst-case drought lasting up to five years. Streamflow metrics of analysis include average streamflow, fraction of stream network that is dry, no-flow duration, average groundwater flow to streams and time to recovery following the drought. Results and implications are presented in a paper submitted to Geophysical Research Letters titled, "The role of bedrock circulation depth and porosity in mountain streamflow response to prolonged drought" by Rosemary WH. Carroll, Andrew H. Manning and Kenneth H Williams. A Readme.txt file provides instructions on how to download all model files and execute each model scenario. In addition to the GSFLOW output/prms/copper_drought.csv file containing daily basin water stores and fluxes (refer to GSFLOW manual) and the output/prms/copper_drought_statvar.dat file with output defined in the gsflow3.control file (refer to GSFLOW Manual), output files also include spatially distributed daily values of total evapotranspiration, canopy evaporation, precipitation, snowfall, infiltration, snow water equivalent, potential evapotranspiration, recharge, sublimation, soil moisture, contributing interflow, water table elevations, changes in groundwater storage, groundwater evapotranspiration, interbasin groundwater flow (limited to the alluvium below the stream outlet), and surface-groundwater exchanges within the river system.

54 ENVIRONMENTAL SCIENCES↗

Sequential Precipitation Input Tagging (SPIT) to Estimate Water Transit Times and Hydrologic Tracer Dynamics Within Water‐Tagging Enabled Hydrologic Models

Determining the age distribution of water exiting a catchment is important for understanding groundwater storage and mixing. New water-tagging capabilities within models track precipitation events as they move through simulated storages, yet forward modeling of individual events may not systematically capture the full transit time distribution (TTD). Here, we present a “sequential precipitation input tagging” (SPIT) framework to tag all input precipitation at regular intervals during extended model simulations. Monthly tags over 7 years were applied at six National Ecological Observatory Network sites to calculate TTDs and derive mean virtual tracer age, $\overline{T_{V}}$, fractions of young water, F yw , and hydrologic tracer concentrations (water isotopes δ 18 O and δ 2 H) within a tagging enabled version of the Weather Research and Forecast hydrologic model (WRF-Hydro). Throughout seven simulation years, the fraction of simulated discharge derived from tagged events, F tag , increased each year, with the final year's F tag ranging from 66% to 100% and highlights the need to apply SPIT over many years to understand TTDs. When the F tag was >75%, simulated $\overline{T_{V}}$ ranged 179–923 days and F yw 0.6%–23.9%, with daily values exhibiting a power-law relationship with precipitation, discharge, and groundwater. Through implementation of SPIT, we find this hydrologic model configuration performs poorly in estimation of $\overline{T_{V}}$ and F yw (root mean squared error of 469 days and 14.4% respectively), suggesting it misrepresents subsurface mixing. Thus, the SPIT framework provides a reproducible approach to calculate watershed transit times within tagging enabled models and thereby assess and improve representation of hydrologic processes.

fraction of young water↗

Spatio–Temporal Machine Learning for Regional to Continental Scale Terrestrial Hydrology

Integrated hydrologic models can simulate coupled surface and subsurface processes but are computationally expensive to run at high resolutions over large domains. Here we develop a novel deep learning model to emulate subsurface flows simulated by the integrated ParFlow–CLM model across the contiguous US. We compare convolutional neural networks like ResNet and UNet run autoregressively against our novel architecture called the Forced SpatioTemporal RNN (FSTR). The FSTR model incorporates separate encoding of initial conditions, static parameters, and meteorological forcings, which are fused in a recurrent loop to produce spatiotemporal predictions of groundwater. We evaluate the model architectures on their ability to reproduce 4D pressure heads, water table depths, and surface soil moisture over the contiguous US at 1 km resolution and daily time steps over the course of a full water year. The FSTR model shows superior performance to the baseline models, producing stable simulations that capture both seasonal and event–scale dynamics across a wide array of hydroclimatic regimes. The emulators provide over 1,000× speedup compared to the original physical model, which will enable new capabilities like uncertainty quantification and data assimilation for integrated hydrologic modeling that were not previously possible. Our results demonstrate the promise of using specialized deep learning architectures like FSTR for emulating complex process–based models without sacrificing fidelity.

54 ENVIRONMENTAL SCIENCES↗

2019 Meander C and Meander Z floodplain groundwater chemistry from the East River Watershed, CO, USA

This dataset includes groundwater geochemistry data from floodplain piezometers collected as a part of the Watershed Function Scientific Focus Area (SFA) located in the Upper Colorado River Basin. The data were collected in order to investigate the role of hyporheic exchange and other river corridor processes on riverine export of solutes. Data includes samples from two intra-meander zones: Meander C, in the Pumphouse vicinity, and Meander Z, just upstream of the confluence with Brush Creek. Floodplain piezometers installed along two transects across Meander C (MCP and MCB wells) and Meander Z (MZA and MZB wells) were sampled on daily to weekly time scales during summer-fall 2019. Some river water grab samples are also included. Data includes in-field measurements (pH, electrical conductivity [EC], oxidation reduction potential [ORP], dissolved oxygen [DO], and groundwater level) along with laboratory measurements (dissolved inorganic carbon [DIC], dissolved organic carbon [DOC], metals and major cations, anions [chloride, sulfate, nitrate], and dissolved ammonium). Files are included in this dataset include: sample locations and depths in both a kmz file which can be opened in Google Earth and a csv file, aqueous geochemistry data in a csv files for Meander C and Meander Z, and analytical detection limits in a csv file. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.

54 ENVIRONMENTAL SCIENCES↗

Field and Model Data Associated with the Manuscript “Drivers of Streamflow Intermittency in Humid Regions: 2. Evaluating Controls on Flow Persistence in an Urbanized Catchment”

This package contains field data, modeling files, and scripts supporting the investigation of the drivers of streamflow intermittency in an urbanized catchment. It includes the field data collected from electrical resistivity tomography (ERT) surveys, distributed temperature sensing (DTS), continuous self-potential (SP) monitoring, groundwater and stilling well. In addition, it contains the data and results of the coupled water- and electrical-flow model developed using the COMSOL Multiphysics and Advanced Terrestrial Simulator (ATS), as well as software files and Jupyter notebooks used to process the data and generate figures in the manuscript submitted for peer review. The data archive is organized in the following directories: 1) Climate Includes hourly precipitation and daily evapotranspiration time series (2024 – 2025) provided as CSV files, alongside a text file detailing dataset units. 2) Coupled_model Field_Application subfolder contains the ATS XML input scripts, data files, output data for the SP site. It also contains the Jupyter notebook (Plot_final_calib.ipynb) to visualize the results of the modeled SP, stream-groundwater exchange and moisture content. The flow model simulation is executed using the ATS XML scripts and the included Python script (generate_data_set.py) to convert ATS output to COMSOL-ready input. COMSOL Multiphysics template (.m can only be used with COMSOL with MATLAB) is executed using the ATS output data to simulate the potential field. 3) Discharge Includes the electrical conductivity (EC) time series (provided as CSV files) from salt slug injections. It also includes the Jupyter notebook (Discharge_process.ipynyb) used to estimate discharge. All discharge measurements collated into rating_curve_processed.csv 4) DTS Contains collated DTS data including raw Stokes and anti-Stokes measurement (provided as .h5 file). It also includes DTS processing.ipynb, a Jupyter notebook for calibrating the DTS data using dts_calibration Python package. cooler_calibration.csv is the DTS calibration CSV used in the calibration sequence. 5) ERT Contains raw resistivity data (provided as CSV files), spatial location of each of the electrodes (provided as CSV files), and files used for the resistivity inversion. 6) Slug_test Includes the slug test data at all the groundwater wells provided as CSV files, as well as the Jupyter notebook (Slug_test.ipynb) for calculating hydraulic conductivity. 7) SP Contains the SP data collected in field at the SP sites (provided as CSV files). 8) Well_data Contains two subfolders: 1) Raw, which provides unprocessed pressure, electrical conductivity and temperature timeseries downloaded from the loggers in all the groundwater and stilling wells, and 2) Processed, which contains sorted, QA/QC timeseries data for each well. The data archive also contains data_process.ipynb, a Jupyter notebook used for field data analysis and generating figures (plotting well, SP, climate, and discharge data, as well as calculating head gradient at sites with nested groundwater wells). Note: Code files (.ipynb, .py, .xml) can be opened in any standard code editor, .exo file can be viewed using Paraview, .h5 files can be opened using HDFView software and h5py Python package, and .resipy file can be opened with the open-source ResIPy software.

ATS↗

Opportunistic Short‐Term Water Uptake Dynamics by Subalpine Trees Observed via In Situ Water Isotope Measurements

Abstract Variations in tree water sources are important to understand in semi‐arid ecosystems because climatic shifts towards lower snowpack and increased drought affect water availability in subalpine forests of the western US. Here, we use daily in situ measurements of stable isotopes ( 2 H & 18 O) in soil and tree stem water, soil matric potential and sap flow to study tree water uptake dynamics. We instrumented three soil profiles down to 90 cm, as well as three aspen and engelmann spruce trees near Gothic, Colorado, in the East River watershed. We observed the fate of natural isotopic variations in rainfall, soil, and plants from June to October 2022, and in August 2023 we conducted a 2 H labeled irrigation experiment. Our observations showed that all studied aspen trees compensated for water scarcity in the shallow soil by shifting the dominant water source at 60(±20) cm to ⅔ of uptake from 90 cm within a few days of a dry period. Both species relied on snowmelt stored in the subsoil to sustain transpiration. Intense rainfall caused the plant water uptake to shift partially to top soil layers within 2 days. Spruce transpiration was lower and relied more on snowmelt, because rainfall infiltration was low in the spruce stand due to high canopy interception. Our findings highlight the important role of snowmelt stored in the deep soil layers for subalpine forest drought response and the dominant fate of monsoonal rainfall to become transpiration rather than recharging groundwater and streams in the Upper Colorado River. Plain Language Summary There is a need to understand how trees in mountainous regions respond to dry conditions that lead to water scarcity, because climate projections suggest that such conditions will become more frequent in the future. Here we present a novel data set of measurements of daily stable isotopes of water across soil profiles and in tree stems of aspen and spruce. Our data show that when the upper soil dried out, aspen trees shifted to using water from deeper layers (beneath 60 cm) to keep transpiring. For spruce trees the uptake pattern is less clear, but both types of trees mainly used snowmelt stored in the deeper soil layers to survive the dry summer. After heavy rain, aspen and spruce trees switched to using water from the top 20 cm of soil. However, for spruce, only some rain reached the soil because the dense tree canopy intercepted it, so spruce trees stayed more dependent on snowmelt and used less water overall. This study shows how important deep snowmelt water is for helping forests survive dry periods and suggests that most summer rain is quickly used by trees rather than replenishing streams and groundwater in the headwaters of the Colorado River. Key Points Tree water resources changed within a few days from snow dominated to higher share of rainfall as soils wetted up after a dry period Compensatory plant water uptake by aspen from the deep layer (90 cm), while uptake from soil depths that became drier (60 cm) declined Strong differences between water sources and availability beneath aspen and spruce, respectively

Sprenger, Matthias↗

Exploring Water System Vulnerabilities in California's Central Valley Under the Late Renaissance Megadrought and Climate Change

Abstract California faces cycles of drought and flooding that are projected to intensify, but these extremes may impact water users across the state differently due to the region's natural hydroclimate variability and complex institutional framework governing water deliveries. To assess these risks, this study introduces a novel exploratory modeling framework informed by paleo and climate‐change based scenarios to better understand how impacts propagate through the Central Valley's complex water system. A stochastic weather generator, conditioned on tree‐ring data, produces a large ensemble of daily weather sequences conditioned on drought and flood conditions under the Late Renaissance Megadrought period (1550–1580 CE). Regional climate changes are applied to this weather data and drive hydrologic projections for the Sacramento, San Joaquin, and Tulare Basins. The resulting streamflow ensembles are used in an exploratory stress test using the California Food‐Energy‐Water System model, a highly resolved, daily model of water storage and conveyance throughout California's Central Valley. Results show that megadrought conditions lead to unprecedented reductions in inflows and storage at major California reservoirs. Both junior and senior water rights holders experience multi‐year periods of curtailed water deliveries and complete drawdowns of groundwater assets. When megadrought dynamics are combined with climate change, risks for unprecedented depletion of reservoir storage and sustained curtailment of water deliveries across multiple years increase. Asymmetries in risk emerge depending on water source, rights, and access to groundwater banks.

Gupta, Rohini S. [School of Civil and Environmenta↗

Field and Model Data Associated with the Manuscript “Drivers of Streamflow Intermittency in Humid Regions: 1. Evaluating Above- and Below-ground Controls of Flow Persistence in a Forested Catchment”

This package contains field data, modeling files, and scripts supporting the investigation of the drivers of streamflow intermittency in a forested catchment. It includes the field data collected from electrical resistivity tomography (ERT) surveys, ground penetrating radar (GPR), continuous self-potential (SP) monitoring, electromagnetic (EM) imaging, groundwater and stilling well. In addition, it contains the data and results of the coupled water- and electrical-flow model developed using the COMSOL Multiphysics and Advanced Terrestrial Simulator (ATS), as well as software files and Jupyter notebooks used to process the data and generate figures in the manuscript submitted for peer review. The data archive is organized in the following directories: 1) Climate Includes hourly precipitation and daily evapotranspiration time series (2024 – 2025) provided as CSV files, alongside a text file detailing dataset units. 2) Coupled_model Contains two subfolders: Synthetic and Field_Application subfolder. Synthetic subfolder contains the ATS XML input script (can be opened using any code editor) for the four synthetic hydrological cases tested (Connected and gaining, Connected and losing, Disconnected and losing, and dry stream). It also includes other experimental cases to test the influence of precipitation and concentration gradient. For each synthetic case, the flow model simulation is executed using the ATS XML scripts and the included Python script (generate_data_set.py) to convert ATS output to COMSOL-ready input. COMSOL Multiphysics template (.mph can be opened with the commercial software COMSOL and requires a license) is executed using the ATS output data to simulate the potential field. It also includes the Synthetic_model_plot.ipynb (can be opened using any code editor) to visualize the SP result and generate manuscript figures. The data subfolder contains mesh files to run both the ATS (.exo and .stl files can be viewed using Paraview; .h5 files can be opened using HDFView software and h5py Python package) and COMSOL models. Field_Application subfolder contains two subfolders: ES_MDA_inversion and Final_Model. ES_MDA_inversion contains the Python script (.py can be opened using any code editor) and SP observation data used to run the Ensemble Smoother with Multiple Data Assimilation (ES-MDA) inversion sequence to get the optimal model parameters. The Final_model subfolder contains the ATS XML input scripts, data files, output data for the two SP sites. The same workflow steps outlined for the Synthetic subfolder apply here. It also contains the Jupyter notebook (Plot_final_calib.ipynb) to visualize the results of the modeled SP, stream-groundwater exchange and moisture content. 3) Discharge Includes the electrical conductivity (EC) time series (provided as CSV files) from salt slug injections. It also includes the Jupyter notebook (Discharge_process.ipynyb) used to estimate discharge. All discharge measurements collated into rating_curve_processed.csv 4) EM Contains the CSV file of the EM data from the DUALEM-42, including spatial coordinates (x, y, z), apparent conductivity, and in-phase measurements at 2 m coil separations for horizontal coplanar (HCP) and perpendicular (PRP) geometries. 5) ERT Contains raw resistivity data (provided as CSV files), spatial location of each of the electrodes (provided as CSV files), and files used for the resistivity inversion (.resipy can be opened with the open-source ResIPy software). 6) GPR Includes GPR field datasets collected at 100 MHz and 250 MHz antenna frequencies, along with the processing/interpretation project file (GPR_process.gpz can be viewed using EKKO_Project 6, a commercial software by Sensors & Software that requires a license). 7) Slug_test Includes the slug test data at all the groundwater wells provided as CSV files, as well as the Jupyter notebook (Slug_test.ipynb) for calculating hydraulic conductivity. 8) SP Contains the SP data collected in field at the two SP sites (one in the perennial reach and the other in the intermittent reach), provided as DAT files. 9) Well_data Contains two subfolders: 1) Raw, which provides unprocessed pressure, electrical conductivity and temperature timeseries downloaded from the loggers in all the groundwater and stilling wells, and 2) Processed, which contains sorted, QA/QC timeseries data for each well. The data archive also contains data_process.ipynb, a Jupyter notebook used for field data analysis and generating figures (plotting well, SP, climate, and discharge data, as well as calculating head gradient at sites with nested groundwater wells). It also includes DTW.ipynb, a Jupyter notebook containing the code for the dynamic time warping (DTW) with sliding window to evaluate SP signal synchronicity.

ATS↗

BSEC ecohydrological and water quality fluxes from RHESSys Simulations in USGS gauged watersheds

Baltimore Environmental Social Collaborative (BSEC) Water and Water Quality Simulations from RHESSys Model The repository contains RHESSys (Tague & Band, 2004; source code) simulated ecohydrological and nutrient (nitrogen only) fluxes at daily, basin-average (RHESSys_basin_output) and monthly, grid (RHESSys_patch_output) levels. We currently simulated the following 8 watersheds in Baltimore: Dead Run Baisman Run Scotts Level Branch Moores Run Powder Mill Run Maidens Choice Run Stony Run The watershed boundaries of all studied watersheds are stored in Watershed_Boundary folder. Variables and their units are listed in the metadata. Spatial projection, NAD83 / UTM zone 18N (EPSG:26918) is used for patch-level, netCDF-format files. For more information, please contact Ruoyu Zhang (rz3jr@virginia.edu).

Baltimore MD↗

Transportability of exogenous microbial community correlates with interwell connectivity in deep aquifers

Subsurface resource engineering operations often utilize continuous injection of externally-sourced water into geological reservoirs for formation pressure maintenance, resource recovery or energy/waste storage. Such injected water generally contains naturally occurring microbes. Little is known, however, about how the injectate microbes transport through geological media as a community, how such transportability is affected by injector-producer connectivity, and whether such knowledge can be utilized for flowpath characterization. In this study, we analyzed daily-to-weekly timeseries microbial community data from the injected- and produced-fluids of a ten-month flow test at a deep, well-characterized engineered aquifer. We found that the injectate microbial community was distinct from the indigenous community at the amplicon sequence variant (ASV) level, and that the transportability of injectate community towards a given producer, quantified by an “nASV-Overlap” metric we propose, had strong and significant positive correlation with known injector-producer connectivities at our site. This suggests that the better the connectivity, the higher the probability for more injectate species to flow through the interwell region and arrive at a producer. Because interwell connectivity is an important yet usually unknown parameter in subsurface resource engineering, such correlation in turn points to nASV-Overlap as a useful indicator of interwell connectivity for aquifer characterization and long-term monitoring. Based on our findings, an nASV-Overlap-based microbial tracing approach was developed for characterizing and monitoring the relative connectivities across multiple producers with a given injector. A side-by-side comparison between the new nASV-Overlap approach and traditional artificial tracer methods is presented, and their respective strengths and limitations are discussed.

Deep biosphere↗

HarDWR - Harmonized Water Rights Records

A dataset within the Harmonized Database of Western U.S. Water Rights (HarDWR). For a detailed description of the database, please see the meta-record v2.0. Changelog v2.0 - Recalculated based on data sourced from WestDAAT - Changed using a Site ID column to identify unique records to using aa combination of Site ID and Allocation ID - Removed the Water Management Area (WMA) column from the harmonized records. The replacement is a separate file which stores the relationship between allocations and WMAs. This allows for allocations to contribute to water right amounts to multiple WMAs during the subsequent cumulative process. - Added a column describing a water rights legal status - Added "Unspecified" was a water source category - Added an acre-foot (AF) column - Added a column for the classification of the right's owner v1.02 - Added a .RData file to the dataset as a convenience for anyone exploring our code. This is an internal file, and the one referenced in analysis scripts as the data objects are already in R data objects. v1.01 - Updated the names of each file with an ID number less than 3 digits to include leading 0s v1.0 - Initial public release Description Here we present an updated database of Western U.S. water right records. This database provides consistent unique identifiers for each water right record, and a consistent categorization scheme that puts each water right record into one of seven broad use categories. These data were instrumental in conducting a study of the multi-sector dynamics of inter-sectoral water allocation changes though water markets (Grogan et al., *in review*). Specifically, the data were formatted for use as input to a process-based hydrologic model, Water Balance Model (WBM), with a water rights module (Grogan et al., *in review*). While this specific study motivated the development of the database presented here, water management in the U.S. West is a rich area of study (e.g., Anderson and Woosly, 2005; Tidwell, 2014; Null and Prudencio, 2016; Carney et al., 2021) so releasing this database publicly with documentation and usage notes will enable other researchers to do further work on water management in the U.S. West. We produced the water rights database presented here in four main steps: (1) data collection, (2) data quality control, (3) data harmonization, and (4) generation of cumulative water rights curves. Each of steps (1)-(3) had to be completed in order to produce (4), the final product that was used in the modeling exercise in Grogan et al. (*in review*). All data in each step is associated with a spatial unit called a Water Management Area (WMA), which is the unit of water right administration utilized by the state in which the right came from. Steps (2) and (3) required use to make assumptions and interpretation, and to remove records from the raw data collection. We describe each of these assumptions and interpretations below so that other researchers can choose to implement alternative assumptions an interpretation as fits their research aims. Motivation for Changing Data Sources The most significant change has been a switch from collecting the raw water rights directly from each state to using the water rights records presented in WestDAAT, a product of the Water Data Exchange (WaDE) Program under the Western States Water Council (WSWC). One of the main reasons for this is that each state of interest is a member of the WSWC, meaning that WaDE is partially funded by these states, as well as many universities. As WestDAAT is also a database with consistent categorization, it has allowed us to spend less time on data collection and quality control and more time on answering research questions. This has included records from water right sources we had previously not known about when creating v1.0 of this database. The only major downside to utilizing the WestDAAT records as our raw data is that further updates are tied to when WestDAAT is updated, as some states update their public water right records daily. However, as our focus is on cumulative water amounts at the regional scale, it is unlikely most records updates would have a significant effect on our results. The structure of WestDAAT led to several important changes to how HarWR is formatted. The most significant change is that WaDE has calculated a field known as `SiteUUID`, which is a unique identifier for the Point of Diversion (POD), or where the water is drawn from. This separate from `AllocationNativeID`, which is the identifier for the allocation of water, or the amount of water associated with the water right. It should be noted that it is possible for a single site to have multiple allocations associated with it and for an allocation to be able to be extracted from multiple sites. The site-allocation structure has allowed us to adapt a more consistent, and hopefully more realistic, approach in organizing the water right records than we had with HarDWR v1.0. This was incredibly helpful as the raw data from many states had multiple water uses within a single field within a single row of their raw data, and it was not always clear if the first water use was the most important, or simply first alphabetically. WestDAAT has already addressed this data quality issue. Furthermore, with v1.0, when there were multiple records with the same water right ID, we selected the largest volume or flow amount and disregarded the rest. As WestDAAT was already a common structure for disparate data formats, we were better able to identify sites with multiple allocations and, perhaps more importantly, allocations with multiple sites. This is particularly helpful when an allocation has sites which cross WMA boundaries, instead of just assigning the full water amount to a single WMA we are now able to divide the amount of water between the number of relevant WMAs. As it is now possible to identify allocations with water used in multiple WMAs, it is no longer practical to store this information within a single column. Instead the stAllocationToWMATab.csv file was created, which is an allocation by WMA matrix containing the percent Place of Use area overlap with each WMA. We then use this percentage to divide the allocation's flow amount between the given WMAs during the cumulation process to hopefully provide more realistic totals of water use in each area. However, not every state provides areas of water use, so like HarDWR v1.0, a hierarchical decision tree was used to assign each allocation to a WMA. First, if a WMA could be identified based on the allocation ID, then that WMA was used; typically, when available, this applied to the entire state and no further steps were needed. Second was the spatial analysis of Place of Use to WMAs. Third was a spatial analysis of the POD locations to WMAs, with the assumption that allocation's POD is within the WMA it should belong to; if an allocation still had multiple WMAs based on its POD locations, then the allocation's flow amount would be divided equally between all WMAs. The fourth, and final, process was to include water allocations which spatially fell outside of the state WMA boundaries. This could be due to several reasons, such as coordinate errors / imprecision in the POD location, imprecision in the WMA boundaries, or rights attached with features, such as a reservoir, which crosses state boundaries. To include these records, we decided for any POD which was within one kilometer of the state's edge would be assigned to the nearest WMA. Other Changes WestDAAT has Allowed In addition to a more nuanced and consistent method of assigning water right's data to WMAs, there are other benefits gained from using the WestDAAT dataset. Among those is a consistent categorization of a water right's legal status. In HarDWR v1.0, legal status was effectively ignored, which led to many valid concerns about the quality of the database related to the amounts of water the rights allowed to be claimed. The main issue was that rights with legal status' such as "application withdrawn", "non-active", or "cancelled" were included within HarDWR v1.0. These, and other water rights status' which were deemed to not be in use have been removed from this version of the database. Another major change has been the addition of the "unspecified water source category. This is water that can come from either surface water or groundwater, or the source of which is unknown. The addition of this source category brings the total number of categories to three. Due to reviewer feedback, we decided to add the acre-foot (AF) column so that the data may be more applicable to a wider audience. We added the ownerClassification column so that the data may be more applicable to a wider audience. File Descriptions The dataset is a series of various files organized by state sub-directories. In addition, each file begins with the state's name, in case the file is separate from its sub-directory for some reason. After the state name is the text which describes the contents of the file. Here is each file described in detail. Note that st is a placeholder for the state's name. stFullRecords_HarmonizedRights.csv: A file of the complete water records for each state. The column headers for each of this type of file are: state - The name of the state to which the allocations belong to. FIPS - The two digit numeric state ID code. siteID - The site location ID for POD locations. A site may have multiple allocations, which are the actual amount of water which can be drawn. In a simplified hypothetical, a farm stead may have an allocation for "irrigation" and an allocation for "domestic" water use, but the water is drawn from the same pumping equipment. It should be noted that many of the site ID appear to have been added by WaDE, and therefore may not be recognized by a given state's water rights database. allocationID - The allocation ID for the water right. For most states this is the water right ID, and what is recommended to use should a right be looked up on a given state's water rights database. The water amounts associated with these IDs tend to be finer scaled than those associated with siteID. It should be noted that some allocations may be extracted from multiple sites, particularly for larger Places of Use. ownerClassification - A classification of the types of owners for water rights. The most common is `Private` which incorporates a wide range of entities. Several classifications would be grouped into a government category, most of which are for the U.S. Federal Government. These allocations could be listed as "Federal", "United States of America", or as the names of any number of federal agencies. The last major grouping of entities is for "Native American"s. priorityDate - The date we use as the water right priority date for our modeling analysis. This is the legal priority date when it is available. However, for some rights, specifically from California and New Mexico, we used a pseudo priority date (e.g. well completion date or start of well drilling date) when a legal priority date was not available. The most questionable dates come from New Mexico, where the only date associated with certain water right records was the date the allocation was recorded in the database. As the allocation record creation tended to be within a few months of the filing of the application of the water right, from manually double checking the water rights, and our analysis focuses on aggregating water rights on the timescale of years, we determined it was acceptable to use such dates to include as many records as possible. primaryBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories WestDAAT. This column is the original WaDE category for the primary water use at the PoD site. allocationBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories for WestDAAT. This column is the original WaDE category

Economics↗