Search NASA⌕ Search

SEARCH · Search NASA

Results for “pattern dictionary”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

A Pattern Dictionary Method for Anomaly Detection

In this paper, we propose a compression-based anomaly detection method for time series and sequence data using a pattern dictionary. The proposed method is capable of learning complex patterns in a training data sequence, using these learned patterns to detect potentially anomalous patterns in a test data sequence. The proposed pattern dictionary method uses a measure of complexity of the test sequence as an anomaly score that can be used to perform stand-alone anomaly detection. We also show that when combined with a universal source coder, the proposed pattern dictionary yields a powerful atypicality detector that is equally applicable to anomaly detection. The pattern dictionary-based atypicality detector uses an anomaly score defined as the difference between the complexity of the test sequence data encoded by the trained pattern dictionary (typical) encoder and the universal (atypical) encoder, respectively. We consider two complexity measures: the number of parsed phrases in the sequence, and the length of the encoded sequence (codelength). Specializing to a particular type of universal encoder, the Tree-Structured Lempel–Ziv (LZ78), we obtain a novel non-asymptotic upper bound, in terms of the Lambert W function, on the number of distinct phrases resulting from the LZ78 parser. This non-asymptotic bound determines the range of anomaly score. As a concrete application, we illustrate the pattern dictionary framework for constructing a baseline of health against which anomalous deviations can be detected.

97 MATHEMATICS AND COMPUTING↗

Report on PTT Imaging of Defects in AM Metallic Materials-Part 2

Metal Additive Manufacturing (AM) is a promising method for cost-efficient fabrication of complex shape structures for applications in harsh environment, such as in a nuclear reactor. However, internal defects (pores) occur in high-strength AM alloys, which are manufactured with Laser Powder Bed Fusion (LPBF) AM method. Pulsed Infrared Thermography (PIT) is an efficient nondestructive evaluation (NDE) method to examine actual structures, because this method offers one-sided non-contact measurements, and fast processing of large sample areas. However, imaging of material defects, particularly defects with sizes at microscopic level, is challenging. In this report, we benchmark the performance of several Unsupervised Learning (UL) algorithms designed to enhance imaging of microscopic defects in metals with PIT. UL aims to learn the latent principal patterns (dictionaries) in PIT data to detect defects with minimal human supervision. Performance of Independent Component Analysis (ICA), Sparse Coding (SC), Principal Component Analysis (PCA) and Exploratory Factor Analysis (EFA) was compared using F-score, UL model training time and defects reconstruction time. We obtained the average F-score of 0.75, and a highest F-score of 0.89 for the EFA algorithm. Overall, EFA outperforms other UL algorithms considered in this study.

36 MATERIALS SCIENCE↗

Pulsed Thermal Tomography Nondestructive Examination of Additively Manufactured Reactor Materials and Components (Final Technical Report)

Metal Additive Manufacturing (AM) is a promising method for cost-efficient fabrication of complex shape structures for applications in harsh environment, such as in a nuclear reactor. However, internal defects (pores) occur in high-strength AM alloys, which are manufactured with Laser Powder Bed Fusion (LPBF) AM method. Pulsed Infrared Thermography (PIT) is an efficient nondestructive evaluation (NDE) method to examine actual structures, because this method offers one-sided non-contact measurements, and fast processing of large sample areas. However, imaging of material defects, particularly defects with sizes at microscopic level, is challenging. In this report, we benchmark the performance of several Unsupervised Learning (UL) algorithms designed to enhance imaging of microscopic defects in metals with PIT. UL aims to learn the latent principal patterns (dictionaries) in PIT data to detect defects with minimal human supervision. Performance of Independent Component Analysis (ICA), Sparse Coding (SC), Principal Component Analysis (PCA) and Exploratory Factor Analysis (EFA) was compared using F-score, UL model training time and defects reconstruction time. We obtained the average F-score of 0.75, and a highest F-score of 0.89 for the EFA algorithm. Overall, EFA outperforms other UL algorithms considered in this study. In another approach, we investigate Thermal Tomography (TT), which is a computational method for reconstruction of depth profile of internal material defects from PIT nondestructive evaluation (NDE). TT algorithm obtains depth reconstructions of thermal effusivity, which has been shown to provide visualization of subsurface internals defects in metals. In many applications, one needs to determine the defect shape and orientation from reconstructed effusivity images. Interpretation of TT images is non-trivial because of blurring, which increases with depth due to heat diffusion-based nature of image formation. We have developed a deep learning convolutional neural network (CNN) to classify size and orientation of subsurface material defects in TT images. CNN was trained with TT images produced with computer simulations of 2D metallic structures (thin plates) containing elliptical subsurface voids. Performance of CNN was investigated using test TT images developed with computer simulations of plates containing elliptical defects, and defects with shape imported from scanning electron microscopy (SEM) images. CNN demonstrated the ability to classify radii and angular orientation of elliptical defects in previously unseen test TT images. We have also demonstrated that CNN trained on TT images of elliptical defects is capable of classifying shape and orientation of irregular defects. Training the CNN on irregular defect shapes instead of on elliptical shapes would make the resulting classifications more descriptive of actual defect shapes. However, this requires a much higher volume of SEM images of material defects, which are difficult to obtain because of random occurrence of defects in LPBF. To address this challenge, we developed a generative adversarial network (GAN) to augment the existing dataset of SEM defect images. The GAN model is demonstrated to create novel yet realistic defect shapes that can be used as input for simulated PTT images to train CNN. We also investigate several approaches based on Gaussian Random Circle and Bezier Curves for constructing parametric models of irregular-shape defects.

36 MATERIALS SCIENCE↗

Historic climate, cosmogenic 10Be, denudation-rate, and geospatial datasets from the Pikes Peak region, Colorado, USA

This data package contains geographic information system (GIS) layers and tabular datasets associated with the study of elevation-dependent denudation rates on Pikes Peak in the Front Range of the Rocky Mountains, Colorado, USA. The package includes GIS layers used to produce the study-area map, including sample locations, sample watershed boundaries, the Pikes Peak batholith, Pleistocene glacier extent, weather station locations, and elevation and hillshade rasters, together with comma-separated value (CSV) tables and matching CSV data dictionaries. These mapped layers provide the geographic framework for interpreting denudation patterns across the Pikes Peak region and for relating sample locations to watershed geometry, bedrock setting, glacial history, and nearby climate stations. The first group of tables reports climate and geospatial context for the study area. These files include station-based temperature and precipitation data used to characterize elevational gradients in mean annual climate and monthly climate seasonality, sample locations, denudation-rate and topographic metrics, fixed frost-cracking model parameters, frost-cracking intensity and precipitation-frequency metrics, and stream-power inversion results. Together, these data provide the basis for evaluating how denudation varies with elevation, climate, and landscape form across sampled catchments on Pikes Peak. The second group of tables reports cosmogenic nuclide and erosion-model results used in the denudation analysis. Included files contain accelerator mass spectrometry (AMS) measurements for in situ-produced cosmogenic beryllium-10 (10Be), including sample identifiers, measured 10Be:9Be ratios, analytical uncertainties, carrier mass, quartz mass, blank corrections, blank-group statistics, and calculated 10Be concentrations and uncertainties. Additional tables summarize stream-power-law inversion results for sampled catchments, including optimized model parameters, predicted erosion rates, residual metrics, channel-pixel counts, and convergence status, as well as regression equations and summary statistics used to evaluate relationships among elevation, climate, frost cracking, precipitation forcing, and denudation rate. The package contains GIS files, comma-separated value files (.csv), Microsoft Excel files (.xlsx), CSV data dictionaries, a file-level metadata table, and a readme text file.

10Be cosmogenic nuclides↗

Data from "A Bayesian Record Linkage Approach to Applications in Tree Demography Using Overlapping LiDAR Scans"

Processed LiDAR data and environmental covariates from 2015 and 2019 LiDAR scans in the Vicinity of Snodgrass Mountain (Western Colorado, USA), in a geographic subset used in primary analysis for the research paper.This package contains LiDAR-derived canopy height maps for 2015 and 2019, crown polygons derived from the height maps using a segmentation algorithm, and environmental covariates supporting the model of forest growth. Source datasets include August 2015 and August 2019 discrete-return LiDAR point clouds collected by Quantum Geospatial for terrain mapping purposes on behalf of the Colorado Hazard Mapping Program and the Colorado Water Conservation Board. Both datasets adhere to the USGS QL2 quality standard. The point cloud data were processed using the R package lidR to generate a canopy height model representing maximum vegetation height above the ground surface, using a pit-free algorithm.This dataset was compiled to assess how spatial patterns of tree growth in montane and subalpine forests are influenced by water and energy availability. Understanding these growth patterns can provide insight into forest dynamics in the Southern Rocky Mountains under changing climatic conditions.This dataset contains .tif, .csv, and .txt files. This dataset additionally includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.

54 ENVIRONMENTAL SCIENCES↗

Model Inputs, Outputs, and Scripts associated with: “Combined effects of stream hydrology and land use on basin-scale hyporheic zone denitrification in the Columbia River Basin”

This data package is associated with the publication “Combined effects of stream hydrology and land use on basin‐scale hyporheic zone denitrification in the Columbia River Basin”, published in Water Resource Research (Son et al.2022) available at https://doi.org/10.1029/2021WR031131. This data package includes the key model inputs/outputs of the river corridor model for the Columbia River Basin (CRB) and the model source codes used in the manuscript. The model is a carbon-nitrogen-coupled river corridor model (RCM), and the model is used to quantify hyporheic zone (HZ) denitrification at the NHDPLUS stream reach scales. The RCM used in this study combines empirical substrate models derived from observations and three microbially driven reactions, including two-step denitrification and aerobic respiration, are considered within the HZ. The key input data of the model are exchange flux, residence time, and stream solute (dissolved organic carbon (DOC), dissolved oxygen (DO), and nitrate concentrations). These inputs are constant over time and represent long-term averaged values. This study uses the RCM to explore the spatial patterns of HZ denitrification across reaches with different sizes and land use in the CRB. Our main objective is to use the RCM as a virtual reality model, and the machine-learning models as surrogates that encapsulate the complexities of the physics-based model while identifying the importance of different variables that are not evident in the model conceptualization. We do not include a direct comparison of the modeled HZ denitrification and measurements; however, the RCM can capture the overall spatial patterns of the HZ denitrification because the model inputs and its reaction networks are based on well-established theory and a physical-based model. The combination of the model-based predictions and a machine-learning approach (e.g., random forest) is used to improve our understanding of what variables of the model are associated with spatial patterns of the modeled denitrification across reaches with different sizes and land uses, and to develop a proxy model using measurable variables to reproduce the simulated patterns.This dataset contains five folders: (1) model_inputs, (2) model_outputs, (3) Rscripts, (4) figures, and (5) model_codes. It also contains a readme, file level metadata (FLMD), and data dictionary (dd). Please see the FLMD for a list of all the files contained in this data package and descriptions for each. The model_inputs folder contains the model inputs used to drive the model simulations. The model_outputs folder contains key model output files from the river corridor model. The Rscripts folder contains the Rscripts for pre- and post- processing model results. The figures folder contains the raw figures associated with the manuscript. The model_codes folder includes key model source codes/input files. All files are .jpg, .jpeg, .out, .e, .od, .dat, .sub, .F90, .0, .R, .sbx, .cpg, .sbn, .shx, .shp, .dbf, .prj, .tfw, .tif, .xml, .pdf, or .csv.

54 ENVIRONMENTAL SCIENCES↗

Temperature, Humidity, and Time-Lapse Video Data from the East River Watershed, Water Year 2024

A new version of this dataset is available at doi:10.15485/3001338 and is the first citation in the 'Related References' section. It is expands on this dataset by appending another water year of data collection and additional logger sites.This dataset contains time-lapse imagery and distributed measurements of air temperature, relative humidity, dew point, and soil temperature across the East River basin from 3 October 2023 to 12 August 2024. Instruments were deployed at 14 sites as part of the DOE Grant: Seasonal Cycles Unravel Mysteries of Missing Mountain Water organized by Jessica Lundquist (University of Washington), Rosemary Carroll (Desert Research Institute), and Ethan Gutmann (National Center for Atmospheric Research). The data are published to support studies of surface climate or hydrologic processes in complex terrain. Measurements were collected with low-cost data loggers installed 2 m high on evergreen trees or buried just below the soil surface. Time-lapse cameras were deployed at three sites. Imagery from sites AP BONUS and AP5 provides insight into large-scale seasonal snow cover variability. Imagery from site EL2 shows smaller-scale snow patterns across a nearby meadow.Dataset files are organized by site and variable (air measurements, ground measurements, or time-lapse video). Air and ground measurements are packaged in LoggerData.zip, and time-lapse imagery is compiled into short videos stored in TimelapseVideos.zip. File-level metadata contains details for each file included in the dataset. A data dictionary provides units and descriptions for column or row names in all files. The locations metadata file describes site characteristics, locations, and associated GPS methods.Dataset update 2025-03-03: Resolved header and datetime formatting inconsistencies within LoggerData.zip files KP1_Air, KP3_Air, KP3_Ground, KP6_Ground, AP3_Ground, AP4_Air, and AP6_Air.Dataset update 2025-11-18: Modified abstract and related references sections to include new version of dataset.

54 ENVIRONMENTAL SCIENCES↗

Temperature, Humidity, and Time-Lapse Video Data from the East River Watershed, Water Years 2024 and 2025

This dataset contains time-lapse imagery and distributed measurements of air temperature, relative humidity, dew point, and soil temperature across the East River basin from 3 October 2023 to 8 August 2025. Instruments were deployed at 19 sites as part of the DOE Grant: Seasonal Cycles Unravel Mysteries of Missing Mountain Water organized by Jessica Lundquist (University of Washington), Rosemary Carroll (Desert Research Institute), and Ethan Gutmann (National Center for Atmospheric Research). The data are published to support studies of surface climate or hydrologic processes in complex terrain. Measurements were collected with low-cost data loggers installed 2 m high on evergreen trees or buried just below the soil surface. Time-lapse cameras were deployed at three sites. Imagery from sites AP BONUS and AP5 (Avery Picnic) provides insight into large-scale seasonal snow cover variability. Imagery from site EL2 (Emerald Lake) shows smaller-scale snow patterns across a nearby meadow. Dataset files are organized by site and variable (air measurements, ground measurements, or time-lapse video). Air and ground measurements are packaged in LoggerData.zip, and time-lapse imagery is compiled into short videos stored in TimelapseVideos.zip. File-level metadata contains details for each file included in the dataset. A data dictionary provides units and descriptions for column or row names in all files. The locations metadata file describes site characteristics, locations, and associated GPS methods.

54 ENVIRONMENTAL SCIENCES↗

Coarse-grained fixed-point tensor networks and holographic reflected entropy in 3D gravity

We use the framework of fixed-point BCFT tensor networks to present a microscopic CFT derivation of the correspondence between reflected entropy (RE) and entanglement wedge cross section (EW) in AdS 3 /CFT 2 , for both bipartite and multipartite settings. These fixed-point tensor networks, obtained by triangulating Euclidean CFT path integrals, allow us to explicitly construct the canonical purification via cutting-and-gluing CFT path integrals. Employing modular flow in the large-c limit, we demonstrate that these intrinsic CFT manipulations reproduce bulk geometric prescriptions, without assuming the AdS/CFT dictionary. The emergence of bulk geometry is traced to coarse-graining over heavy states in the large-c limit. Universal coarse-grained BCFT data for compact 2D CFTs, through the relation to Liouville theory with ZZ boundary conditions, yields hyperbolic geometry on the Cauchy slice. The corresponding averaged replica partition functions reproduce all candidate EWs, arising from different averaging patterns, with the dominant one providing the correct RE and EW. In this way, many heuristic tensor-network intuitions in toy models are made precise and established directly from intrinsic CFT data.

AdS-CFT correspondence↗

Geophysical survey associated with NEON AOP survey, East River, CO 2018

The package contains data layers developed and used in Falco et al. 2024: “EcoImaging: Advanced Sensing to Investigate Plant and Abiotic Hierarchical Spatial Patterns in Mountainous Watersheds". The package is part of the DOE Watershed Function Science Focus Area (SFA) project and includes geophysical measurements collected at the East River, Colorado, in conjunction with the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP) survey conducted in June 2018. This dataset provide soil geophysical information and were used to investigate soil-plant relationships. The dataset consists of: - NEON_2018_EMI_survey.zip: the electromagnetic induction (EMI) survey as shape-file; - NEON_plot_TDR.csv: plot‑level data from Time‑Domain Reflectometry (TDR) measurements, providing: * volumetric water content (VWC) in percent (%); * soil temperature in degrees Celsius (°C); - file level metadata (flmd.csv) - data dictionary (dd.csv) file This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

A gradient-based deep neural network model for simulating multiphase flow in porous media

We report simulation of multiphase flow in porous media is crucial for the effective management of subsurface energy and environment-related activities. The numerical simulators used for modeling such processes rely on spatial and temporal discretization of the governing mass and energy balance partial-differential equations (PDEs) into algebraic systems via finite-difference/volume/element methods. These simulators usually require dedicated software development and maintenance, and suffer low efficiency from a runtime and memory standpoint for problems with multi-scale heterogeneity, coupled-physics processes or fluids with complex phase behavior. Therefore, developing cost-effective, data-driven models can become a practical choice, and in this work, we choose deep learning approaches as they can handle high dimensional data and accurately predict state variables with strong nonlinearity. In this paper, we describe a gradient-based deep neural network (GDNN) constrained by the physics related to multiphase flow in porous media. We tackle the nonlinearity of flow in porous media induced by rock heterogeneity, fluid properties, and fluid-rock interactions by decomposing the nonlinear PDEs into a dictionary of elementary differential operators. We use a combination of operators to handle rock spatial heterogeneity and fluid flow by advection. Since the augmented differential operators are inherently related to the physics of fluid flow, we treat them as first principles prior knowledge to regularize the GDNN training. We use the example of pressure management at geologic CO 2 storage sites, where CO 2 is injected in saline aquifers and brine is produced, and apply GDNN to construct a predictive model that is trained with physics-based simulation data and emulates the physics process. We demonstrate that GDNN can effectively predict the nonlinear patterns of subsurface responses, including the temporal and spatial evolution of the pressure and saturation plumes. We also successfully extend the GDNN to convolutional neural network (CNN), namely gradient-based CNN (GCNN), and validate its capability to improve the prediction accuracy. GDNN has great potential to tackle challenging problems that are governed by highly nonlinear physics and enable the development of data-driven models with higher fidelity.

42 ENGINEERING↗

Dendrometer data at The Morton Arboretum Forestry Plots 2019-2023

We are collecting long-term dendrometer data at The Morton Arboretum to determine seasonal growth patterns in trees. This data on stem growth patterns will eventually be integrated with other ongoing data streams to paint a broader picture of plant phenology. This data package contains the outputs from three types of dendrometer devices: ICT band dendrometer data is in the "ICT_data2020_2023.csv" file, TreeHugger band dendrometer data is in the "TreeHugger_data2019_2020.csv" file, and TOMST point dendrometer data is in the "TOMST_data2021_2023.csv" file. The TreeHugger and TOMST files contain corrected and raw uncorrected values, whereas the ICT data file contains only raw data. Initial tree size data at device installation is located in the "Initial_Tree_Size.csv" file. Additional information on units are contained within each data file's respective data dictionary, and the location metadata file contains the geographic locations of the 23 forestry plots at The Morton Arboretum as well as the measured tree species at each location.

54 ENVIRONMENTAL SCIENCES↗

A Simulation Atlas of Tidal Features in Galaxies

Detailed simulations of tidally induced structure in disk galaxies have either concentrated on specific systems or consisted of a few encounters with relatively small numbers of particles and no self-gravity. Observers need a 'dictionary' of simulations that covers many encounter parameters with fine morphological resolution and includes effects of self-gravitation. Observers can then search the dictionary for the parameters that best match a particular observed morphology. Alternatively, the dictionary can be used with observational samples for statistical studies of system parameters. To fill this need, we present a survey of model tidal encounters using a self-gravitating, 180,000 particle, two-component ('stars' and 'gas') disk. A wide variety of fascinating morphologies results. There are 86 different encounters that vary orbit tilt, perigalacticon distance, galaxy to companion mass ratio, and the amount of halo dark matter relative to the disk. For morphological comparisons, over 1700 images of the entire survey are available in video form. While there is a rich variety of tidal structure covering much of this parameter space, some general patterns may be remarked. There is a strong orbital inclination dependence of the symmetry of tidal patterns, most symmetric for planar orbits and nearly one-sided for polar encounters. Retrograde encounters produce only broad fanlike global patterns, but rich small-scale internal structure. In both kinds of encounter, our numerical resolution allows us to track internal spiral structure driven by the outer material arms, especially in the lighter halo simulations. We note also that polar encounters generate series of expanding, essentially non-rotating loops resembling shell structures in some respects.

Howard, Sethanne↗

Leaf phenology data at The Morton Arboretum Forestry Plots 2019-2023

We are collecting long-term leaf phenology data at The Morton Arboretum to determine seasonal patterns of leaf production in trees. This data on leaf phenology will be integrated with other ongoing data streams to create a connection between above- and below-ground tree processes. This data package contains raw and smooth outputs from phenology data, as well as extracted phenophase dates (i.e., start, peak, and end of season): the raw and smooth outputs from the PhenoCam GUI can be found in the "leafRaw.csv" and "leafSmooth.csv" files, respectively, and the extracted phenophase dates can be found in the "leafPhenophaseDates2019-2023.csv" file. Extracted phenophase dates for evergreen species in 2023 are currently unavailable, and the files will be updated once they are extracted. Additional information on units and other file-level metadata can be found within each data file's respective data dictionary, and metadata for each of the 23 surveyed plots can be found within the "Location_metadata.csv" file. While the "leafRaw.csv" and the "leafSmooth.csv" files contain all data for all species, the "leafPhenophaseDates2019-2023.csv" file currently excludes the dates for evergreen species in 2023. Another version of the file will be added as those dates are extracted.

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with "Coupled primary production and respiration in a large river contrasts with smaller rivers and streams."

This data package is associated with the publication "Coupled primary production and respiration in a large river contrasts with smaller rivers and streams." in review at Limnology and Oceanography (Roley et al. 2023). This study focuses on understanding ecosystem metabolism for the Hanford Reach of the Columbia River in Washington state, a free-flowing stretch with a substantial discharge. Large rivers have been overlooked compared to small and medium rivers due to the challenges associated with measurements. Our study presents novel ways to address these challenges and highlights that metabolism patterns in large rivers differ from those observed in small-medium rivers and requires the application of knowledge and tools beyond those implemented for smaller rivers.This data package includes the data and R scripts for the analyses described in Roley et al. 2023. It includes dissolved oxygen and temperature data from a dissolved oxygen HOBO sensor, light data collected from the National Solar Radiation Database (https://nsrdb.nrel.gov/) and hydrologic variables estimated from the MASS-1 model (Niehus et al.; 2014). It also includes metabolism estimates (gross primary production, ecosystem respiration, and net ecosystem production) estimated via streamMetabolizer (Appling et al.; 2018). All analyses in the paper can be replicated with these data and scripts.The data package is comprised of one main data folder. The folder includes (1) file-level metadata (flmd); (2) a data dictionary (dd) for each data file; (3) data files; and (4) R scripts for metabolism estimates and data analysis. All files are .R, .csv, or .pdf.

54 ENVIRONMENTAL SCIENCES↗

Streamflow measurements from four sites on the Tuolumne River in Yosemite National Park from Water Years 2002 to 2021

Regions with remote and complex terrain experience spatially varying streamflow patterns, but are often poorly sampled due to difficult access. This data package includes streamflow measurements collected using low-visibility and low-impact installations at four sites on the Tuolumne River in Yosemite National Park, for water years 2002 to 2021. The resulting data set offers a unique opportunity to explore hydrologic processes in complex terrain.This data package contains half-hourly recordings of unvented pressure, vented pressure, and water temperature are measured and used to estimate discharge and stage height. Discharge flags provide insight into data anomalies. This dataset is formatted in accordance with ESS-Dive's Hydrologic Monitoring and File Level Metadata Formats. It contains the following files:1) Folder containing four csv files of time series streamflow measurements (unvented pressure, vented pressure, estimated discharge, water temperature, stage height, and discharge flag) from four locations on the Tuolumne River2) Data dictionary (dd.csv) containing units, definitions, human readable column names, and data type for all column headers throughout the dataset3) File-level metadata (FLMD.csv) containing metadata for files contained in the dataset4) Installation methods (InstallationMethods.csv) containing metadata on sensor installation

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with “Allometric scaling of hyporheic respiration across basins in the Pacific Northwest USA"

This data package is associated with the publication “Allometric scaling of hyporheic respiration across basins in the Pacific Northwest USA” submitted to JGR-Biogeosciences (Regier et al. 2025).This study used reach-scale modeled estimates of hyporheic aerobic respiration made by the River Corridor Model (Fang et al. 2020) and watershed characteristics across the Willamette and Yakima River basins to explore potential allometric scaling (i.e., power-law relationships between size and function) of cumulative hyporheic respiration across catchment-to-basin scales. Scaling was explored quantitatively via the R2, slope, and y-intercept of relationships between cumulative hyporheic respiration and watershed area, divided into hyporheic exchange flux (HEF) quantiles. We also explored relationships between allometric scaling and other watershed characteristics through linear regression, spatial patterns, and mutual information analyses. Our results also suggest variability of hyporheic respiration allometry for middle exchange flux quantiles, and in relation to land-cover. Our findings provide initial evidence that allometric scaling may be useful for predicting hyporheic biogeochemical dynamics across watersheds from reach to basin scales. This data package is associated with the GitHub repository found at https://github.com/peterregier/rc_wrb_yrb_scaling. The data package is organized into several key directories. The “data” folder contains multiple CSV files, including landscape heterogeneity, scaling analysis, and watershed boundary data. The “figures” folder has all figure files in both PDF and PNG formats. Core analysis scripts and figure generation scripts are in the “scripts” directory, systematically numbered for sequential execution. The root directory includes essential project files; please see the file ending in “flmd.csv” for a list and description of all files contained in this data package and the file ending in “dd.csv” for data dictionaries used to describe tabular column headers.

54 ENVIRONMENTAL SCIENCES↗