Search NASA⌕ Search

SEARCH · Search NASA

Results for “science data processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29

AMMT 2025 Milestone

Idaho National Laboratory initiated the examination of nickel-based alloys manufactured laser powder directed energy deposition additive manufacturing for potential applications in nuclear, high temperature structural components. With the rapid push towards additive manufacturing, codes do not exist that definitively define what is or is not tolerable for each process and application, such as with conventional, wrought products. This report contains the initial work to understand possible manufacturing methods for high temperature alloys, and specifically, void formation, microstructure evolution, and mechanical properties. To generate mechanical test data, specimens were tested irrespective of voids and microstructures were analyzed to better understand how to negate/improve these issues. The preliminary results showed major decreases in mechanical performance for material tested. Test specimens will continue to be produced to further improve additive manufacturing processes, quantify void acceptance, and better understand the most suitable high temperature alloys receptive to additive manufacturing and high temperature nuclear applications.

36 - MATERIALS SCIENCE↗

Towards a Robust Adaptive Digital Twin for Fusion Applications

The development of a digital twin system for fusion applications is essential for enhancing the prediction, analysis, and optimization of complex plasma processes. Machine learning (ML), particularly deep learning has demonstrated strong capabilities in modeling such highly nonlinear and intricate systems. However, two critical challenges limit the deployment of deep learning-based digital twins: Uncertainty Quantification (UQ) and data drift. UQ is vital for ensuring trustworthy predictions, especially in decision-support scenarios. Additionally, data-driven models are often sensitive to changes in the underlying data distribution, such as shot-to-shot variations in fusion experiments, which can lead to performance degradation over time. To address these challenges, we are developing an uncertainty-aware, adaptive digital twin framework. Our approach incorporates deep learning models enhanced with Gaussian Process approximations for predictive uncertainty estimation, coupled with an online learning mechanism that enables continuous model adaptation to new experimental data. This adaptive capability allows the data driven models to respond effectively to evolving plasma behaviors and equipment conditions. Specifically, to mitigate the effects of shot-to-shot drift, our system updates itself incrementally as new data becomes available, improving both robustness and fidelity. Our vision is to evolve this data driven model into a self-sustaining digital twin system that leverages UQ based feedback to continuously refine itself and potentially support real-time decision making. This presentation will cover a brief background on uncertainty quantification for ML, our ongoing effort on development of UQ capabilities for ML, our data science pipeline from data collection to model development and analysis and online learning framework for modeling coil deflection at DIII-D. I will also briefly touch upon opportunities and challenges in development of digital twin framework.

Sammuli, Brian [General Atomics]↗

PFLOTRAN modeling data and scripts associated with “Refining the Hydrogeologic Framework of a Large River Corridor Model Using Waterborne Transient Electromagnetics”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the publication “Refining the Hydrogeologic Framework of a Large River Corridor Model Using Waterborne Transient Electromagnetics” submitted to Water Resources Research (Terry et al. 2025). The data package contains the groundwater modeling dataset from PFLOTRAN software. It includes the python script for mesh generation, boundary condition setting, PFLOTRAN input deck formation and postprocessing. It couples groundwater flow and species transport for Hanford Reach river corridor and pipelines the model generation and processing. This model can be used to easily generate the model and analysis for Hanford site. It can also be adjusted to other hydrologic area with ease. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. The data package consists of 6 folders: (1) “data” contains all necessary data as input and intermediate data for processing; (2) “mesh” contains all mesh related files to generate mesh in Hanford Reach river corridor; (3) “model_run” contains the generated script for PFLOTRAN modeling; (4) “notebooks” contains all the Python script to generate the model; (5) “output” contains all the output from the computation; (6) “postprocessing” contains the Python script to generate scientific figure for manuscript. All files are .csv (comma-separated values), .h5 (HDF5 format), .in (input files), .ipynb (Jupyter notebooks), .p (Python pickle), .png (images), .PNG (images), .py (Python scripts), .pyc (Python bytecode), .r (R scripts), .sh (shell scripts), .txt (text files), .vtu (3D mesh/visualization format), .xz (compressed archive), or .zip (compressed archive).

54 ENVIRONMENTAL SCIENCES↗

CryoTEN: efficiently enhancing cryo-EM density maps using transformers

Abstract Motivation Cryogenic electron microscopy (cryo-EM) is a core experimental technique used to determine the structure of macromolecules such as proteins. However, the effectiveness of cryo-EM is often hindered by the noise and missing density values in cryo-EM density maps caused by experimental conditions such as low contrast and conformational heterogeneity. Although various global and local map-sharpening techniques are widely employed to improve cryo-EM density maps, it is still challenging to efficiently improve their quality for building better protein structures from them. Results In this study, we introduce CryoTEN—a 3D UNETR++ style transformer to improve cryo-EM maps effectively. CryoTEN is trained using a diverse set of 1295 cryo-EM maps as inputs and their corresponding simulated maps generated from known protein structures as targets. An independent test set containing 150 maps is used to evaluate CryoTEN, and the results demonstrate that it can robustly enhance the quality of cryo-EM density maps. In addition, automatic de novo protein structure modeling shows that protein structures built from the density maps processed by CryoTEN have substantially better quality than those built from the original maps. Compared to the existing state-of-the-art deep learning methods for enhancing cryo-EM density maps, CryoTEN ranks second in improving the quality of density maps, while running >10 times faster and requiring much less GPU memory than them. Availability and implementation The source code and data are freely available at https://github.com/jianlin-cheng/cryoten.

Biochemistry & Molecular Biology↗

Compressed sensing methods with applications to advanced air sampling

Environmental sampling methods developed by the Savannah River National Laboratory (SRNL) employ collectors with sorbent media tubes set at various locations to collect airborne emissions. Laboratory analyses of these tubes results in one-dimensional signals regarding what chemicals are being released and transported within the atmosphere. The analysis process is time consuming especially when analyzing a full year’s worth of tubes (hourly sample collection results in nearly 9,000 tubes per year). Using a signal processing method such as compressed sensing allows for recreation of the full signal while greatly reducing the number of analyzed samples required. Due to the sparsity of data retrieved from the air tubes, it is possible to use measurements a fraction of the size of the original data to gain much of the same information. This would improve the overall time and cost of analysis when modeling one-dimensional sampling signals.

54 ENVIRONMENTAL SCIENCES↗

SRBench++: Principled Benchmarking of Symbolic Regression With Domain-Expert Interpretation

Symbolic regression searches for analytic expressions that accurately describe studied phenomena. The main promise of this approach is that it may return an interpretable model that can be insightful to users, while maintaining high accuracy. The current standard for benchmarking these algorithms is SRBench, which evaluates methods on hundreds of datasets that are a mix of real-world and simulated processes spanning multiple domains. At present, the ability of SRBench to evaluate interpretability is limited to measuring the size of expressions on real-world data, and the exactness of model forms on synthetic data. In practice, model size is only one of many factors used by subject experts to determine how interpretable a model truly is. Furthermore, SRBench does not characterize algorithm performance on specific, challenging sub-tasks of regression such as feature selection and evasion of local minima. In this work, we propose and evaluate an approach to benchmarking SR algorithms that addresses these limitations of SRBench by 1) incorporating expert evaluations of interpretability on a domain-specific task, and 2) evaluating algorithms over distinct properties of data science tasks. We evaluate 12 modern symbolic regression algorithms on these benchmarks and present an in-depth analysis of the results, discuss current challenges of symbolic regression algorithms and highlight possible improvements for the benchmark itself.

97 MATHEMATICS AND COMPUTING↗

Baltimore Social-Environmental Collaborative (BSEC) Doppler Lidar & Derived Products

This repository contains all processed Doppler‐lidar outputs from the PSU lidar deployed for the Baltimore Social‐Environmental Collaborative (BSEC) project. Vertical Stare Scans (fixed‐beam, vertical profiling): 1 Hz backscatter intensity (m⁻¹ sr⁻¹), signal‐to‐noise ratio (unitless), and Doppler vertical‐velocity (m s⁻¹) on ~30 m range gates, stored as CF-compliant NetCDF. Wind Profiles (horizontal‐wind retrieval): daily NetCDF outputs of retrieved horizontal wind speed (m s⁻¹) and direction (degrees), computed from the angled‐scan returns. Profile Statistics (summary statistics on the vertical velocity): 15 min windows (default) of mean, variance, skewness, kurtosis, high-frequency variance, etc., as a function of height; saved as CF-compliant NetCDF files. Boundary Layer Height (BLH) (fuzzy-logic output): 15 min BLH estimates (m), with lower/upper fuzzy bounds (m) and a quality flag (0–4) indicating data status (e.g., no data, good, below range, ran out of signal, cloud-topped). Cloud Base Height (Haar-gradient detection): 15 min estimates of cloud-base height (m) with a cloud-detection quality flag (0–3: none, low, moderate, high). All five product streams are organized by year and date under their own top-level folders (01_Vertical_Stare_Scans/ through 05_Cloud_Height/). Each folder contains a data_ /YYYY/ subdirectory with daily CF-compliant NetCDF outputs (96 windows per day at 15 min intervals). Global attributes in each file include creation history, version (2.0.0), institution, and source. Instrument & MeasurementsThe PSU Doppler Lidar samples aerosol backscatter (m⁻¹ sr⁻¹), signal-to-noise ratio, and radial velocity at ~1 Hz. Vertical stare scans point the beam straight up; after collecting angled scans through multiple elevation angles, the "Wind Profiles" product contains the fully retrieved horizontal wind speed and direction. Data were collected continuously at ~30 m range resolution, with a typical height ceiling of ~12 km. How to Use Open any NetCDF with Python's xarray, MATLAB, or similar CF-compliant tools. Stare scans and angled-scan retrievals (Wind Profiles) are CF-compliant daily NetCDF files. Profile-Statistics, BLH, and Cloud Height files are daily 15 min summaries (96 time steps per file). Inspect the included variables (e.g., vertical_velocity_variance, wind_speed, BLH, cloud_base_height) for your analyses. Use the quality flags (BLH_flag, cloud_flag) to filter out poor-quality retrievals. For more information or questions about processing methods, please contact:Nicholas E. Prince ⟨nec5299@psu.edu⟩Penn State Department of Meteorology & Atmospheric Science

Air Quality↗

Taming nuclear mass models with Gaussian processes

We propose a new set of nuclear mass predictions based on multiple theoretical mass models. By employing Gaussian process regression with the Matérn kernel, we achieved root-mean-square (rms) deviations below 100 keV for the training dataset. The best-performing mass models achieved rms deviations below 150 keV for the new precise mass data from AME2020, whereas the ensemble average showed robust performance across the nuclear chart. Our approach uniquely combines: (1) systematic refinement of eight mass models through their residuals, (2) physics-informed features, including magic numbers, nucleon parity numbers, neutron excess, and nuclear collectivity, and (3) theory-to-theory validation demonstrating robust extrapolation capability. We find that the Matérn kernel provides superior uncertainty quantification compared to the RBF kernel, with a length-scale analysis revealing enhanced inter-nuclei correlations. We provide complete mass predictions for all unknown nuclides in AME2020, offering valuable constraints for nuclear structure studies and astrophysical modeling when used with proper uncertainty propagation.

Gaussian processes↗

Using in situ UO 2 bicrystal sintering to understand grain boundary dislocation nucleation kinetics and creep

Capillary evolution at bicrystal UO 2 grain boundaries is characterized using in situ transmission electron microscopy. The discontinuous nature of the densification process, both particle rotation and axial strain, along with the large activation stress for densification support a hypothesis that grain boundary strain in UO 2 follows nucleation rate limited kinetics at low to intermediate stresses, that is, less than ≈ 10 8 PA. Further, the temperature dependence of the average activation stress for sintering agrees well with analysis of bulk sintering data and creep data reported within the literature when analyzed in the context of a grain boundary dislocation nucleation rate limited kinetic model.

36 MATERIALS SCIENCE↗

Leveraging data mining, active learning, and domain adaptation for efficient discovery of advanced oxygen evolution electrocatalysts

Developing advanced catalysts for acidic oxygen evolution reaction (OER) is crucial for sustainable hydrogen production. This study presents a multistage machine learning (ML) approach to streamline the discovery and optimization of complex multimetallic catalysts. Our method integrates data mining, active learning, and domain adaptation throughout the materials discovery process. Unlike traditional trial-and-error methods, this approach systematically narrows the exploration space using domain knowledge with minimized reliance on subjective intuition. Then, the active learning module efficiently refines element composition and synthesis conditions through iterative experimental feedback. The process culminated in the discovery of a promising Ru-Mn-Ca-Pr oxide catalyst. Our workflow also enhances theoretical simulations with domain adaptation strategy, providing deeper mechanistic insights aligned with experimental findings. By leveraging diverse data sources and multiple ML strategies, we demonstrate an efficient pathway for electrocatalyst discovery and optimization. This comprehensive, data-driven approach represents a paradigm shift and potentially benchmark in electrocatalysts research.

Science & Technology - Other Topics↗

Lost and Found: Rediscovering Microbiome-Associated Phenotypes that Reshape Agricultural Sustainability

Modern agriculture faces an urgent need to improve nutrient use efficiency while reducing environmental impacts. Here, we show that ancestral traits controlling rhizosphere microbiome functions can be reintroduced into elite maize through targeted teosinte introgressions. Using near-isogenic lines, we mapped microbiome-associated phenotypes (MAPs) derived from teosinte that suppress nitrification and denitrification—key microbial processes contributing to nitrogen loss. These introgressions altered root exudate chemistry, resulting in distinct microbial assemblies and enhanced nitrogen retention. We identified candidate loci and metabolites responsible for suppressive activity and demonstrated their functional effects in vitro. Our findings reveal a genetic and biochemical basis for rewilding microbiome-mediated ecosystem services in crops, offering a scalable path toward sustainable nutrient management in global agriculture. ---- These maize root exduate metabolomics data are a subset of this larger project and make up a phenotyping for candidate lines.

Favela, Alonso [School of Plant Sciences, Universi↗

Equipment Qualification Report Environmental Qualification of GNB Absolyte Valve Regulated Lead Acid (VRLA) 1600 Ah100G33 Battery Rack Assembly (24590-QL-POA-EDB0-00001-11-00002_00A)

Greenberry Environmental Qualification Report 550001.001-35.0.5 provides basis for assignment of qualified life for: GNB Absolyte Valve Regulated Lead Acid (VRLA) 1600 Ah 100G33 Battery Rack Assembly in accordance with the requirements specified 24590-WTP-3PS-G000-T0015 (Rev 2) and Environmental Qualification Plan 550001.001-35.0.1 (Rev. 3). The equipment qualification basis represents the most conservative capability of the equipment. The analysis performed for the qualification is not less conservative than the bounding environmental conditions detailed in contract documents issued to Greenberry in contract 24509-QL-POA-EDB0-00001 Rev.0. The qualified life of 10 years has been established based upon an end-condition objective of the equipment condition indicators that correlate to the ability of equipment to perform its safety function. The VRLA Battery Cell Assembly was aged by 10 years (minimum) in accordance with conditions specified by Bechtel Equipment Qualification Datasheet 24590-LAW-EUQ-UPE 00003 Rev. 3 and the process conditions specified by the Instrument Data Sheet 24590-LAW EUD-UPE-00009 Rev. 2. Greenberry Industrial has contracted with GNB Industrial Battery Co located at 4115 S Zero St, Fort Smith, AR 72908 to perform age conditioning, monitoring, and capacity testing in accordance with Environmental Qualification Plan 550001.001-35.0.1 Rev. 1. The required process at the stated conditions set by the parameters established by the plan were completed satisfactorily. The details of the of the test process observed by Greenberry is detailed in the attached Seismic Test Log 550001.001-7.0.3, including examples of the objective evidence collected during the qualification process.

54 ENVIRONMENTAL SCIENCES↗

SAIL Radar b1 Data Processing: Corrections, Calibrations, and Processing Report

The U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) user facility deployed the second ARM Mobile Facility (AMF2) near Crested Butte, Colorado for the Surface Atmosphere Integrated Field Laboratory (SAIL) campaign. The SAIL campaign occurred from September 1, 2021, to June 15 2023. To study the water cycle in the East River Watershed, ARM deployed a vertically pointing Ka-band ARM Zenith radar (KAZR) and a scanning X-band precipitation radar managed by Colorado State University (CSU XPRECIP), as shown in Figure 1.

54 ENVIRONMENTAL SCIENCES↗

Lost and Found: Rediscovering Microbiome-Associated Phenotypes that Reshape Agricultural Sustainability

Overview Code and data repository for NIL Manuscript. Documentation includes sequence processing examples and data analysis. Supplemental sequence processing and R statistical analysis for publication, which compares the microbiome of teosinte-B73 Near Isogenic Lines. Sample Data Amplicon sequence data for 16S rRNA genes, the fungal ITS2 region, and nitrogen-cycling functional genes are available through the NCBI Sequence Read Archive (SRA) under accession number PRJNA1042643(https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1042643). Raw metabolomic data are available on Metabolomics Workbench, Project ID: PR002654. This study is available at the NIH Common Fund's National Metabolomics Data Repository (NMDR) website, the Metabolomics Workbench, https://www.metabolomicsworkbench.org where it has been assigned Study ID ST004211. The data can be accessed directly via its Project DOI: http://dx.doi.org/10.21228/M8KV8T.

Near Isogeneic Lines↗

GNSS-based Vegetation Optical Depth, Tree Sway, and Evapotranspiration data from the Niwot Ridge Subalpine Forest (US-NR1) AmeriFlux site

This data package contains data and information about Global Navigation Satellite System (GNSS)-based Vegetation Optical Depth (VOD), tree sway motion, and eddy-covariance evapotranspiration (ET) data collected at the Niwot Ridge Subalpine Forest AmeriFlux site (US-NR1). The raw GNSS data were collected between May 2022 and August 2023. Other processed datasets such as tree sway motion and ET data are also included. The goal was to study the water content within a subalpine forest and, more specifically, examine the canopy evaporation process. This data archive includes all data that were used within the following Biogeosciences discussion paper that further summarizes the research objectives and conclusions:Burns, S.P., V. Humphrey, E.D. Gutmann, M.S. Raleigh, D.R. Bowling, and P.D. Blanken, 2025: Using GNSS-based vegetation optical depth, tree sway motion, and eddy-covariance to examine evaporation of canopy-intercepted rainfall in a subalpine forest. EGUsphere [preprint],https://doi.org/10.5194/egusphere-2025-1755This data archive also supplements the 30-min Lawrence Berkeley National Laboratory (LBNL) AmeriFlux dataset for US-NR1 (i.e., https://doi.org/10.17190/AMF/1246088) and updates what was in the 2020 ESS-DIVE US-NR1 archive (https://doi.org/10.15485/1671825) to include data from the years 2020-2025. More specifically, the following updates are provided: (i) five-minute statistics (means, variances, covariances) of all data measured by the US-NR1 data system between Sep 2020 and Jun 2025 in netCDF format, (ii) the electronic logbook of US-NR1 site visits, (iii) a web calendar (in HTML format) documenting activity at the site (a replica of https://urquell.colorado.edu/calendar/), (iv) photos taken at the site between years 2020 and present day (Aug 2025), and (v) several auxiliary datasets, primary related to trees near the site, soil properties, soil moisture and soil temperature, and subcanopy radiation data. The data package is setup so that the web calendar, photos, and electronic logbook can be easily accessed on a local computer using a web browser. The provided data files are in either BINEX or SBF format (for the raw GNSS data), netCDF, CSV, ASCII, or MATLAB format. To obtain a better understanding about the archive, please start by reading the following PDF which is included within the data archive:README_ESS_DIVE_USNR1_2025_readme_first.pdf.

54 ENVIRONMENTAL SCIENCES↗

The Pan-Arctic Vegetation Cover (PAVC) database v1.1

The Pan-Arctic Vegetation Cover (PAVC) database contains synthesized field-data observations of vegetation cover from 978 Arctic Alaska plots with observations from 2010 to 2021. The cover datasets contain plot data at both the plant functional type (PFT) and species-level resolution, with standardized PFT definitions and species names. We synthesized publicly available point-intercept and visual estimate plots from the Arctic Vegetation Archive of Alaska, the Alaska Vegetation Plots Database, the North Slope Science Catalog, and the National Ecological Observatory Network; as well as previously unpublished data from the Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic).Users will find four synthesized datasets, 4 associated data descriptor (dd) files, and 1 metadata file in the PAVC database:synthesized_species_fcover.csv contains fractional cover (fcover) for unique accepted species names, where names include vegetation identified at the family, genus, species, subspecies, and variety levels, as well as general functional types across all 5 data sources. The synthesized_species_fcover_dd.csv accompanies this dataset with header information.synthesized_pft_fcover.csv contains fcover for the following PFTs: non-vascular plants with lichen and bryophyte subcategories, trees with deciduous and evergreen subcategories, shrubs with deciduous and evergreen subcategories, graminoids (grasses), and forbs (herbaceous flowering plants) measured as total cover. Litter and “other” cover are also included as total cover. Additional “types” include water and bare ground, which were measured as top cover. The synthesized_pft_fcover_dd.csv accompanies this dataset with header information.species_pft_checklist.csv is a lookup table containing the translation from a dataset species name to an accepted species name and to a PFT. This table can be used to clarify our species to PFT adjudications, and to aid users in assigning their own PFTs. Any issues found in this checklist should be reported in the Issues tab of our github.survey_unit_information.csv contains auxiliary information about the plots synthesized in this database. It contains useful information for filtering plots of interest based on temporal, geospatial, and contextual information about the plot surveys.flmd.csv contains metadata information about each file in the database.This research was performed as a part of the NGEE Arctic project. The NGEE Arctic project was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

A robust synthetic data generation framework for machine learning in high-resolution transmission electron microscopy (HRTEM)

Machine learning techniques are attractive options for developing highly-accurate analysis tools for nanomaterials characterization, including high-resolution transmission electron microscopy (HRTEM). However, successfully implementing such machine learning tools can be difficult due to the challenges in procuring sufficiently large, high-quality training datasets from experiments. In this work, we introduce Construction Zone, a Python package for rapid generation of complex nanoscale atomic structures which enables fast, systematic sampling of realistic nanomaterial structures and can be used as a random structure generator for large, diverse synthetic datasets. Using Construction Zone, we develop an end-to-end machine learning workflow for training neural network models to analyze experimental atomic resolution HRTEM images on the task of nanoparticle image segmentation purely with simulated databases. Further, we study the data curation process to understand how various aspects of the curated simulated data—including simulation fidelity, the distribution of atomic structures, and the distribution of imaging conditions—affect model performance across three benchmark experimental HRTEM image datasets. Using our workflow, we are able to achieve state-of-the-art segmentation performance on these experimental benchmarks and, further, we discuss robust strategies for consistently achieving high performance with machine learning in experimental settings using purely synthetic data. Construction Zone and its documentation are available at https://github.com/lerandc/construction_zone.

36 MATERIALS SCIENCE↗

Old Woman Creek Wetland Sediment and Electrochemical Sensor Microbial Community, 2023

We are developing a technique to monitor microbiological activities referred to as zero resistance ammetry, which entails the deployment of graphite electrodes in sediments. Measurement of current between electrodes of contrasting redox regimes and/or predominant terminal electron accepting processes can be used as an indicator of the extents of microbiological activity. We deployed an electrode array at depths of 2 mm, 4 mm, 76 mm, 78 mm, 152 mm, 154 mm, 227 mm, and 229 mm below the wetland sediment water interface in the Old Woman Creek National Estuarine Research Center, Huron, OH, USA (Lat. = 41.380833, Long. = -82.508889). A core was collected from adjacent sediment and subsamples were collected from depth intervals of 0 – 25 mm, 25 – 127 mm, 127 – 128 mm, and below 178 mm. To determine if the microbial communities attached to the electrodes were reflective of the adjacent sediment-associated microbial community, we conducted a 16S rRNA gene-based (V4 region) survey of these respective materials. This data package contains the results of these surveys, including metadata on the depths from which samples were collected (samples.csv), DNA extraction and sequencing information (OWC_DEPTH_AMPLICON_SEQUENCING_METADATA), sequence processing information (OWC_DEPTH_BIOINFORMATIC_METADATA.csv), an operational taxonomic unit (OTU) table (OWC_DEPTH_97OTUS_TABLE.csv), and nucleotide sequences of OTUs (OWC_DEPTH_97OTUS_SEQS.fasta). All files can be opened using a text-editing application. The fasta file is compatible with bioinformatics applications.

54 ENVIRONMENTAL SCIENCES↗