Search NASASearch

SEARCH · Search NASA

Results for “Biological databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Addressing the dynamic nature of reference data: a new nucleotide database for robust metagenomic classification

Accurate metagenomic classification relies on comprehensive, up-to-date, and validated reference databases. While the NCBI BLAST Nucleotide (nt) database, encompassing a vast collection of sequences from all domains of life, represents an invaluable resource, its massive size—currently exceeding 10 12 nucleotides—and exponential growth pose significant challenges for researchers seeking to maintain current nt-based indices for metagenomic classification. Recognizing that no current nt-based indices exist for the widely used Centrifuge classifier, and the last public version currently available was released in 2018, we addressed this critical gap by leveraging advanced high-performance computing resources. We present new Centrifuge-compatible nt databases, meticulously constructed using a novel pipeline incorporating different quality control measures, including reference decontamination and filtering. These measures demonstrably reduce spurious classifications, as shown through our reanalysis of published metagenomic data where Plasmodium annotations were dramatically reduced using our decontaminated database, highlighting how database quality can significantly impact research conclusions. Through temporal comparisons, we also reveal how our approach minimizes inconsistencies in taxonomic assignments stemming from asynchronous updates between public sequence and taxonomy databases. These discrepancies are particularly evident in taxa such as Listeria monocytogenes and Naegleria fowleri, where classification accuracy varied significantly across database versions. These new databases, made available as pre-built Centrifuge indexes, respond to the need for an open, robust, nt-based pipeline for taxonomic classification in metagenomics. Applications such as environmental metagenomics, forensics, and clinical metagenomics, which require comprehensive taxonomic coverage, will benefit from this resource. Our work highlights the importance of treating reference databases as dynamic entities, subject to ongoing quality control and validation akin to software development best practices. This approach is crucial for ensuring accuracy and reliability of metagenomic analysis, especially as databases continue to expand in size and complexity.

59 BASIC BIOLOGICAL SCIENCES

Enabling high-throughput enzyme discovery and engineering with a low-cost, robot-assisted pipeline

Abstract As genomic databases expand and artificial intelligence tools advance, there is a growing demand for efficient characterization of large numbers of proteins. To this end, here we describe a generalizable pipeline for high-throughput protein purification using small-scale expression in E. coli and an affordable liquid-handling robot. This low-cost platform enables the purification of 96 proteins in parallel with minimal waste and is scalable for processing hundreds of proteins weekly per user. We demonstrate the performance of this method with the expression and purification of the leading poly(ethylene terephthalate) hydrolases reported in the literature. Replicate experiments demonstrated reproducibility and enzyme purity and yields (up to 400 µg) sufficient for comprehensive analyses of both thermostability and activity, generating a standardized benchmark dataset for comparing these plastic-degrading enzymes. The cost-effectiveness and ease of implementation of this platform render it broadly applicable to diverse protein characterization challenges in the biological sciences.

36 MATERIALS SCIENCE

RNA language models predict mutations that improve RNA function

Structured RNA lies at the heart of many central biological processes, from gene expression to catalysis. RNA structure prediction is not yet possible due to a lack of high-quality reference data associated with organismal phenotypes that could inform RNA function. We present GARNET (Gtdb Acquired RNa with Environmental Temperatures), a new database for RNA structural and functional analysis anchored to the Genome Taxonomy Database (GTDB). GARNET links RNA sequences to experimental and predicted optimal growth temperatures of GTDB reference organisms. Using GARNET, we develop sequence- and structure-aware RNA generative models, with overlapping triplet tokenization providing optimal encoding for a GPT-like model. Leveraging hyperthermophilic RNAs in GARNET and these RNA generative models, we identify mutations in ribosomal RNA that confer increased thermostability to the Escherichia coli ribosome. The GTDB-derived data and deep learning models presented here provide a foundation for understanding the connections between RNA sequence, structure, and function.

59 BASIC BIOLOGICAL SCIENCES

Building a framework to genetically characterize “feather spots” and understand demographic impacts of solar energy sites on migratory bird populations

The lack of data on the impact of utility-scale solar facilities on avian species and populations adds to the cost of siting and operation. As much as 32 percent of the avian biological material (feathers and carcasses) recovered from solar facilities remain unidentified, because they often take the form of “feather spots”. Feather spots are remains of impacted animals that can be separated into two broad categories: 1) those remains that may be visually identified to a species, or 2) those that cannot be visually identified to a species due to degradation from the environment and/or scavenger activity (listed as “unknown”). Even when feather spots can be identified to species, they cannot be visually assigned to particular breeding populations. In some cases, it is unknown whether multiple feather spots represent single or multiple individuals. This project’s objectives were to: 1. Use a developed, genetic-based technique to identify and determine the species, population of origin, and number of individuals found in feather spots recovered from solar facilities. 2. Implement collected data and resulting analyses to develop a publicly accessible web-based decision-making tool that can be used by the solar industry, regulators and other stakeholders to inform siting, mitigation, and conservation management efforts. 3. Establish a not-for-profit fee-for-service center at UCLA to ensure collection and identification of feather spots continue after the project period of performance. During the Project Period, we proposed to establish a pipeline for collecting, transporting, and storing of avian biological material collected at solar facilities and the collection and identification of feather spots to species and individual. We proposed the development of a genetic-based framework that would recover viable DNA from feather spots, amplify this DNA (i.e., make millions of copies of the original DNA), and use it to match the resulting sequences to a national database of known species of birds. The result would be the identification of feathers spots that were previously unidentified, and the incorporation of these samples into a larger database that included all samples recovered from solar facilities. The resulting report (below) details the result of this work and its alignment with proposed activities. We proposed the use of the data collected to assess the comparative risk to specific species or populations of species from solar facilities. For some species, we have already identified genomic markers of specific breeding populations and developed “genoscapes,” maps of unique genetic variation across the full breeding range of a species. We used these (previously and newly developed) genoscapes to probabilistically link a feather spot to the specific breeding populations from which it originated (assignment probabilities range from 75%-100% depending on species and population groups). For those species without genoscapes, we developed a vulnerability and susceptibility estimate that determines the relative local and regional risk to populations that are in geographic proximity to solar facilities, using citizen science data (Breeding Bird Survey (BBS) and eBird). These two feather spot processing pipelines (see Figure 1 below) provide quantitative estimates as to the numbers of individuals from a given population of origin that are affected by solar facilities, and ultimately can reduce costs to the consumer by reducing the industry costs associated with mitigation and siting strategies for future solar energy development.

14 SOLAR ENERGY

Simulating water dynamics related to pedogenesis across space and time: Implications for four-dimensional digital soil mapping

Digital soil mapping (DSM) relies on machine-learning and geostatistics to represent soil property observations across space. DSM techniques are powerful but often empirical, being limited to the quality and density of point samples. Water dynamics are closely related to soil variability, and the physics that govern water movement are well known. Hydrological properties can hence be simulated by physical models through space and time, unveiling key characteristics about soils. We propose the use of hydrologic models to map soils across the surface (2D), depth (1D), and time (1D)–which provides a 4D approach to digital soil mapping (4DSM). The Distributed Hydrology Soil Vegetation Model (DHSVM) was applied to a watershed currently under pasture. Moisture sensors and wells were installed at different depths in the watershed on summit, sideslope and toeslope positions to validate the model. DHSVM simulations of soil moisture distribution and depth to saturation were performed during the hydrological year (October 2008-September 2009). Clusters of similar pixels based on soil moisture values were determined using Dynamic Time Warping (DTW) to align temporal data and K-means. Clustering was performed both seasonally and for the entire year. Temporal patterns simulated by DHSVM matched measurements given by moisture sensors and wells. Seasonal clusters differed from the annual cluster. Distinct clusters were observed for each season and with depth, showing that spatiotemporal soil variability is lost when statically assessing soils. Spatiotemporal clusters corroborated field observations of fragipan occurrence not explicitly spatially mapped by Soil Survey Geographic Database (SSURGO). If a connection can be made between water and soils, static and dynamic soil variability can be predicted using physically based hydrologic models. Hydrologic models can benefit soil mapping by enabling reliable 4D simulation of water dynamics, which are fundamental to soil variability and soil classification and directly relate to biological, physical and chemical soil processes not captured by typical soil sampling protocols.

54 ENVIRONMENTAL SCIENCES

Expansion of the tmRNA sequence database and new tools for search and visualization

Abstract Transfer–messenger RNA (tmRNA) contributes essential tRNA-like and mRNA-like functions during the process of trans-translation, a mechanism of quality control for the translating bacterial ribosome. Proper tmRNA identification benefits the study of trans-translation and also the study of genomic islands, which frequently use the tmRNA gene as an integration site. Automated tmRNA gene identification tools are available, but manual inspection is still important for eliminating false positives. We have increased our database of precisely mapped tmRNA sequences over 50-fold to 97 179 unique sequences. Group I introns had previously been found integrated within a single subsite within the TψC-loop; they have now been identified at four distinct subsites, suggesting multiple founding events of invasion of tmRNA genes by group I introns, all in the same vicinity. tmRNA genes were found in metagenomic archaeal genomes, perhaps a result of misbinning of bacterial sequences during genome assembly. With the expanded database, we have produced new covariance models for improved tmRNA sequence search and new secondary structure visualization tools.

59 BASIC BIOLOGICAL SCIENCES

Geochemical Phosphorus Sequestration in Tundra Soils Impedes Delivery of Bioavailable Phosphorus to the Kuparuk River, Alaska, USA: Implications for the Broader Arctic Region

Long-term river monitoring of the Kuparuk River (North Slope, Alaska, USA) confirms significant increases in solutes that are indicative of active layer thickening due to thawing permafrost. However, there is no evidence of an increase in total dissolved phosphorus (TDP) or soluble reactive phosphorus (SRP), the nutrient that limits primary production in this and similar rivers in the region. Here, we show that Mehlich-3 extractable iron (Fe) and aluminum (Al) in active layer soils impart high P geochemical sorption capacities across a range of landscape features that we would expect to promote lateral movement of water and solutes to headwater streams in our study watershed. Reanalysis of a recently published pan-Arctic soils database that includes active layer and permafrost soil samples suggests that this high P sorption capacity could be common in other parts of the Arctic region. We conclude that soil minerals enhance P retention on hillslopes and propose pedogenic secondary Fe and Al minerals may continue to retain P in these soils and limit biological productivity in the adjacent river even as active layer thickening increases potential P mobility in the watershed. We suggest that similar interactions may occur in other areas of the Arctic where comparable geochemical conditions prevail.

Sutor, Frederick W. [Univ. of Vermont, Burlington,

Development of Hydropower Biological Evaluation Toolset (HBET): V2.1.9 Release Notes for HBET

The following release notes reflect changes made to HBET for proposed changes to be released in July 2024. Notes are broken up into three sections: 1) Key Improvements, 2) Bug Fixes, and 3) Data Changes • Key Improvements: primary features added and changes to existing features that affect the user experience. • Bug Fixes: Issues discovered or reported that were fixed in the proposed work to be released. • Data Changes: Any work done on the databases directly or the process to calculate data for the system.

13 HYDRO ENERGY

Birth of protein folds and functions in the virome

The rapid evolution of viruses generates proteins that are essential for infectivity and replication but with unknown functions, due to extreme sequence divergence. Here, using a database of 67,715 newly predicted protein structures from 4,463 eukaryotic viral species, we found that 62% of viral proteins are structurally distinct and lack homologues in the AlphaFold database. Among the remaining 38% of viral proteins, many have non-viral structural analogues that revealed surprising similarities between human pathogens and their eukaryotic hosts. Structural comparisons suggested putative functions for up to 25% of unannotated viral proteins, including those with roles in the evasion of innate immunity. In particular, RNA ligase T-like phosphodiesterases were found to resemble phage-encoded proteins that hydrolyse the host immune-activating cyclic dinucleotides 3',3'- and 2',3'-cyclic GMP-AMP (cGAMP). Experimental analysis showed that RNA ligase T homologues encoded by avian poxviruses similarly hydrolyse cGAMP, showing that RNA ligase T-mediated targeting of cGAMP is an evolutionarily conserved mechanism of immune evasion that is present in both bacteriophage and eukaryotic viruses. Together, the viral protein structural database and analyses presented here afford new opportunities to identify mechanisms of virus–host interactions that are common across the virome.

59 BASIC BIOLOGICAL SCIENCES

Isolation of genome-predicted Caldatribacterium ( Atribacterota ) reveals pervasive microbial cultivation problem due to folate precipitation

Most bacterial phyla have few or no pure cultures, including Atribacterota , comprised of ubiquitous anaerobes. Here, we report genome-guided enrichment and isolation of two Atribacterota species representing a new family, Caldatribacterium saccharofermentans from a hot spring, and Caldatribacterium inferamans from a deep aquifer. Both were co-enriched with sulfate-reducing bacteria and initially resisted isolation, which we link to inadvertent removal of precipitated folic acid by filter-sterilization of unbuffered Wolin’s vitamin solution. We then predict folate auxotrophy across the Atribacterota and ~29% of all bacteria, with extensive auxotrophy in 27% of phyla. Since ≥604 of 791 ( ≥ 76%) media with folic acid additions in the MediaDive database use unbuffered vitamin solutions in which folic acid is likely removed during filter-sterilization, we propose that folate auxotrophy limits culturability in defined media en masse. We also uncover unusual features of Caldatribacterium , including three lipid membrane-like layers (LMLs), with the inner LML surrounding the nucleoid, and a high percentage of secreted proteins, supporting a unique cell biology of Atribacterota .

Biological and medical sciences

Embedded EPICS server for PowerPMAC motion controllers

An embedded server layer of Experimental Physics and Industrial Control System (EPICS) for PowerPMAC motion controllers has been developed and deployed at two undulator beamlines of the National Institute of General Medical Sciences and the National Cancer Institute (GM/CA) Structural Biology Facility at the Advanced Photon Source (APS). This compact, open source solution makes the power and versatility of PowerPMAC motion controls directly accessible to distributed EPICS clients. At GM/CA the system controls about 200 servo and stepper motors — both encoded and unencoded — and multiple digital and analog I/O accessories. The server stack comprises two sublayers: a lower-level driver and database that communicates directly with PowerPMAC, and a facility-specific soft sublayer built on top. The paper describes installing EPICS on PowerPMAC, the implementation of both layers and client examples, including on-the-fly scanning.

EPICS

1000 Soils Pilot Dataset, version 8, May 2025

This record hosts data generated by the 1000 Soils Pilot. Data will be updated as more become available. Please see the most recent data upload for current data. A beta visualization tool is available for some data types at https://shinyproxy.emsl.pnnl.gov/app/1000soils. Please submit any suggestions or comments through the 'contact' tab. We are actively working to improve visualizations and value all feedback. Data completed include: Geochemistry, texture, respiration, and enzyme activities FTICR-MS organic matter chemistry Microbial biomass C and N TOC/TDN of water-extractable OM X-ray computed tomography (derived metrics available here, raw data available upon request) Metagenomes; a variety of data formats are available upon request Soil hydraulic properties Data in progress: LC-MS/MS in development, timeline TBD, inquire for status 1000S_processed_BGC_summary.csv contains all available biogeochemical data; microbial biomass C and N; and TOC/TDN of water-extractable OM; and 1000S_Tomography.xslx contains a summary of data generated via X-ray computed tomography. icr_v2_corems2.csv contains FTICR-MS data processed by CoreMS version 2. These data are merged by formula across instrument runs to enable cross-sample comparisons. Technical replicates are merged by retaining peaks present in 2 out of 3 replicates. 1000Soils_Metadata_Site_Mastersheet_v1.csv contains site information. Soil Hydraulics_corrected_02042025.xlsx contains soil hydraulics information. Readme File_v4.xlsx is the readme file. Please contact the MONet project (monet.emsl@pnnl.gov) or Emily Graham (emily.graham@pnnl.gov) with questions. The following file and all raw data are available upon request: icr_by_mass_for_single_sample_analysis_only.csv contains FTICR-MS data processed by CoreMS and is intended for usage in the calculation of biochemical transformations within samples only. These data are not acceptable for cross-sample comparison of masses because they are from multiple instrument runs. For more information, please see: https://www.emsl.pnnl.gov/monet and https://sc-data.emsl.pnnl.gov/monet Acknowledgment: Soil data were provided by the Molecular Observation Network (MONet) at the Environmental Molecular Sciences Laboratory (https://ror.org/04rc0xn13), a DOE Office of Science user facility sponsored by the Biological and Environmental Research program under Contract No. DE-AC05-76RL01830. The work (proposal: 10.46936/10.25585/60008970) conducted by the U.S. Department of Energy, Joint Genome Institute (https://ror.org/04xm1d337), a DOE Office of Science user facility, is supported by the Office of Science of the U.S. Department of Energy operated under Contract No. DE-AC02-05CH11231. The Molecular Observation Network (MONet) database is an open, FAIR, and publicly available compilation of the molecular and microstructural properties of soil. Data in the MONet open science database can be found at https://sc-data.emsl.pnnl.gov/.

biogeochemistry

Viroid-like “obelisk” agents are widespread in the ocean and exceed the abundance of RNA viruses in the prokaryotic fraction

Abstract “Obelisks” are recently discovered ribonucleic acid (RNA) viroid-like elements present in diverse environments with no phylogenetic similarity to any known biological agent. obelisks were first identified in the human gut and in a commensal bacterium acting as a replicative host. They have a circular ∼1 kb RNA genome, rod-like secondary structures, and the encoding of a protein superfamily called “Oblins”. We performed a large-scale search of obelisks in the ocean using the Pebblescout program and the transcriptomic Sequence Archive Read databases, revealing the biogeography and abundance of these viroid-like RNA elements. We detected 55 obelisk genomes resulting in 35 marine clusters at the species level. These obelisks were detected in the prokaryotic fraction and to a lesser extent in the eukaryotic fraction, and distributed across all the oceans from surface to mesopelagic including the Arctic, and even in the coldest seawater of Earth beneath the Antarctic Ross Ice Shelf. The obelisk hallmark protein Oblin-1 confirmed by 3D models was found in various marine samples. Some of the detected marine obelisks harbor hammerhead self-cleaving ribozymes in both polarities. In the prokaryotic, but not the eukaryotic, fraction of the Tara Ocean dataset, relative abundance of obelisks calculated by transcriptomic fragment recruitment indicated that they are abundant in marine samples, reaching or even exceeding the relative abundance of the previously discovered uncultured RNA viruses. In conclusion, obelisks are abundant and widespread viroid-like elements that should be included in ocean biogeochemical models.

Environmental Sciences & Ecology

Genome collection processing for “Conserved upper thermal limits and small safety margins in soil copiotrophic bacteria”

We extracted the genomic DNA of 400 randomly selected isolates using a Quick-DNA Microprep Kit (Zymo Research D3020) according to the manufacturer’s protocol. We then submitted the extracted gDNA samples for short-read Illumina sequencing (200 Mbp) at SeqCoast Genomics (Portsmouth, NH, USA). After preprocessing the sequences using Trimmommatic (Bolger et al. 2014), we assembled the genomes using SPADES (Bankevich et al. 2012) and checked the quality of each assembly using QUAST (Gurevich et al. 2013). We processed the genome assemblies using a KBase (v1.4.0) pipeline (Allen et al. 2017; Arkin et al. 2018). Briefly, we used DRAM (v0.1.2) with default settings to annotate the genome assemblies. We then evaluated genome quality and possible contamination levels using CheckM (v1.0.18) (Parks et al. 2015) and retained genomes with completeness above 98% and contamination below 5% (n = 354), following the authors' guidelines. We then obtained taxonomic assignments for all remaining isolates using the Genome Taxonomy Database tool GTDB-Tk (v2.3.2, database version r214) (Chaumeil et al. 2019). We constructed a phylogenetic tree using the tool SpeciesTree (v2.2.0). We then trimmed the tree (using Trim SpeciesTree to GenomeSet- v1.4.0), retaining only tips within our collection with measured thermal performance.

59 BASIC BIOLOGICAL SCIENCES

Metagenome-assembled genomes from Slate River floodplain sediments near Crested Butte, CO, USA (June to October 2020)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken June to October 2020 at two locations (OBJ1 and OBJ2) near the confluence of the Oh-Be-Joyful Creek and Slate River. The site is one of the field sites in focus for the SLAC National Accelerator Laboratory Groundwater Quality Science Focus Area (SFA) program. Sediment samples from a deep soil pit were collected from 30 cm depth below surface to just above the cobble layer (~190-250 cm depth) at discrete depths every 40 cm for microbial analyses. A total of 35 metagenomes were sequenced through the Joint Genome Institute (JGI) and can be found under Genomes Online Database (GOLD) sequencing project: Gs0142591. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 2848 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES

Metagenome-assembled genomes from Slate River floodplain sediments near Crested Butte, CO, USA (September 2019)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken September 2019 at one locations (OBJ1) near the confluence of the Oh-Be-Joyful Creek and Slate River. The site is one of the field sites in focus for the SLAC National Accelerator Laboratory Groundwater Quality Science Focus Area (SFA) program. Sediment samples from a deep soil pit were collected from 50 to 150 cm depth below surface at discrete depths every 20 cm for microbial analyses. A total of 6 metagenomes were sequenced through the Joint Genome Institute (JGI) and can be found under Genomes Online Database (GOLD) sequencing project: Gs0142591. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 2562 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES

Metagenome-assembled genomes from Slate River floodplain sediments near Crested Butte, CO, USA (June 2018)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken June 2018 at two locations (OBJ1 and OBJ2) near the confluence of the Oh-Be-Joyful Creek and Slate River. The site is one of the field sites in focus for the SLAC National Accelerator Laboratory Groundwater Quality Science Focus Area (SFA) program. Sediment samples from a deep soil pit were collected from 50 to 150 cm depth below surface at discrete depths every 20 cm for microbial analyses. A total of 12 metagenomes were sequenced through the Joint Genome Institute (JGI) and can be found under Genomes Online Database (GOLD) sequencing project: Gs0142591. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 1233 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (June to September 2019)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2019 in June (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) near the Meander C/Pumphouse floodplain sites of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (30-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (70-80 cm depth below surface). Sediments were homogenized from the 10 cm core for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0141020. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 436 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES