Search NASA⌕ Search

SEARCH · Search NASA

Results for “omics data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Protein N -Glycans in Healthy and Sclerotic Glomeruli in Diabetic Kidney Disease

Diabetes is expected to directly affect renal glycosylation; yet to date, there has not been a comprehensive evaluation of alterations in N-glycan composition in the glomeruli of patients with diabetic kidney disease (DKD). Here, we used untargeted mass spectrometry imaging to identify N-glycan structures in healthy and sclerotic glomeruli in formalin-fixed paraffin-embedded sections from needle biopsies of five patients with DKD and three healthy kidney samples. Regional proteomics was performed on glomeruli from additional biopsies from the same patients to compare the abundances of enzymes involved in glycosylation. Secondary analysis of single-nucleus RNA sequencing (snRNAseq) data were used to inform on transcript levels of glycosylation machinery in different cell types and states. We detected 120 N-glycans, and among them, we identified 12 of these protein post-translated modifications that were significantly increased in glomeruli. All glomeruli-specific N-glycans contained an N-acetyllactosamine epitope. Five N-glycan structures were highly discriminant between sclerotic and healthy glomeruli. Sclerotic glomeruli had an additional set of glycans lacking fucose linked to their core, and they did not show tetra-antennary structures that were common in healthy glomeruli. Orthogonal omics analyses revealed lower protein abundance and lower gene expression involved in synthesizing fucosylated and branched N-glycans in sclerotic podocytes. In snRNAseq and regional proteomics analyses, we observed that genes and/or proteins involved in sialylation and N-acetyllactosamine synthesis were also downregulated in DKD glomeruli, but this alteration remained undetectable by our spatial N-glycomics assay. Integrative spatial glycomics, proteomics, and transcriptomics revealed protein N-glycosylation characteristic of sclerotic glomeruli in DKD.

60 APPLIED LIFE SCIENCES↗

Metagenome-assembled genomes from Wind River Basin floodplain sediments Riverton, Wyoming site (May to September 2017)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken roughly every month in the period May 18 to September 13 in 2017 at a location (Pit2) close to DOE Legacy Management well 855 at the Riverton, Wyoming floodplain site in the Wind River Basin (WRB). The groundwater at this site exhibits persistent U, Mo, and sulfate plumes and is one of the field sites in focus for the SLAC Groundwater Quality SFA program. Cores were taken with a hand-auger and separated into 5-20 cm segments based on soil horizonation down to 150 cm depth below surface. Each segment was subsampled for microbial analyses. Corresponding 16S rRNA gene amplicon data is available at the NCBI Single Read Archive (SRA) Database BioProject ID PRJNA626616, and soil geochemistry data at doi:10.15485/1631972. 40 metagenomes were sequenced through JGI and can be found under Gold sequencing project: Gs0142591. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 6993 MAG fasta files and a csv file with quality, taxonomic classification (GTDB RS220), and metagenome accessions for MAGs generated from the Wind River Basin (WRB). This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled genomes from topsoils along a hillslope water gradient across early snowmelt to late summer in East River, CO

Drought is changing the American Mountain West at unprecedented rates with unknown consequences to soil microbiome composition and function. As a part of LBNL Watershed Science Focus Area (SFA), we investigated shifts in microbial community and transcriptional activity on a subalpine conifer-meadow transition zone throughout the summer of 2023 as soil dried down. This work took place in Crested Butte, CO on Snodgrass mountain, using a proxy for drought conditions.Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal community at 0-10cm from three sites along a hillslope water gradient across five timepoints from early snowmelt to late summer. 42 metagenomes were sequenced at Joint Genome Institute (JGI) and can be found under the JGI GOLD (Genomes Online Database) sequencing project Gs0166660. Metagenomes were assembled through an inhouse pipeline (see methods), binned using four autobinners (concoct, maxbin2, metabat2, and vamb) and consolidated using dastool. The consolidated bins from all metagenomes were pooled, filtered by completeness (>70%) and contamination (<10%), and dereplicated at 95% ANI using drep. This dataset (1) a zip file of 157 MAGs (as fasta files, Gs0166660_bins_tar.gz), (2) sample metadata file with sample IGSNs (International Generic Sample Numbers) (samples.csv), (3) bounding box coordinates for the sampled locations (Gs0166660.kml), (4) metagenome assembly and coassembly metadata file listing IMG/M (Integrated Microbial Genomes/Metagenomes) metagenome accessions linking samples to metagenomes (EastRiver_Drought_ESSDive_Metadata.csv), (5) location metadata file (locations.csv), (6) file-level metadata file (flmd.csv) and (7) data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled genomes from Slate River floodplain sediments near Crested Butte, CO, USA (June 2018)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken June 2018 at two locations (OBJ1 and OBJ2) near the confluence of the Oh-Be-Joyful Creek and Slate River. The site is one of the field sites in focus for the SLAC National Accelerator Laboratory Groundwater Quality Science Focus Area (SFA) program. Sediment samples from a deep soil pit were collected from 50 to 150 cm depth below surface at discrete depths every 20 cm for microbial analyses. A total of 12 metagenomes were sequenced through the Joint Genome Institute (JGI) and can be found under Genomes Online Database (GOLD) sequencing project: Gs0142591. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 1233 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES↗

Unveiling the Arsenal of Apple Bitter Rot Fungi: Comparative Genomics Identifies Candidate Effectors, CAZymes, and Biosynthetic Gene Clusters in Colletotrichum Species

The bitter rot of apple is caused by Colletotrichum spp. and is a serious pre-harvest disease that can manifest in postharvest losses on harvested fruit. In this study, we obtained genome sequences from four different species, C. chrysophilum, C. noveboracense, C. nupharicola, and C. fioriniae, that infect apple and cause diseases on other fruits, vegetables, and flowers. Our genomic data were obtained from isolates/species that have not yet been sequenced and represent geographic-specific regions. Genome sequencing allowed for the construction of phylogenetic trees, which corroborated the overall concordance observed in prior MLST studies. Bioinformatic pipelines were used to discover CAZyme, effector, and secondary metabolic (SM) gene clusters in all nine Colletotrichum isolates. We found redundancy and a high level of similarity across species regarding CAZyme classes and predicted cytoplastic and apoplastic effectors. SM gene clusters displayed the most diversity in type and the most common cluster was one that encodes genes involved in the production of alternapyrone. Our study provides a solid platform to identify targets for functional studies that underpin pathogenicity, virulence, and/or quiescence that can be targeted for the development of new control strategies. With these new genomics resources, exploration via omics-based technologies using these isolates will help ascertain the biological underpinnings of their widespread success and observed geographic dominance in specific areas throughout the country.

59 BASIC BIOLOGICAL SCIENCES↗

Metagenome-assembled genomes from Slate River floodplain sediments near Crested Butte, CO, USA (September 2019)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken September 2019 at one locations (OBJ1) near the confluence of the Oh-Be-Joyful Creek and Slate River. The site is one of the field sites in focus for the SLAC National Accelerator Laboratory Groundwater Quality Science Focus Area (SFA) program. Sediment samples from a deep soil pit were collected from 50 to 150 cm depth below surface at discrete depths every 20 cm for microbial analyses. A total of 6 metagenomes were sequenced through the Joint Genome Institute (JGI) and can be found under Genomes Online Database (GOLD) sequencing project: Gs0142591. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 2562 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled genomes from Wind River Basin floodplain sediments Riverton, Wyoming site (June to October 2019)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken at three time points from June 12, 2019 to October 23,2019 at a location (PTT1) close to DOE Legacy Management well 855 at the Riverton, Wyoming floodplain site in the Wind River Basin (WRB). The groundwater at this site exhibits persistent U, Mo, and sulfate plumes and is one of the field sites in focus for the SLAC Groundwater Quality SFA program. Sediment samples were collected from 60 to 180 cm below surface every 30cm for microbial analyses through metagenomic sequencing. 15 metagenomes were sequenced through JGI and can be found under Gold sequencing project: Gs0131241. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 780 MAG fasta files and a csv file with quality, taxonomic classification (GTDB RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type. A sample metadata file (samples.csv) that contains site information has also been included.

54 ENVIRONMENTAL SCIENCES↗

Enhancing Biopreparedness through a Model System to Understand the Molecular Mechanisms that Lead to Pathogenesis and Disease Transmission: NW-BRaVE

The science of biopreparedness to counter biological threats hinges on understanding the fundamental principles and molecular mechanisms that lead to pathogenesis and disease transmission. Our vision to address this challenge is to create a powerful and user-friendly platform to elucidate the fundamental principles of how molecular interactions drive pathogen-host relationships and host shifts. We will enable groundbreaking discoveries by integrating a wide range of structural, genomics, proteomics, and other advanced omics measurements, along with evolutionary and artificial intelligence predictions. To make sure the system is applicable to real-world problems, we will develop it in the context of a tractable model system, the small, abundant, and accessible photosynthetic cyanobacteria and their constantly co-adapting viral pathogens, cyanophages. This model will maintain the system’s applicability to real-world problems and techniques, but the overall focus will be on elucidating general principles of detecting, assessing, and surveilling molecular interaction, adaptation, and coevolution that are system agnostic and therefore extensible to other viral-host interactions. Our overall objectives are to (1) identify the molecular complexes that comprise the cyanobacteria redox macromolecular subsystem and how they dynamically change with bacteriophage infection in situ, using cryo-electron tomography; (2) profile regulatory changes during infection using proteomics, multiomics, and experimental validation, and integrate the data with in situ structures; (3) use genomics and metagenomics to determine environmental and population factors across time scales that impact the interactions between marine cyanobacteria and their cyanophage parasites, predicting the evolutionary origins of in situ structural and functional interactions, convergence and coevolution; and (4) develop a data integration and transformation platform that facilitates the integration of in situ, proteomic, and evolutionary measurements of molecular interactions to surveil diverse hosts and parasites in various environmental contexts. These objectives address Focus Area 2 Reveal Molecular Interactions Across Biological Scales for Design of Targeted Interventions. Our powerful and user-friendly platform will enhance connections between the often-siloed fields of structure, molecular phenotype, and evolutionary genomics that are key to biopreparedness, but in need of integration (Figure 1). We will build an integrated navigation tool to facilitate the effective use of globally distributed experimental data for integrated analysis and predictive modeling. The project will develop, implement, and test a platform to assess host-pathogen molecular interactions, adaptation to hosts and host shifts, and coevolution between hosts and pathogens, successfully impacting the research community by revolutionizing abilities to study any host-pathogen interaction, encourage diverse community contributions, and gain fundamental insights into how proteins adapt to new contexts. This ability will be critical for designing early interventions to address future threats. We will build surveillance training capability, aiming for a fair and equitable response to future pandemics and biothreats.

59 BASIC BIOLOGICAL SCIENCES↗

Metagenome-assembled genomes from Slate River floodplain sediments near Crested Butte, CO, USA (June to October 2020)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken June to October 2020 at two locations (OBJ1 and OBJ2) near the confluence of the Oh-Be-Joyful Creek and Slate River. The site is one of the field sites in focus for the SLAC National Accelerator Laboratory Groundwater Quality Science Focus Area (SFA) program. Sediment samples from a deep soil pit were collected from 30 cm depth below surface to just above the cobble layer (~190-250 cm depth) at discrete depths every 40 cm for microbial analyses. A total of 35 metagenomes were sequenced through the Joint Genome Institute (JGI) and can be found under Genomes Online Database (GOLD) sequencing project: Gs0142591. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 2848 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES↗

Coupling Metabolic Source Isotopic Pair Labeling and Genome Wide Association for Metabolite and Gene Annotation in Plants (Final Technical Report)

In this project, we applied our labeling pipeline to Arabidopsis and sorghum by feeding tissues with isotopically labeled versions of commercially available amino acids to identify all metabolite features that incorporate the label. In sorghum, we fed five accessions, sampled across the diversity of sorghum, to identify the precursor-of-origin for metabolites that vary between accessions as well as those that may be missing from a single reference genotype. This provided us with precursor-of-origin annotation for thousands of unknown metabolites. We then used GWA to map genes responsible for the synthesis of precursor-of-origin classified metabolites. For sorghum leaf and root ducible metabolites, we performed untargeted metabolomics on leaf and root tissues from 300 diverse genotyped sorghum inbred lines. The amino acid precursor-of-origin metabolite library were then used to identify the corresponding metabolites in the GWA data sets and to identify novel gene-metabolite associations. Finally, we utilized existing and newly generated sequenced EMS mutants of sorghum to validate the predicted gene-metabolite relationships that our labelling analysis identified. In parallel, we conducted similar feeding experiments in Arabidopsis to categorize metabolites based on precursor-of-origin, identify those that vary across our existing Arabidopsis metabolite GWA dataset, and identify genes required for the synthesis of each metabolite. To provide an independent test of gene annotation and pathway involvement, we tested the GWA gene-metabolite associations in Arabidopsis by analyzing the metabolic phenotypes of gene knockouts. Genes of particular interest from both sorghum and Arabidopsis were studied in detail by directly measuring the activity of the corresponding enzymes following heterologous expression. In summary, this work classified as-yet-unknown amino acid-derived metabolites and identified genes involved in their production generated through “omics” technologies. This information was used to validate gene function and identify new metabolism in Arabidopsis and sorghum.

09 BIOMASS FUELS↗

Single‐Cell Nanodroplet Processing Proteomics Pipeline for Analysis of Human‐Derived Microglia

Single-cell omics tools provide unique insights into heterogeneous cell populations and their responses to stimuli. For example, single-cell RNA sequencing has identified several transcriptionally distinct populations of microglia, which are resident immune cells of the central nervous system (CNS) that are responsive to CNS injury, infection, and neurodegeneration. To date, single-cell studies of microglia have focused on RNA-sequencing or cytometry by time of flight (CyTOF), which provide indirect readouts of protein abundance or quantification of a limited number of targets. Herein, we present a workflow based on FACS-assisted isolation, cryopreservation, and nanodroplet-based processing for single-cell mass spectrometry proteomics analysis of the postmortem human brain cortex-derived microglia. From a single microglial cell, 1039 proteins could be identified on average. As a proof-of-principle, we applied single-cell proteomics for exploring the heterogeneity of brain microglia at the cellular level. This pilot proteomics data partially recapitulates the prior microglia subtypes. Specifically, we determined that mitochondrial proteins, in particular members of NADH dehydrogenase (Complex I), cytochrome b-c1 (Complex III), cytochrome c oxidase (Complex IV), F1-ATPase (Complex V), and Na+/K+-ATPase complex, drive variation across microglia. This pipeline offers the potential for identifying functionally and analytically relevant protein targets for microglia in Alzheimer's disease and other neurological disorders.

59 BASIC BIOLOGICAL SCIENCES↗

Unlocking the distinctive enzymatic functions of the early plant biomass deconstructive genes in a brown rot fungus by cell-free protein expression

ABSTRACT Saprotrophic fungi that cause brown rot of woody biomass evolved a distinctive mechanism that relies on reactive oxygen species (ROS) to kick-start lignocellulosic polymers’ deconstruction. These ROS agents are generated at incipient decay stages through a series of redox relays that shuttle electrons from fungus’s central metabolism to extracellular Fenton chemistry. A list of genes has been suggested encoding the enzyme catalysts of the redox processes involved in ROS’s function. However, navigating the functions of the encoded enzymes has been challenging due to the lack of a rapid method for protein synthesis. Here, we employed cell-free expression system to synthesize four redox or degradative enzymes, which were identified, by transcriptomic data, as conserved players of the ROS oxidation phase across brown rot fungal species. All four enzymes were successfully expressed and showed activities that enable confident assignment of function, namely, benzoquinone reductase (BQR), ferric reductase, α-L-arabinofuranosidase (ABF), and heme-thiolate peroxidase (HTP). Detailed analysis of their catalytic features within the context of brown rot environments allowed us to interpret their roles during ROS-driven wood decomposition. Specifically, we validated the functions of BQR as the driver redox enzyme of Fenton cycles and reconstructed its interactions with the co-occurring HTP or laccase and ABF. Taken together, this research demonstrated that the cell-free expression platform is adequate for synthesizing functional fungal enzymes and provided an alternative route for the rapid characterization of fungal proteins, escalating our understanding of the distinctive biocatalyst system for plant biomass conversion. IMPORTANCE Brown rot fungi are efficient wood decomposers in nature, and their unique degradative systems harbor untapped catalysts pursued by the biorefinery and bioremediation industries. While the use of “omics” platforms has recently uncovered the key “oxidative-hydrolytic” mechanisms that allow these fungi to attack lignocellulose, individual protein characterization is lagging behind due to the lack of a robust method for rapid synthesis of crucial fungal enzymes. This work delves into the studies of biochemical functions of brown rot enzymes using a rapid, cell-free expression platform, which allowed the successful depictions of enzymes’ catalytic features, their interactions with Fenton chemistry, and their roles played during the incipient stage of brown rot when fungus sets off the reactive oxygen species for oxidative degradation. We expect this research could illuminate cell-free protein expression system’s use to fulfill the increasing need for functional studies of fungal enzymes, advancing the discoveries of novel biomass-converting catalysts.

60 APPLIED LIFE SCIENCES↗

Hyaloscypha finlandica Metabolome Repository

This repository provides the curated data tables, manuscript figure and table exports, dependency records, and workflow scripts supporting an integrated comparative genomics and untargeted LC-MS/MS metabolomics analysis of Hyaloscypha finlandica strain PMI 746, a root-associated dark septate endophyte of poplar. The repository includes genome-mining summaries from antiSMASH, FunBGCeX, BGC-Prophet, and BiG-SCAPE; processed metabolomics inputs; metabolite annotation evidence; statistical outputs; and publication-facing figures and tables. Raw LC-MS/MS spectra, full genome/protein downloads, and large generated tool outputs are referenced through public archive/accession records and are not stored in Git.

59 BASIC BIOLOGICAL SCIENCES↗

Associations between SARS-CoV-2 Infection or COVID-19 Vaccination and Human Milk Composition: A Multi-Omics Approach

Background: The risk of contracting severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) via human milk-feeding is virtually nonexistent. Adverse effects of coronavirus disease 2019 (COVID-19) vaccination for lactating individuals are not different from the general population, and no evidence has been found that their infants exhibit adverse effects. Yet, there remains substantial hesitation among this population globally regarding the safety of these vaccines. Objectives: Herein, we aimed to determine if compositional changes in milk occur following SARS-CoV-2 infection or COVID-19 vaccination, including any evidence of vaccine components. Methods: An extensive multiomics approach was taken using a subset of milk samples obtained as part of our broad studies examining the effects on milk of SARS-CoV-2 infection and COVID-19 vaccination. Results: We found that compared with unvaccinated individuals, SARS-CoV-2 infection was associated with significant compositional differences in 67 proteins, 385 lipids, and 13 metabolites. In contrast, COVID-19 vaccination was not associated with any changes in lipids or metabolites, although it was associated with changes in 13 or fewer proteins. Compositional changes in milk differed by vaccine. Changes following vaccination were greatest after 1–6 h for the mRNA-based Moderna vaccine (8 changed proteins), 3 d for the mRNA-based Pfizer (4 changed proteins), and adenovirus-based Johnson and Johnson (13 changed proteins) vaccines. Proteins that changed after both natural infection and Johnson and Johnson vaccine were associated mainly with systemic inflammatory responses. In addition, no vaccine components were detected in any milk sample. Conclusions: Together, our data provide evidence of only minimal changes in milk composition because of COVID-19 vaccination, with much greater changes after natural SARS-CoV-2 infection.

60 APPLIED LIFE SCIENCES↗

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (June to September 2019)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2019 in June (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) near the Meander C/Pumphouse floodplain sites of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (30-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (70-80 cm depth below surface). Sediments were homogenized from the 10 cm core for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0141020. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 436 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (May to September 2018)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2018 in May (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) near the Meander C/Pumphouse floodplain sites of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (30-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (70-80 cm depth below surface). Sediments were homogenized from the 10 cm core for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0141020. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 478 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (June to September 2017)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2017 in June (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) in an active meander (Meander C) of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (15-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (50-88 cm depth below surface). Sediments were homogenized from the ~10 cm cores for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0151851. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 405 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled genomes from topsoils collected during NEON campaign in East River, CO (06/14/2018-06/28/2018)

The Watershed Function Science Focus Area (WF SFA) at Lawrence Berkeley National Lab is working to build a mechanistic understanding of the distribution and dynamics of biogeochemical processes in mountainous watersheds and their response to perturbation. In June 2018, the NEON (National Ecological Observatory Network) Airborne Observatory Platform (AOP) performed a taskable airborne imaging campaign to collect visible to shortwave infrared (VSWIR) imaging spectroscopy and LiDAR data across 330 km2 in the Upper East River at Crested Butte, CO. We conducted a parallel ground sampling campaign to sample vegetation traits, as well as soil physical, chemical, and microbiological characteristics. We collected these samples from 438 sites across 12 locations spanning much of the elevation, topographic, and geologic variability across the study area. A subset of 250 samples were used for soil metagenomics which is presented here. In addition, at each site, vegetation samples were collected to measure species-specific leaf water content and leaf mass area, foliar elemental composition and foliar CN stable isotope ratios. Soil samples were collected to measure soil physical properties which include bulk density and soil texture analysis. A suite of soil chemical properties was measured from the samples collected at each site, including pH, organic matter, concentrations exchangeable cations, total elemental composition, and the concentrations of extractable N pools (e.g. total free amino acids, ammonium, nitrate, dissolved organic N, and total dissolved N). Additionally, we have measured soil microbial biomass CN stoichiometry. Here, we present 1982 metagenome-assembled genomes (MAGs) for the bacterial and archaeal community from topsoil collected from during NEON 2018 campaign. All metagenomes were sequenced at JGI (Joint Genome Institute) (GOLD Study ID: Gs0149986). Metagenomes were assembled using JGI Metagenome Workflow (10.1128/mSystems.00804-20). The dataset includes (1) zip files for 1982 MAG fasta files (neon_genomes1-5.tar.gz, split into 5 tarballs to keep tarballs under 0.5 GB), (2) neon_Gs0149986_samples_soilproperties_metagenomes.csv: the sample information together with the accession numbers for the underlying metagenomes and the associated soil physical and chemical measurements in NMDC (National Microbiome Data Collaborative) compliant format, (3) neon_Gs0149986.kml: location bounding box file for the sampled locations, (4) samples.csv: sample metadata file used to register Internationall Generic Sample Numbers (IGSNs), (5) flmd.csv: file level metadata file, and (6) dd.csv: data dictionary file. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗