Search NASA⌕ Search

SEARCH · Search NASA

Results for “Sequence Function Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Metagenome-assembled genomes from topsoils along a hillslope water gradient across early snowmelt to late summer in East River, CO

Drought is changing the American Mountain West at unprecedented rates with unknown consequences to soil microbiome composition and function. As a part of LBNL Watershed Science Focus Area (SFA), we investigated shifts in microbial community and transcriptional activity on a subalpine conifer-meadow transition zone throughout the summer of 2023 as soil dried down. This work took place in Crested Butte, CO on Snodgrass mountain, using a proxy for drought conditions.Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal community at 0-10cm from three sites along a hillslope water gradient across five timepoints from early snowmelt to late summer. 42 metagenomes were sequenced at Joint Genome Institute (JGI) and can be found under the JGI GOLD (Genomes Online Database) sequencing project Gs0166660. Metagenomes were assembled through an inhouse pipeline (see methods), binned using four autobinners (concoct, maxbin2, metabat2, and vamb) and consolidated using dastool. The consolidated bins from all metagenomes were pooled, filtered by completeness (>70%) and contamination (<10%), and dereplicated at 95% ANI using drep. This dataset (1) a zip file of 157 MAGs (as fasta files, Gs0166660_bins_tar.gz), (2) sample metadata file with sample IGSNs (International Generic Sample Numbers) (samples.csv), (3) bounding box coordinates for the sampled locations (Gs0166660.kml), (4) metagenome assembly and coassembly metadata file listing IMG/M (Integrated Microbial Genomes/Metagenomes) metagenome accessions linking samples to metagenomes (EastRiver_Drought_ESSDive_Metadata.csv), (5) location metadata file (locations.csv), (6) file-level metadata file (flmd.csv) and (7) data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Spatiotemporal forecasting of the edge localized modes in tokamak plasmas using neural networks

Artificial intelligence techniques have been increasingly adopted by the plasma and fusion science to address problems like plasma reconstruction, surrogate modeling, and tokamak/stellarator optimization. A key focus in sustained fusion research is the prediction and mitigation of edge-localized-modes (ELMs), instabilities that occur in short, periodic bursts and can cause erosion to the tokamak vessel wall. Recent research has demonstrated the power of neural networks in approximating continuous functions. In this work, we build spatiotemporal forecasting models that can predict the onset of ELMs and their evolution at early stages. We leverage recent advances in generative modeling, sequence-to-sequence modeling, and Fourier neural operators to propose architectures and training strategies that can learn to forecast short to long term dynamics of the noisy signals due to ELMs. We benchmark the developed model against a state-of-the-art foundation model using the beam emission spectroscopy (BES) data that captures the plasma fluctuations due to ELMs over a 8 x 8 spatial grid. Our models demonstrate high accuracy, outperforming the baselines, in predicting the evolution of BES signals during ELM events. Furthermore, the developed models exhibit high accuracy in predicting the rapid rise and relaxation of the signals due to ELMs within 30–80 µs.

edge localized modes↗

Dataset_for_Molecular_Motion_Below_the_Glass_Transition_A_Solid-State_NMR_Study_of_Siloxane_Polymer_Dynamics Study

This dataset contains solid-state 1H and 13C NMR relaxometry data, differential scanning calorimetry (DSC) data, and size exclusion chromatography (SEC/GPC) data supporting the study of sub-glass-transition (sub-Tg) molecular dynamics in a composition- and sequence-controlled series of diphenyl-substituted polysiloxanes (PDMS, 14Ph, 33Ph, 50Ph, 67Ph, and 100Ph; 0–100% diphenylsiloxane content by mole).All solid-state NMR data were acquired on a 200 MHz Bruker Avance III HD spectrometer using a static 7 mm HX probe or a 4 mm HX probe under 4 kHz magic-angle spinning. Raw Bruker TopSpin experiment folders are included for: (1) variable-temperature 1H lineshape measurements used to determine linewidth (FWHM) as a function of temperature across the glass transition; (2) 1H T1 (saturation recovery with solid-echo detection), probing nanosecond-scale dynamics near the 1H Larmor frequency; (3) 1H T1rho (direct spin-lock, 62.5 kHz), probing microsecond-scale segmental dynamics; (4) 13C-detected Lee–Goldburg cross-polarization 1H T1rho (LGCPH T1rho) for 33Ph and 50Ph, resolving aromatic and aliphatic proton environments; and (5) 13C T1 relaxation for 33Ph and 50Ph. Differential scanning calorimetry data (TA Instruments DSC 25, −150 to +120 °C, up to +300 °C for 100Ph, 10 °C/min) are included for all six compositions and support the glass-transition temperatures in Table 1 and Figure 1. Size exclusion chromatography data (Agilent 1200 Series, PL-Gel 300 mixed-C column, THF mobile phase, polystyrene calibration standards) are included for the three synthesized copolymers (33Ph, 50Ph, 67Ph) and support the number-average molecular weights in Table 1. Processed data include per-composition relaxation-time summaries (Excel), curve-fitting and Bloembergen-Purcell-Pound (BPP) model analysis notebooks (Jupyter/Python), and Igor Pro (.pxp) master files used to generate the manuscript's figures.

Bloembergen-Purcell-Pound theory↗

Structure of the E. coli nucleoid-associated protein YejK reveals a novel DNA binding clamp

Abstract Nucleoid-associated proteins (NAPs) play central roles in bacterial chromosome organization and DNA processes. The Escherichia coli YejK protein is a highly abundant, yet poorly understood NAP. YejK proteins are conserved among Gram-negative bacteria but show no homology to any previously characterized DNA-binding protein. Hence, how YejK binds DNA is unknown. To gain insight into YejK structure and its DNA binding mechanism we performed biochemical and structural analyses on the E. coli YejK protein. Biochemical assays demonstrate that, unlike many NAPs, YejK does not show a preference for AT-rich DNA and binds non-sequence specifically. A crystal structure revealed YejK adopts a novel fold comprised of two domains. Strikingly, each of the domains harbors an extended arm that mediates dimerization, creating an asymmetric clamp with a 30 Å diameter pore. The lining of the pore is electropositive and mutagenesis combined with fluorescence polarization assays support DNA binding within the pore. Finally, our biochemical analyses on truncated YejK proteins suggest a mechanism for YejK clamp loading. Thus, these data reveal YejK contains a newly described DNA-binding motif that functions as a novel clamp.

Biochemistry & Molecular Biology↗

The landscape of regulatory element evolution in a C4 perennial grass

Gene regulatory evolution is a well-known source of phenotypic diversity and adaptive evolution. Although cis-regulatory elements (CREs) play a vital role in gene expression evolution, the molecular evolution of CREs remains mostly unknown due to the difficulty in identifying and characterizing these functional elements. Comparative genomic analyses of noncoding DNA can be leveraged to identify conserved noncoding sequences (CNS), many of which may harbor functional CREs conserved by purifying selection. However, purely computational inference of CREs from putative CNS can be erroneous due to the complex genomic architecture in plants. One promising experimental approach to identify CREs is by profiling accessible chromatin regions (ACRs) that are often associated with the location of CREs. In this study, we use comparative genomics along with the profiling of ACRs to study the molecular evolution of putative functional noncoding regulatory regions in Panicoid grasses. We identified sets of CNS that varied in relationship to the degree of evolutionary divergence among the studied taxa, including identifying core-Panicoid-CNS. We augmented this analysis by profiling ACRs in Panicum hallii ecotypes using ATAC-seq. ACRs had low SNP density at the summit, harbored a high frequency of core-Panicoid-CNS, and were enriched with expression QTL. These data help to annotate the P. hallii genome for putative functional elements and suggest that a large proportion of these ACRs are evolving under purifying selection. Turnover in CNS and ACR between ecotypes of P. hallii identifies a small set of putatively divergent CREs that may underlie differences in gene regulation between genotypes from inland and coastal habitats. In summary, we profiled ACRs in Panicoid grasses and integrated this data with our putative CNS prediction framework, which provides unique insight into patterns of polymorphism and divergence in CREs in C4 perennial grasses.

59 BASIC BIOLOGICAL SCIENCES↗

Methane-cycling microbial communities from Amazon floodplains and upland forests respond differently to simulated climate change scenarios

Seasonal floodplains in the Amazon basin are important sources of methane (CH 4 ), while upland forests are known for their sink capacity. Climate change effects, including shifts in rainfall patterns and rising temperatures, may alter the functionality of soil microbial communities, leading to uncertain changes in CH 4 cycling dynamics. To investigate the microbial feedback under climate change scenarios, we performed a microcosm experiment using soils from two floodplains (i.e., Amazonas and Tapajós rivers) and one upland forest. We employed a two-factorial experimental design comprising flooding (with non-flooded control) and temperature (at 27 °C and 30 °C, representing a 3 °C increase) as variables. We assessed prokaryotic community dynamics over 30 days using 16S rRNA gene sequencing and qPCR. These data were integrated with chemical properties, CH 4 fluxes, and isotopic values and signatures. In the floodplains, temperature changes did not significantly affect the overall microbial composition and CH 4 fluxes. CH 4 emissions and uptake in response to flooding and non-flooding conditions, respectively, were observed in the floodplain soils. By contrast, in the upland forest, the higher temperature caused a sink-to-source shift under flooding conditions and reduced CH 4 sink capability under dry conditions. The upland soil microbial communities also changed in response to increased temperature, with a higher percentage of specialist microbes observed. Floodplains showed higher total and relative abundances of methanogenic and methanotrophic microbes compared to forest soils. Isotopic data from some flooded samples from the Amazonas river floodplain indicated CH 4 oxidation metabolism. This floodplain also showed a high relative abundance of aerobic and anaerobic CH 4 oxidizing Bacteria and Archaea. Taken together, our data indicate that CH 4 cycle dynamics and microbial communities in Amazonian floodplain and upland forest soils may respond differently to climate change effects. We also highlight the potential role of CH 4 oxidation pathways in mitigating CH 4 emissions in Amazonian floodplains.

16S rRNA sequencing↗

Modeling of Stress and Temperature Effects on Creep of Reduced Activation Ferritic-Martensitic Steel Alloy F82H (Tertiary Creep Modeling of RAFM Steel)

A Bayesian optimization procedure is presented for calibrating a multi-mechanism micromechanical model for creep to experimental data of F82H steel. Reduced activation ferritic martensitic (RAFM) steels based on are the most promising candidates for some fusion reactor structures. Although there are indications that RAFM steel could be viable for fusion applications at temperatures up to 600 °C, the maximum operating temperature will be determined by the creep properties of the structural material and the breeder material compatibility with the structural material. Due to the relative paucity of available creep data on F82H steel compared to other alloys such as Grade 91 steel, micromechanical models are sought for simulating creep based on relevant deformation mechanisms. As a point of departure, this work recalibrates a model form that was previously proposed for Grade 91 steel to match creep curves for F82H steel. Due to the large number of parameters (9) and cost of the nonlinear simulations, an automated approach for tuning the parameters is pursued using a recently developed Bayesian optimization for functional output (BOFO) framework [1]. Incorporating extensions such as batch sequencing and weighted experimental load cases into BOFO, a reasonably small error between experimental and simulated creep curves at two load levels is achieved in a reasonable number of iterations. Validation with an additional creep curve provides confidence in the fitted parameters obtained from the automated calibration procedure to describe the creep behavior of F82H steel at 600 °C. The model is further extended using a temperature dependent scaling law approach to simulate creep response between 550 °C and 650 °C. The efficacy of this extension is compared with the previously used scaling law approach for Grade 91 steel.

36 MATERIALS SCIENCE↗

Data from: A high-quality genome assembly of the tetraploid Teucrium chamaedrys unveils a recent whole genome duplication and a large biosynthetic gene cluster for diterpenoid metabolism

Teucrium is well known for making clerodane-type diterpenoids that are produced from the backbone kolavanyl diphosphate. In order to begin to elucidate some of the complex biosynthetic pathways of these medicinal compounds, we identified and functionally characterized several kolavanyl diphosphate synthases from T. chamaedrys . Along the way, we discovered the genome of this species to be one of the largest genomes published from the Lamiaceae family, to which it belongs. This tetraploid, 3 Gbp genome is especially rich in diterpene synthase genes, with 74 putative sequences identified.

biosynthetic gene cluster (BGC)↗

Physics-Informed Machine Learning Model for Ceramic Matrix Composite Creep

A physics-informed recurrent neural network (RNN) based surrogate model is developed to emulate the nonlinear, time-dependent constitutive behavior of ceramic matrix composites (CMCs) driven by matrix damage and constituent creep at the microscale. Physics-informed constraints are introduced into the surrogate model through regularization to ground the prediction in physics and improve its predictive capabilities. Training data is generated using the high-fidelity generalized method of cells (HFGMC) approach which calls appropriate creep and damage models for each of the constituents. This coupling permits simulating the nonlinear behavior of CMCs based on constituent response at the microscale along with microstructural features such as fiber and porosity volume fraction and fiber radius. The microscale repeating unit cell is loaded under creep fatigue conditions to replicate the material loading experienced in a turbine engine. Therefore, the RNN-based surrogate model is tasked with predicting, as a function of variable input stress sequence, temperature, and microstructural features, the resulting strain history response while satisfying physical constraints related to creep rate, isochoric inelastic deformation, and strain energy density. The trained surrogate model is shown to effectively match the strain history over quantified distributions of microstructural features and relevant loading regimes and temperatures. Neural network based surrogate models can offer efficient alternatives to running computationally intensive multiscale material models to simulate the nonlinear response of large structural models. Therefore, the presented work provides evidence towards the feasibility of developing, training, and running such models for CMCs with complex microstructures, nonlinear time-dependent material response, and under non-monotonic loading conditions.

ceramic matrix composites↗

Produced Water DNA Database (PW-DNA): Utilizing KBase to generate an environmental specific curated molecular database

The deep subsurface is estimated to host the majority of Earth’s microbial biomass yet remains one of the most challenging environments to access and study. One common approach to investigate these microbial communities is through the analysis of produced water from subsurface reservoirs, where researchers can assess water and gas chemistry along with molecular (DNA/RNA) sequence data. Advances in high-throughput sequencing have greatly expanded our understanding of these environments and their biotechnological potential. However, further progress requires large-scale, integrative meta-analyses across diverse datasets. To address this need, we developed the Produced Water-DNA (PW-DNA) Database, a curated, publicly available resource that consolidates microbial DNA/RNA sequences, geochemical data, and relevant metadata from in situ hydrocarbon environments such as coal beds, oil reservoirs, and natural gas systems. The PW-DNA database delivers three core benefits to the research community: (1) it improves data sharing by linking environmental microbial datasets with corresponding geochemical parameters, enabling more robust filtering and analysis; (2) it connects with complementary research databases to promote broader dissemination and interoperability; and (3) it supports technological innovation by serving as a resource for identifying microbial trends and exploring genetic potential. While individual studies have highlighted basin-specific microbial communities and functional redundancy in biogeochemical cycling, a comprehensive, system-wide perspective is needed to better understand connectivity and novelty across subsurface ecosystems. By designing the PW-DNA in the KBase platform, we provide a reproducible, visual framework for integrating large-scale genomic and geochemical data, enabling researchers to perform more informed analyses and experimental design. Ultimately, this resource enhances the ability to identify, characterize, and interpret microbial functions across diverse subsurface environments, thereby accelerating discovery in subsurface microbiology and biotechnology.

59 BASIC BIOLOGICAL SCIENCES↗

Missing microbial eukaryotes and misleading meta-omic conclusions

Meta-omics is commonly used for large-scale analyses of microbial eukaryotes, including species or taxonomic group distribution mapping, gene catalog construction, and inference on the functional roles and activities of microbial eukaryotes in situ. Here, we explore the potential pitfalls of common approaches to taxonomic annotation of protistan meta-omic datasets. We re-analyze three environmental datasets at three levels of taxonomic hierarchy in order to illustrate the crucial importance of database completeness and curation in enabling accurate environmental interpretation. We show that taxonomic membership of sequence clusters estimates community composition more accurately than returning exact sequence labels, and overlap between clusters can address database shortcomings. Clustering approaches can be applied to diverse environments while continuing to exploit the wealth of annotation data collated in databases, and selecting and evaluating these databases is a critical part of correctly annotating protistan taxonomy in environmental datasets. We argue that ongoing curation of genetic resources is crucial in accurately annotating protists in in situ meta-omic datasets. Moreover, we propose that precise taxonomic annotation of meta-omic data is a clustering problem rather than a feasible alignment problem.

59 BASIC BIOLOGICAL SCIENCES↗

Canted antiferromagnetism and spin reorientation in corner-shared single chain quasi-one-dimensional Ba 2 ⁢FeSe 3

Here, we report the canted antiferromagnetic (AFM) structure together with a spin reorientation in a single chain quasi-one-dimensional (Q-1D) iron chalcogenide Ba 2⁢ FeSe 3 . Ba 2 ⁢FeSe 3 crystallizes in Pnma (No. 62) orthorhombic structure with linear single iron chains consisting of corner-shared distorted FeSe 4 tetrahedra along the 𝑏 axis. Ba 2 ⁢FeSe 3 is a narrow-gap semiconductor and orders AFM below 60 K. Modeling of neutron powder diffraction data reveals a canted AFM ground state of magnetic space group 𝑃⁢𝑎⁢21/𝑐 (BNS No. 14.80) with commensurate propagation vector 𝐤 =(0, $\frac{1}{2}$, 0), where the Fe ion spins are AFM aligned with up-down-up-down (↑−↓−↑−↓) sequence along the Q-1D chain direction of the 𝑏 axis. In the magnetically ordered state, the canting of magnetic moments reorients from the 𝑎⁢𝑐 plane to the 𝑎⁢𝑏 plane below 30 K, with a 10° tilting angle toward the 𝑎 axis, and the magnetic moment does not induce a net moment in either orientation. The density functional theory results indicate that an ↑−↓−↑−↓ AFM state is stabilized along the chain direction. In this work, we elucidate the unique canted AFM of the iron chalcogenide and pave the way for searching exotic physics in Q-1D Ba 2⁢ FeSe 3 .

Gao, Fei [Univ. of Texas at Dallas, Richardson, TX↗

Montane Conifer, Aspen, Meadow, and Sagebrush Metagenome Resolved Genomes and Traits in East River Watershed, Colorado, USA

Climate change is driving vegetation shifts in mountain watersheds, with unknown impacts on biogeochemical cycles. We hypothesize that these shifts will reshape soil microbiomes and associated biogeochemical processes. As a part of Lawrence Berkeley National Laboratory (LBNL) Watershed Science Focus Area (SFA), we assessed microbiome and microbial functional trait differences between soils under conifer, aspen, forby meadows, and sagebrush across the East River Watershed, CO, controlling for elevation and aspect.Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from soils 0-20cm in depth across three locations in the watershed—Headwaters, Upper Reaches, and Lower Reaches from August 3-11th 2016. Each location was further subdivided into two blocks, with one block on a west facing aspect, and two on the east aspect of the valley. Within blocks, two samples per vegetation type were taken (one at each depth). This resulted in 66 samples, which were sequenced at JGI and can be found under the Joint Genome Institute (JGI) Genomes Online Database (GOLD) sequencing project Gs0118068. Metagenomes were assembled through an inhouse pipeline (see methods), binned using four autobinners (concoct, maxbin2, metabat2, and vamb) and consolidated using dastool. The consolidated bins from all metagenomes were pooled, filtered by completeness (>75%) and contamination (<25%), and dereplicated at 95% ANI using drep. The dataset includes a zip file of 687 genomes (Vegtype_MAGS.zip), the accession numbers for the underlying metagenomes, a csv file with MAG quality metrics and taxonomy from Genome Taxonomy Database (GTDB) and National Center for Biotechnology Information (NCBI) taxonomic representative genome proteins (EastRiver_Vegtype_drep_genome_info.csv), and a file containing MAG quality metrics and taxonomy (gtdb_drep_bin_taxonomy.csv). The dataset additionally includes a sample metadata file (EastRiver_Vegtype_sample_metadata.csv), a metadata file used to register associated samples with IGSNs (International Generic Sample Numbers) (samples.csv), a Google KML file for the sampled locations (sample_collection_sites.kml), a location metadata file (locations.csv), a file-level metadata file (flmd.csv), and a data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Assembly of small silica nanoparticles using lipid-tethered DNA ‘bonds’

Single-stranded DNA molecules modified with cholesterol functional groups are physically tethered to silica nanoparticles (diameter 25 nm) that are encapsulated in a lipid bilayer. Such tethering increases the azimuthal mobility of the DNA molecules across the nanoparticle surface and enables nonspecific bonding, eliminating the need for specialized surface chemistries (such as silane or thiol ligands). To induce assembly, double-stranded DNA ‘bridge’ molecules are then added with complementary nucleotides to the DNA ‘anchor’ molecules that are physically tethered to the lipids on the surface of the particles. Assembly is observed to occur at room temperature and without the need for temperature annealing. Using automated liquid handling tools, assemblies are created in high throughput and rapidly characterized using SAXS. It is determined that the relative concentration of DNA-to-silica and the ionic strength of the solution are important parameters that affect the resulting assembly. Analysis of SAXS data is performed using coarse-grained particle dynamics simulations. The results support the spontaneous formation of semi-crystalline particle assemblies by particle condensation, where the interparticle distance is tuned by the sequence of the DNA ‘bridge’ used to link the particles. Crystallinity analysis performed on the resulting simulations, optimized to match SAXS observations, suggest that particle clusters display increased crystallinity in the center of the clusters, but their maximum size remains relatively small (sub-micron) before settling occurs, which limits the extent of crystallization.

Chiang, Huat Thart [Univ. of Washington, Seattle, ↗

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (June to September 2019)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2019 in June (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) near the Meander C/Pumphouse floodplain sites of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (30-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (70-80 cm depth below surface). Sediments were homogenized from the 10 cm core for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0141020. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 436 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (June to September 2017)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2017 in June (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) in an active meander (Meander C) of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (15-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (50-88 cm depth below surface). Sediments were homogenized from the ~10 cm cores for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0151851. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 405 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (May to September 2018)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2018 in May (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) near the Meander C/Pumphouse floodplain sites of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (30-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (70-80 cm depth below surface). Sediments were homogenized from the 10 cm core for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0141020. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 478 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled genomes measured at 3 depths during snowmelt period in East River, CO (March, May, and June, September 2017)

Snowmelt is a critical biogeochemical period that accounts for large nitrogen (N) export events from high-elevation watersheds. Soil microbial populations bloom and immobilize N during snowmelt, yet the population size crashes in spring, which releases a pulse of soil N. We sought to discover the N sources fueling this microbial bloom and determine the fate of N following microbial die-off. Here, focusing on the snowmelt period within a headwater catchment of the Upper Colorado River Basin (East River, CO), we deployed strain-resolved metagenomics to identify the metabolic pathways and processes that mobilize soil N during and after snowmelt. Soil metagenome samples were taken from 6 snowpits from 3 depths (0-5cm, 5-15cm, >15cm) at 4 time points during snowmelt period (March 2017, May 2017, and June 2017, September 2017) generating 48 metagenomes. We reconstructed 474 metagenome-assembled genomes (MAGs) across all metagenomes.All 48 metagenomes were sequenced at JGI and raw data can be found under JGI (Joint Genome Institute) GOLD Study Gs0135149. Metagenome assemblies from IMG under the same study were used for genome binning. This dataset (1) a zip file of 474 MAGs (as fasta files, Gs0135149_bins_tar.gz), (2) sample metadata file with sample IGSNs (International Generic Sample Numbers) (samples.csv), (3) bounding box coordinates for the sampled locations (Gs0135149.kml), (4) metagenome metadata file listing IMG/M (Integrated Microbial Genomes/Metagenomes) metagenome accessions linking samples to metagenomes (metagenomes.csv), (5) location metadata file (locations.csv), (6) file-level metadata file (flmd.csv) and (7) data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗