Search NASASearch

SEARCH · Search NASA

Results for “Data Quality Office”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

ARM Cloud and Precipitation Measurements and Science Group (CPMSG) 2024 Workshop Report

The mission of the U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) user facility is to improve the understanding and representation of cloud and aerosol processes and their interaction with the Earth's surface in Earth system models (ESMs) by providing comprehensive field observations and supporting advanced data analytics. The ARM Cloud and Precipitation Measurements and Science Group (CPMSG) was chartered in March 2019 to help improve the performance and scientific impact of ARM measurements of clouds and precipitation. The group aims to identify and address gaps in measurement capabilities, maximize the scientific impact of ARM data, and effectively serve the scientific community. To achieve these goals, the group includes experts in cloud and precipitation science, as well as representatives from ARM infrastructure, including instrument mentors, engineers, data quality officers, and data product translators. Prior to CPMSG, early discussions on cloud and precipitation measurements primarily focused on improving radar systems, but have since evolved to include a broader scope involving radiometers and other instruments. Since its formation, the CPMSG has gathered feedback using science traceability matrices. CPMSG aims to keep these as living documents to show the measurement needs, scientific drivers, roadblocks, maturity of measurements and retrievals, and pathways to model improvements. The group meets quarterly to discuss and prioritize measurement and operational improvements.

54 ENVIRONMENTAL SCIENCES

ARM Cloud and Precipitation Measurements and Science Group (CPMSG) 2024 Workshop Report

The mission of the U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) user facility is to improve the understanding and representation of cloud and aerosol processes and their interaction with the Earth's surface in Earth system models (ESMs) by providing comprehensive field observations and supporting advanced data analytics. The ARM Cloud and Precipitation Measurements and Science Group (CPMSG) was chartered in March 2019 to help improve the performance and scientific impact of ARM measurements of clouds and precipitation. The group aims to identify and address gaps in measurement capabilities, maximize the scientific impact of ARM data, and effectively serve the scientific community. To achieve these goals, the group includes experts in cloud and precipitation science, as well as representatives from ARM infrastructure, including instrument mentors, engineers, data quality officers, and data product translators. Prior to CPMSG, early discussions on cloud and precipitation measurements primarily focused on improving radar systems, but have since evolved to include a broader scope involving radiometers and other instruments. Since its formation, the CPMSG has gathered feedback using science traceability matrices. CPMSG aims to keep these as living documents to show the measurement needs, scientific drivers, roadblocks, maturity of measurements and retrievals, and pathways to model improvements. The group meets quarterly to discuss and prioritize measurement and operational improvements.

54 ENVIRONMENTAL SCIENCES

ARM Lead Mentor Selection Process

The Atmospheric Radiation Measurement (ARM) Program was created in 1989 with funding from the U.S. Department of Energy (DOE) to develop several highly instrumented ground stations to study cloud-formation processes and their influence on radiative transfer. This scientific infrastructure provides for fixed sites, mobile facilities, an aerial facility, and a data archive available for use by scientists worldwide through the ARM Climate Research Facility—a scientific user facility. The ARM Climate Research Facility currently operates more than 300 instrument systems that provide ground-based observations of the atmospheric column. To keep ARM at the forefront of climate observations, the ARM infrastructure depends heavily on instrument scientists and engineers, known as Mentors. Mentors must have an excellent understanding of instrumentation theory and operation for their instrument areas and have comprehensive knowledge of critical scale-dependent atmospheric processes. They must also possess the technical and analytical skills to develop new data retrievals that provide innovative approaches for creating research-quality data sets. The ARM Facility seeks the best overall qualified candidate, or team when appropriate, that can fulfill Mentor requirements in a timely manner. The roles and responsibilities of the ARM Instrument Operations Manager are provided in Appendix A. The key role and responsibilities and detailed responsibilities of ARM Lead Mentors are provided in Appendix B and Appendix C, respectively.

47 OTHER INSTRUMENTATION

Plutonium Oxidation State Distribution in the Presence of WIPP-Relevant Organics and Iron Corrosion Products

The oxidation state and solubility of plutonium (Pu) in high ionic strength synthetic WIPP (Waste Isolation Pilot Plant) brines as a function of pC H+ in the presence and absence of WIPP-relevant organic ligands (EDTA [Ethylenediaminetetraacetic acid], oxalate, citrate, acetate) and iron corrosion products (magnetite and metallic iron) at 𝑇 = 23 ± 2 ∘C was thoroughly studied by long-term batch solubility experiments (between approximately 800-1,100 days) from an undersaturation approach. The oxidation state of Pu in the WIPP environment has been a topic of interest since the initial Compliance Certification Application (CCA). This study aims to investigate the solubility of Pu under the expected WIPP conditions and to determine the oxidation state of the solid phase that will control the solubility. One of the most important results of this study is that the Pu oxidation state was analyzed both from the surface area of the corrosion products and in the precipitated solid. The analysis of the Pu oxidation state on the surface of the iron mineral from the ongoing undersaturated experiments is more relevant to the performance assessment of nuclear waste disposal than short term batch (plutonium-iron phase) experiments. The X-ray Absorption Near-Edge Spectroscopy (XANES) analysis showed that Pu oxidation state is different on the metal surface and in the precipitated solid. Pu(III) is the dominant oxidation state in the metallic iron (Fe 0 ) system in the presence and absence of organics. Pu(IV) is the dominant oxidation state in the magnetite system in the presence and absence of organics. Also, organics stabilize Pu(IV) in the magnetite system. Analysis of Pu in the precipitated solid by Extended X-ray Absorption Fine Structure (EXAFS) analysis showed that Pu formed an inner-sphere complex with iron with minor amounts of PuO 2 present. X-ray diffraction (XRD) results indicate that metallic iron and magnetite did not oxidize in three years in the alkaline and high ionic strength system. Under these conditions (8 < pC H+ < 10 at T = (22 ± 2) °C under nitrogen atmosphere, the solubility of Pu changes by up to three orders of magnitude (10 -5 and 10 -8 M). The spread in solubility is highest at pC H+ = 9. Pu(III) and Pu(IV) showed different solubility behavior in the synthetic WIPP brine. A summary of the data collected in this report will be submitted to Sandia National Laboratories as part of a parameter update report which will outline the changes to the OXSTAT parameter for the 2026 Compliance Recertification Application (CRA-2026). The experiments performed were done according to the U.S. Department of Energy (DOE) approved Test Plan entitled “Effects of Radiolysis, Organic Complexation, and Redox Conditions on the Speciation and Oxidation State Distribution of Pu(III/IV)” (LCO-ACP-25). All data reported were obtained under the Los Alamos National Laboratory-Carlsbad Office (LANL-CO) Quality Assurance Program, which is compliant with the DOE Carlsbad Field Office, Quality Assurance Program Document (CBFO/QAPD).

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W

Circularity Futures Workshop Series: Summary Report

The aim of this report is to synthesize key feedback received from the three-part Circularity Futures workshop series held in Spring 2024. The workshop series was conducted by the National Renewable Energy Laboratory (NREL) on behalf of U.S. Department of Energy, Office Energy Efficiency and Renewable Energy (EERE), and was broken into three workshops: Workshop 1 - Circularity Analysis Needs and Priorities; Workshop 2 - Circularity Metrics and Indicators; and Workshop 3 - Circularity Data. Together, the workshops focused on identifying the existing priorities and gaps in the circularity modeling space, understanding different stakeholders' use and interpretation of circularity metrics and indicators, identifying common data gaps and data quality challenges, and assessing the robustness of available solutions. The workshop series brought a diverse group of stakeholders - including representatives from U.S. government offices, national labs, nonprofit organizations, industry, and academia - to collect first-hand feedback on needs, priorities, challenges and opportunities in the circularity modeling and analysis space. The workshop discussions highlighted numerous common needs, priorities and challenges among the interviewed groups. Several topics were frequently discussed, including: 1) Circularity as a pathway for sustainable economic growth: While circularity is generally defined in terms of resource conservation and reducing wasteful disposal of materials, participants agreed that circular strategies should serve broader economic, environmental, and social goals. It is therefore crucial for circularity analysis to look beyond waste reduction and instead evaluate a variety of impact metrics such as cost savings, job creation, air quality, and pollutant emissions. Mutli-criteria decision-making frameworks may be useful for making sense of disparate metrics and evaluating tradeoffs between impact categories.; 2) Economic and social factors are not well understood: Underdevelopment of existing end-of-life (EOL) management infrastructure, inconsistent standardization codes and policy space in reusing recycled content, and suboptimal collection and sorting strategies collectively contribute to uncertainty about the economic potential of circular pathways. The latter observation is consistent among all technologies but more emphasized for renewable energy systems. Social impacts of circularity practices are less understood and less researched than other sustainability aspects.; 3) Inconsistent methods for assessing emerging technologies: LCA and TEA results vary widely depending on the assumptions made with regards to market adoption of new technologies. Emerging technologies suffer limited availability of data needed to conduct a robust circularity analysis. Yet, understanding projected impacts of proposed nascent technology is a key need for different stakeholder groups.; and 4) Lack of temporally and geospatially explicit data: There is a need for open data that represents variations in circularity technologies over time and location. The lack thereof leads to aggregated and potentially misrepresented results in circularity analysis. Sensitivity analyses should be included to verify whether options perceived as more sustainable align with real-world practices.

29 ENERGY PLANNING, POLICY, AND ECONOMY

VA Community Determinants of Health Data Curation Documentation FY26-Q2

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Community Determinants of Health (EDH) Data project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1 km grid) to another (e.g., U.S. Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, U.S. Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1 km grids. Some economic data may only be available at the ZIP code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., U.S. Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Community Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

99 GENERAL AND MISCELLANEOUS

VA Community Determinants of Health Data Curation Documentation FY26-Q1

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Community Determinants of Health (EDH) Data project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Community Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

99 GENERAL AND MISCELLANEOUS

VA Determinants of Health Data Curation Documentation FY25-Q2

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING

VA Determinants of Health Data Curation Documentation FY25-Q3

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, ‘standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING

VA Community Determinants of Health Data Curation Documentation FY25-Q4

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources. With funding from the VA Office of Mental Health and Suicide Prevention (OMHSP), the Determinants of Health (EDH) project has developed innovative datasets associated with specific health outcomes, a methodology for transforming spatiotemporal data from one spatial reference (e.g., a 1km grid) to another (e.g., US Census Tracts), and capabilities for modeling health outcomes. These datasets represent an enhancement of the Agency for Healthcare Research and Quality (AHRQ), addressing key gaps by introducing finer spatial resolution (Census Tract) and additional geographical covariates into existing data. The curation and standardization of these datasets is a complex task since they often originate from various sources and are measured at different spatial and temporal resolutions. For example, US Census data products typically use census blocks, block groups, or counties, while data like weather data are available on 1km grids. Some economic data may only be available at the zip code level. In this context, standardized’ means that all datasets share the same spatial extent (e.g., US Census Tract and/or County), and ‘curated’ implies a repeatable process with data provenance and the use of appropriate methodologies for covariate conversion. The Determinants of Health datasets draw from multiple sources, resulting in variables with varying degrees of availability, patterns of missing data, and methodological considerations across different sources, geographies, and years.

97 MATHEMATICS AND COMPUTING

CROCUS Dual Polarization Ceilometer Data at Northeastern Illinois University Rooftop

This dataset is from the Department of Energy Office of Science funded project, CROCUS Urban Integrated Field Laboratory (https://crocus-urban.org/). The dual-polarization ceilometer (Vaisala CL61) is an autonomous lidar system operating at 910 nm wavelength, providing valuable measurements for understanding atmospheric boundary layer evolution, air quality, and cloud-aerosol interactions. The CL61 measures the backscattered signal intensity alternating between parallel- and cross-polarization signals. The unique depolarization measurement capability improves discrimination between different particle types, such as liquid droplets, ice crystals, and aerosols. The depolarization is highly dependent on the scatterer shape and orientation (spherical vs non-spherical particles), with the linear depolarization ratio providing a measure of dominant backscatter signal component from atmospheric particles at various heights, essentially allowing discrimination between liquid and solid particles. With its efficient optical system, CL61’s improved signal-to-noise ratio compared to traditional ceilometers allows studying detailed vertical profiles of aerosols and clouds up to 15 km height.Datasets are stored in a netCDF data format, and we we encourage users to make use of the associated toolkits available from Unidata (https://www.unidata.ucar.edu/software/netcdf/), Project Pythia (https://foundations.projectpythia.org/core/data-formats/netcdf-cf.html), and our “Instrument Cookbooks” (https://crocus-urban.github.io/instrument-cookbooks) for more information on how to process the metadata-rich datasets.

54 ENVIRONMENTAL SCIENCES

Viziv Wireless Power Transfer Evaluation

Under direction from the DOE Office of Electricity, Sandia National Laboratories performed testing of the Viziv system to evaluate the quality of the Zenneck surface wave and potential application to long range power transfer. This report documents the test methodology as well as the test results. This includes an analysis of prior test data collected by Viziv.

24 POWER TRANSMISSION AND DISTRIBUTION

Does class matter? Understanding differential pandemic recovery via a building typology

This study investigates the recovery of building-level footfall from the COVID-19 pandemic using privacy-preserving mobile devices-based footfall data within 60 downtown areas in the USA and Canada. Using clustering, we identify five distinct building typologies based on their characteristics, including rent, quality and recovery rates. The results reveal significant variation of recovery rates by building features. We find negative relationships with footfall recovery for both the percentage of office and remote work tenants and building quality. In contrast, buildings with traditional work tenants and retail functions achieve higher recovery rates. We also test the ‘flight to quality’ hypothesis via on our typology results. High-quality office buildings (Class A+) continue to have high rents but experience low physical footfall recovery, which suggests that this class is not as resilient as portrayed. The findings thus suggest the importance of considering both economic and footfall resilience in evaluating the performance of office buildings.

Covid-19

Deep reinforcement learning control for co-optimizing energy consumption, thermal comfort, and indoor air quality in an office building

With the recent demand for decarbonization and energy efficiency, advanced HVAC control using Deep Reinforcement Learning (DRL) becomes a promising solution. Due to its flexible structures, DRL has been successful in energy reduction for many HVAC systems. However, only a few researches applied DRL agents to manage the entire central HVAC system and control multiple components in both the water loop and the air loop, owing to its complex system structures. Moreover, those researches have not extended their applications by incorporating the indoor air quality, especially both CO2 and PM2.5concentrations, on top of energy saving and thermal comfort, as achieving those objectives simultaneously can cause multiple control conflicts. What's more, DRL agents are usually trained on the simulation environment before deployment, so another challenge is to develop an accurate but relatively simple simulator. Therefore, we propose a DRL algorithm for a central HVAC system to co-optimize energy consumption, thermal comfort, indoor CO2 level, and indoor PM2.5 level in an office building. To train the controller, we also developed a hybrid simulator that decoupled the complex system into multiple simulation models, which are calibrated separately using laboratory test data. The hybrid simulator combined the dynamics of the HVAC system, the building envelope, as well as moisture, CO2, and particulate matter transfer. Three control algorithms (rule-based, MPC, and DRL) are developed, and their performances are evaluated on the hybrid simulator environment with a realistic scenario (i.e., with stochastic noises). The test results showed that, the DRL controller can save 21.4 % of energy compared to a rule-based controller, and has improved thermal comfort, reduced indoor CO2 concentration. The MPC controller showed an 18.6 % energy saving compared to the DRL controller, mainly due to savings from comfort and indoor air quality boundary violations caused by unmeasured disturbances, and it also highlights computational challenges in real-time control due to non-linear optimization. Finally, we provide the practical considerations for designing and implementing the DRL and MPC controllers based on their respective pros and cons.

Guo, Fangzhou

Metagenome-assembled genomes from Slate River floodplain sediments near Crested Butte, CO, USA (June to October 2020)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken June to October 2020 at two locations (OBJ1 and OBJ2) near the confluence of the Oh-Be-Joyful Creek and Slate River. The site is one of the field sites in focus for the SLAC National Accelerator Laboratory Groundwater Quality Science Focus Area (SFA) program. Sediment samples from a deep soil pit were collected from 30 cm depth below surface to just above the cobble layer (~190-250 cm depth) at discrete depths every 40 cm for microbial analyses. A total of 35 metagenomes were sequenced through the Joint Genome Institute (JGI) and can be found under Genomes Online Database (GOLD) sequencing project: Gs0142591. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 2848 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES

Metagenome-assembled genomes from Slate River floodplain sediments near Crested Butte, CO, USA (September 2019)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken September 2019 at one locations (OBJ1) near the confluence of the Oh-Be-Joyful Creek and Slate River. The site is one of the field sites in focus for the SLAC National Accelerator Laboratory Groundwater Quality Science Focus Area (SFA) program. Sediment samples from a deep soil pit were collected from 50 to 150 cm depth below surface at discrete depths every 20 cm for microbial analyses. A total of 6 metagenomes were sequenced through the Joint Genome Institute (JGI) and can be found under Genomes Online Database (GOLD) sequencing project: Gs0142591. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 2562 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES

Metagenome-assembled genomes from Slate River floodplain sediments near Crested Butte, CO, USA (June 2018)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken June 2018 at two locations (OBJ1 and OBJ2) near the confluence of the Oh-Be-Joyful Creek and Slate River. The site is one of the field sites in focus for the SLAC National Accelerator Laboratory Groundwater Quality Science Focus Area (SFA) program. Sediment samples from a deep soil pit were collected from 50 to 150 cm depth below surface at discrete depths every 20 cm for microbial analyses. A total of 12 metagenomes were sequenced through the Joint Genome Institute (JGI) and can be found under Genomes Online Database (GOLD) sequencing project: Gs0142591. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 1233 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES

Montane Conifer, Aspen, Meadow, and Sagebrush Metagenome Resolved Genomes and Traits in East River Watershed, Colorado, USA

Climate change is driving vegetation shifts in mountain watersheds, with unknown impacts on biogeochemical cycles. We hypothesize that these shifts will reshape soil microbiomes and associated biogeochemical processes. As a part of Lawrence Berkeley National Laboratory (LBNL) Watershed Science Focus Area (SFA), we assessed microbiome and microbial functional trait differences between soils under conifer, aspen, forby meadows, and sagebrush across the East River Watershed, CO, controlling for elevation and aspect.Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from soils 0-20cm in depth across three locations in the watershed—Headwaters, Upper Reaches, and Lower Reaches from August 3-11th 2016. Each location was further subdivided into two blocks, with one block on a west facing aspect, and two on the east aspect of the valley. Within blocks, two samples per vegetation type were taken (one at each depth). This resulted in 66 samples, which were sequenced at JGI and can be found under the Joint Genome Institute (JGI) Genomes Online Database (GOLD) sequencing project Gs0118068. Metagenomes were assembled through an inhouse pipeline (see methods), binned using four autobinners (concoct, maxbin2, metabat2, and vamb) and consolidated using dastool. The consolidated bins from all metagenomes were pooled, filtered by completeness (>75%) and contamination (<25%), and dereplicated at 95% ANI using drep. The dataset includes a zip file of 687 genomes (Vegtype_MAGS.zip), the accession numbers for the underlying metagenomes, a csv file with MAG quality metrics and taxonomy from Genome Taxonomy Database (GTDB) and National Center for Biotechnology Information (NCBI) taxonomic representative genome proteins (EastRiver_Vegtype_drep_genome_info.csv), and a file containing MAG quality metrics and taxonomy (gtdb_drep_bin_taxonomy.csv). The dataset additionally includes a sample metadata file (EastRiver_Vegtype_sample_metadata.csv), a metadata file used to register associated samples with IGSNs (International Generic Sample Numbers) (samples.csv), a Google KML file for the sampled locations (sample_collection_sites.kml), a location metadata file (locations.csv), a file-level metadata file (flmd.csv), and a data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES