Search NASA⌕ Search

SEARCH · Search NASA

Results for “soil databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

49 records · Page 3

ORNLERASE - EPA COUNTS PER MINUTE CALCULATOR CONVERSION FACTORS

This database is used in the EPA's Superfund Counts Per Minute (CPM) Calculator (https://epa-cpm.ornl.gov/index.html). This database contains conversion factors used in the calculator that help users estimate the expected radiation detector reading, in counts per minute (CPM), corresponding to a measured radioactivity level reported in either pCi/cm² or pCi/g. Because surface contamination and volumetric contamination behave differently, the tool uses separate conversion processes, each based on Monte Carlo N-Particle (MCNP)–derived conversion factors. The CPM Calculator supports conversions for a range of common environmental media—including soil, steel, glass, drywall, concrete, and wood—to generate detector-ready CPM values. The primary goal of this tool is to enable more efficient, real-time field measurements, reducing reliance on laboratory analyses and ultimately saving both time and money during Superfund assessments.

Dolislager, Fred [ORNL] (ORCID:0009000325477921)↗

A metagenomic perspective on the microbial prokaryotic genome census

Following 30 years of sequencing, we assessed the phylogenetic diversity (PD) of >1.5 million microbial genomes in public databases, including metagenome-assembled genomes (MAGs) of uncultivated microbes. As compared to the vast diversity uncovered by metagenomic sequences, cultivated taxa account for a modest portion of the overall diversity, 9.73% in bacteria and 6.55% in archaea, while MAGs contribute 48.54% and 57.05%, respectively. Therefore, a substantial fraction of bacterial (41.73%) and archaeal PD (36.39%) still lacks any genomic representation. This unrepresented diversity manifests primarily at lower taxonomic ranks, exemplified by 134,966 species identified in 18,087 metagenomic samples. Our study exposes diversity hotspots in freshwater, marine subsurface, sediment, soil, and other environments, whereas human samples yielded minimal novelty within the context of existing datasets. These results offer a roadmap for future genome recovery efforts, delineating uncaptured taxa in underexplored environments and underscoring the necessity for renewed isolation and sequencing.

59 BASIC BIOLOGICAL SCIENCES↗

MicroFisher: Fungal taxonomic classification for metatranscriptomic and metagenomic data using multiple short hypervariable markers

AbstractProfiling the taxonomic and functional composition of microbes using metagenomic (MG) and metatranscriptomic (MT) sequencing is advancing our understanding of microbial functions. However, the sensitivity and accuracy of microbial classification using genome– or core protein-based approaches, especially the classification of eukaryotic organisms, is limited by the availability of genomes and the resolution of sequence databases. To address this, we propose the MicroFisher, a novel approach that applies multiple hypervariable marker genes to profile fungal communities from MGs and MTs. This approach utilizes the hypervariable regions of ITS and large subunit (LSU) rRNA genes for fungal identification with high sensitivity and resolution. Simultaneously, we propose a computational pipeline (MicroFisher) to optimize and integrate the results from classifications using multiple hypervariable markers. To test the performance of our method, we applied MicroFisher to the synthetic community profiling and found high performance in fungal prediction and abundance estimation. In addition, we also used MGs from forest soil and MTs of root eukaryotic microbes to test our method and the results showed that MicroFisher provided more accurate profiling of environmental microbiomes compared to other classification tools. Overall, MicroFisher serves as a novel pipeline for classification of fungal communities from MGs and MTs.

Wang, Haihua↗

Plant sulfate transporter protein sequences for phylogenetic analysis

Sulfur is an essential macronutrient that supports plant growth, development, and responses to environmental stress. Sulfate is the predominant inorganic form of sulfur in soils, and its uptake by roots and translocation to shoots are facilitated by the sulfate transporter (SULTR) family of proteins. Although the first plant SULTR gene was identified nearly three decades ago, several subfamily members, particularly those in the expansive and angiosperm-specific SULTR3 group, remain poorly characterized. To support comprehensive phylogenetic and sequence-based analyses, we compiled a curated dataset of 262 SULTR protein sequences from 22 plant species spanning the evolutionary breadth of land plants. This collection includes representatives from two basal lineages, two early-divergent angiosperms, six monocots, and ten dicots. All sequences were extracted from genome assemblies available in Phytozome v13 (Joint Genome Institute) and manually curated, with cross-referencing to additional databases such as NCBI when needed. This dataset provides a valuable resource for reconstructing the evolutionary history of the SULTR family, with particular emphasis on the diversification of SULTR3 transporters in flowering plants. This resource may also support functional annotation, comparative genomics, and structural modeling of sulfate transport proteins.

CBI↗

EvoNet: A phylogenomic and systems biology approach to identify genes underlying plant survival in marginal, low‐N soils

The DOE‐BER “EvoNet” project investigates the genetic and molecular basis of plant resilience in extreme environments. We do this by identifying key genes that enable “extreme survivor” species to thrive in the nitrogen-poor soils of Chile’s hyper-arid Atacama Desert. Our collections focus on 32 Atacama extremophile species, including seven grass species with potential biofuel applications. To identify genes-of-importance to survival we compared genomic and transcriptomic profiles of extremophile species that thrive in the Atacama to those of closely related “sister” species from nitrogen-rich arid and mesic regions of California. Deep RNA sequencing and de novo transcriptome assembly across these triplet species sets supported a phylogenomic framework for identifying positively selected genes associated with adaptive divergence. Our integrative analysis combined ecological and environmental data, metagenomics, evolutionary and systems biology, and metabolomics. This enabled us to create an unprecedented framework for systematically understanding how non-model plants have adapted to survive in extreme conditions. Our resulting database of positively selected ortholog groups in the extremophile plants offers promising targets for engineering crop and biofuel species with enhanced resilience to drought and extreme weather. Additionally, our newest dataset explores and exploits a complementary metabolomic approach. This new aspect provides innovative strategies to manipulate plant cell metabolism, further supporting efforts to improve agricultural productivity in the face of extreme climates. Importantly, our combined evolutionary- and metabolomic-based strategies focused on convergent patterns of adaptation, providing a genetic and metabolomic toolkit for improving crop and biofuel resilience across diverse plant species. Finally, our novel exploration of ecological and evolutionary dynamics delivered to the community a phylogenomic computational pipeline called “PhyloGeneious.” Our continued adaptations of this pipeline are publicly available to expedite evolutionary genomic research for future scientific discoveries. In total, our DOE-BER has provided genomic, metabolomic, and computational strategies to understand how extremophile plants provide evolutionary and physiological targets for improving agricultural and biofuel production.

59 BASIC BIOLOGICAL SCIENCES↗

Site Integration and Regulatory Considerations for an NPP Colocated with a Petroleum Refinery, Methanol Plant, and Wood Pulp Plant

This research explores the colocation of nuclear power plants (NPPs) with industrial applications. Three existing industrial sites were considered to demonstrate the siting process and illuminate technological gaps for future work. The three applications demonstrated for colocation here are a petroleum refinery, a methanol production plant, and a pulp and paper plant. This study uses a modified version of the EPRI siting criteria to explore the geological and demographic characteristics of the location of the current industrial site, as well as exploring external hazards from the industrial plant and its surrounding land use. Data was collected from public databases to estimate site characteristics. We then discuss how the site characteristics may impact the ability to colocate an NPP with an industrial application. The application site and 5 additional sites were explored for each application to give a general indication of the siting implications for an NPP in each area. The hazards for each industrial application was also explored to determine how colocation may impact reactor safety. The following gaps have been identified and should be explored in future research on colocation of NPPs with petroleum refineries, methanol plants, and pulp and paper plants: - There is a variety of industrial use, hazards, and pipelines in the surrounding area. A more thorough review of these hazards should be considered for colocation. - In general, the whole region around some applications seems to have softer soil, with implications for large site preparation costs. Further site investigations should prioritize looking into the geotechnical conditions. - Applications along coastlines are susceptible to flooding and hurricanes. The benefits of colocation should be weighed against the potential design implications. - The benefits of natural gas pipeline infrastructure in place should be explored further. If heat supply from the NPP is not required or not feasible due to the distance between the NPP and the application, there may be an opportunity to supply hydrogen to the plant through an existing pipeline. - Because there are several collocated industrial plants in the regions for the refinery and methanol plant, the benefits of sharing resources from the NPP should be explored further. This may open up additional sites for colocation. The following knowledge gaps were identified for the colocation of NPPs with these three industries, and industrial applications in general. These gaps are: - While the STAND tool contains many important characteristics for the reactor siting process, it is not calibrated for the colocation of NPPs with industrial facilities. - There are aspects of both the NPP and industrial application that need to be quantified for a siting analysis. Particularly, we need to understand the water intake requirements for NPPs and each application. - Further work may focus on adapting the STAND site comparison methodology to comparison of sites for co-location. This will involve using the data documented in this report as a starting point and performing a comprehensive and quantitative comparison. - Without spending significant resources, it would be impossible to gather data for each site to evaluate all aspects of siting. One approach to finding data and understanding its implications to siting is looking at FSARs for existing plants. For example, most sites considered in this study have small Vs30 values, indicating soft soil. However, there are NPPs located in the vicinity of most of the sites (e.g., Waterford Steam Electric Station near New Orleans) and reviewing available site characteristics and geotechnical data for these NPPs, might provide further information for siting. - The siting analysis in this study indicates that colocation of the NPP with the industrial site could be difficult based on external hazards, cooling requirements, weather, or population. We need to determine the impact of distance between the two facilities on cost and quality of energy transport. - This study did not touch on socioeconomic impacts for NPP colocation with industrial facilities. The input-output analysis methodology could be applied to the communities referenced in this study to determine the socioeconomic impact of these projects. - Similarly, the impacts of colocation on emergency planning was not explored in this study. The impacts on emergency planning infrastructure are somewhat related to the socioeconomic impacts, and could be explored using a similar methodology. - This study also did not address physical and cybersecurity, which will be important aspects of co-location [ref] . Cybersecurity will be important, regardless of the distance, but physical security will be important if the facilities are located very closely. Physical security might also be important for the steam lines between the plants, unless they are determined to be non-safety significant. - In many site l

08 - HYDROGEN↗

Southwest Regional Partnership on Carbon Sequestration: Phase III (Final Scientific/Technical Report)

The Southwest Regional Partnership on Carbon Sequestration (SWP) is one of 7 regional partnerships formed in 2003 under the U.S. Department of Energy’s (DOE) Regional Carbon Sequestration Partnerships (RCSPs) initiative. The overall purpose of the initiative was to help determine and implement the technology, infrastructure, and regulations most appropriate to promote carbon storage in different regions of the country. Covering Arizona, Colorado, New Mexico, Oklahoma, Utah, and parts of Texas, Wyoming, and Kansas, the SWP evaluated regional carbon storage and utilization potential and focused on technologies and sites that could complement the region’s strong position in energy production. The project progressed through three phases: • Phase I (2003–2005): Characterized regional geologic formations and CO 2 sources, assessed sequestration potential, and identified pilot test sites. • Phase II (2005–2013): Conducted small-scale field tests to validate sequestration methods, including geologic and terrestrial projects. • Phase III (2008–2022): Demonstrated large-scale CO 2 injection at a commercial oil field to test monitoring, verification, and long-term storage strategies. This report covers Phase III. The final project site, the Farnsworth Unit (FWU) in Texas, provided real-world testing of reservoir characterization, monitoring, and risk evaluation tools and processes that could be used in any commercial scale carbon capture, utilization, and storage (CCUS) project. Extensive data collection and analysis helped refine best practices for reservoir characterization, injection monitoring, and storage verification. The SWP contributed to national databases, DOE best practice manuals, and regional geological assessments to support future sequestration efforts. Key lessons learned include the importance of robust data management, strategic site selection, regulatory navigation, and effective industry collaboration. The project’s findings will inform ongoing and future carbon storage initiatives. Task 1 (Regional Characterization) • The SWP continued to participate in national outreach efforts and NATCARB. • The SWP evaluated multiple potential sites before selecting the FWU as the primary field test location. Task 2 (Public Outreach and Education) • The SWP contributed to national databases, DOE best practice manuals, and regional geological assessments to support future sequestration efforts. Task 3 (Permitting and Regulatory Compliance) • The SWP ensured compliance with federal and state regulations, including National Environmental Policy Act (NEPA) requirements. • The SWP obtained all necessary permits for drilling, injection, and monitoring activities. Task 4 (Site Characterization and Planning) • The SWP developed work plans for four key activities: characterization, simulation, monitoring and verification, and risk evaluation. • The SWP collected and synthesized legacy data from multiple sources to build initial static geological models and dynamic reservoir models demonstrating project feasibility. • The SWP conducted an initial risk evaluation and developed mitigation plans. Task 5 (Field Operations and Data Collection) • The SWP drilled, logged, and cored three characterization wells to gather critical subsurface data. • The SWP conducted multiple geophysical surveys, including 3D seismic, crosswell seismic, and vertical seismic profiling, to improve reservoir characterization. Task 6 (Monitoring and Verification) • The SWP performed extensive geological characterization using data from characterization wells and seismic surveys. • The SWP established a surface monitoring network to track CO 2 flux in soil gas, groundwater chemistry, and near-surface atmospheric CO 2 levels. • The SWP built and refined reservoir models to study the effects of relative permeability on simulation behavior and improve calibration with experimental data. Task 7 (Risk Assessment and Model Refinement) • The SWP conducted multiple studies to evaluate reservoir integrity, predict CO 2 plume behavior and improve predictive modeling capabilities. • The SWP refined geological models and used them to enhance the accuracy of simulation models. • The SWP continued quantitative risk assessment of top-ranked risks and strengthened the link between qualitative and quantitative risk methodologies.

02 PETROLEUM↗

Global Geo-processed Data of Aquifer Properties by 0.5° Grid, Country and Water Basins

This repository of global hydrogeologic datasets contains aquifer properties on 0.5° scale, including depth to groundwater (Fan et al., 2013), aquifer thickness (de Graaf et al., 2015), WHYMap aquifer classes (Richts et al., 2011), recharge (Döll and Fiedler, 2008; Gleeson et al., 2016), lakes (Messager et al., 2016), porosity and permeability (Gleeson et al., 2014), digitized and geo-processed from their respective sources. Globally gridded aquifer properties could be used independently to estimate global groundwater availability or used as critical inputs to the superwell model to simulate groundwater extraction and provide estimates of pumped volumes and unit costs under user-specific scenarios. Key resources related to this data are: Niazi, H., Ferencz, S. B., Graham, N. T., Yoon, J., Wild, T. B., Hejazi, M., Watson, D. J., & Vernon, C. R. (2025). Long-term hydro-economic analysis tool for evaluating global groundwater cost and supply: Superwell v1.1. Geoscientific Model Development, 18(5), 1737-1767. https://doi.org/10.5194/gmd-18-1737-2025 superwell model repository which uses this data to simulate groundwater extraction and provides estimates of the global extractable volumes and unit-costs ($/km3) of accessible groundwater production under user-specified extraction scenarios. Repository Overview Main output: aquifer_properties_rec.csv contains all processed outputs, including aquifer properties like porosity, permeability, recharge, lake areas, aquifer thickness, and depth to groundwater. shapefiles.zip: contains all digitized GIS databases and shapefile for all aquifer properties prep_inputs.R and prep_inputs_recharge_lakes.R: R scripts that process the shapefiles to produce the aquifer_properties_rec.csv file plot_inputs.R: R script for plotting the maps and conducting preliminary analysis on the available groundwater volume basin_to_country_mapping.csv, basin_country_region_mapping.csv and continent_county_mapping.csv provide the mapping between continents, 32 energy-economic macro regions, countries, and water basins for post-processing aquifer_properties_rec.csv Maps: Each map visualizes the spatial distribution of one of the aquifer properties across the globe map_in_Porosity.png map_in_Permeability.png map_in_Aquifer_thickness.png map_in_Depth_to_water.png map_in_Recharge.png map_in_Grid_area_km.png map_in_Lake_area_km.png map_in_WHYClass.png Sample inputs sample_inputs.py: this script samples inputs from the aquifer_properties_rec dataset, ensuring the sampled and original inputs maintain the same distributions sampled_data_100.csv contains 100 sampled data points and sampled_data_100.png compares their distributions Dataset Overview The main outputs are consolidated in a comprehensive aquifer_properties_rec.csv file and include the following fields: GridCellID: Unique identifier for each (roughly 0.5°) grid cell Continent: Continent name Country: Country name GCAM_basin_ID: Identifier for GCAM hydrologic basin Basin_long_name: Full name of the basin WHYClass: Hydrogeologic classification based on WHYMap aquifer classes (Richts et al., 2011) Porosity: Soil porosity (%) (Gleeson et al., 2014) Permeability: Soil permeability (in square meters; Gleeson et al., 2014) Aquifer_thickness: Thickness of the aquifer (in meters; de Graaf et al., 2015) Depth_to_water: Depth to groundwater (in meters; Fan et al., 2013) Recharge: long-term annual averaged recharge rates (in m/yr; Döll and Fiedler, 2008; Gleeson et al., 2016) Grid_area: Area of the grid cell (in square meters) Lakes_area: Area of inland lakes (in square meters; Messager et al., 2016) Key References The datasets are digitized versions of global hydrogeologic properties from the following key literature sources: Depth to Groundwater: Fan, Y., Li, H., & Miguez-Macho, G. (2013). Global Patterns of Groundwater Table Depth. Science, 339(6122), 940-943. https://doi.org/10.1126/science.1229881 Aquifer Thickness: de Graaf, I. E. M., Sutanudjaja, E. H., van Beek, L. P. H., & Bierkens, M. F. P. (2015). A high-resolution global-scale groundwater model. Hydrol. Earth Syst. Sci., 19(2), 823-837. https://doi.org/10.5194/hess-19-823-2015 Porosity and Permeability: Gleeson, T., Moosdorf, N., Hartmann, J., & van Beek, L. P. H. (2014). A glimpse beneath earth's surface: GLobal HYdrogeology MaPS (GLHYMPS) of permeability and porosity. Geophysical Research Letters, 41(11), 3891-3898. https://doi.org/10.1002/2014GL059856 Aquifer classes: Richts, A., Struckmeier, W. F., & Zaepke, M. (2011). WHYMAP and the Groundwater Resources Map of the World 1:25,000,000. In J. A. A. Jones (Ed.), Sustaining Groundwater Resources: A Critical Element in the Global Water Crisis (pp. 159-173). Springer Netherlands. https://doi.org/10.1007/978-90-481-3426-7_10 Recharge: Döll, P., & Fiedler, K. (2008). Global-scale modeling of groundwater recharge. Hydrol. Earth Syst. Sci., 12(3), 863-885. https://doi.org/10.5194/hess-12-863-2008; Gleeson, T., Befus, K. M., Jasechko, S., Luijendijk, E., & Cardenas, M. B. (2016). The global volume and distribution of modern groundwater. Nature Geoscience, 9(2), 161-167. https://doi.org/10.1038/ngeo2590 Inland Lakes: Messager, M. L., Lehner, B., Grill, G., Nedeva, I., & Schmitt, O. (2016). Estimating the volume and age of water stored in global lakes using a geo-statistical approach. Nature Communications, 7(1), 13603. https://doi.org/10.1038/ncomms13603 Cite as Niazi, H., Watson, D., Hejazi, M., Yonkofski, C., Ferencz, S., Vernon, C., Graham, N., Wild, T., & Yoon, J. (2024). Global Geo-processed Data of Aquifer Properties by 0.5° Grid, Country and Water Basins. MultiSector Dynamics-Living, Intuitive, Value-adding, Environment. https://doi.org/10.57931/2484226 Contact Reach out to Hassan Niazi or open an issue in the superwell repository for questions or suggestions.

aquifer thickness↗

Estimating irrigation water use from remotely sensed evapotranspiration data: Accuracy and uncertainties at field, water right, and regional scales

Irrigated agriculture is the dominant user of water globally, but most water withdrawals are not monitored or reported. As a result, it is largely unknown when, where, and how much water is used for irrigation. Here, we evaluated the ability of remotely sensed evapotranspiration (ET) data, integrated with other datasets, to calculate irrigation water withdrawals and applications in an intensively irrigated portion of the United States. We compared irrigation calculations based on an ensemble of satellite-driven ET models from OpenET with reported groundwater withdrawals from hundreds of farmer irrigation application records and a statewide flowmeter database at three spatial scales (field, water right group, and management area). At the field scale, we found that ET-based calculations of irrigation agreed best with reported irrigation when the OpenET ensemble mean was aggregated to the growing season timescale (bias = 1.6–4.9%, R 2 = 0.53–0.74), and agreement between calculated and reported irrigation was better for multi-year averages than for individual years. At the water right group scale, linking pumping wells to specific irrigated fields was the primary source of uncertainty. At the management area scale, calculated irrigation exhibited similar temporal patterns as flowmeter data but tended to be positively biased with more interannual variability. Disagreement between calculated and reported irrigation was strongly correlated with annual precipitation, and calculated and reported irrigation agreed more closely after statistically adjusting for annual precipitation. The selection of an ET model was also an important consideration, as variability across ET models was larger than the potential impacts of conservation measures employed in the region. From these results, we suggest key practices for working with ET-based irrigation data that include accurately accounting for changes in soil moisture, deep percolation, and runoff; careful verification of irrigated area and well-field linkages; and conducting application-specific evaluations of uncertainty.

59 BASIC BIOLOGICAL SCIENCES↗

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (May to September 2018)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2018 in May (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) near the Meander C/Pumphouse floodplain sites of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (30-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (70-80 cm depth below surface). Sediments were homogenized from the 10 cm core for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0141020. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 478 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (June to September 2017)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2017 in June (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) in an active meander (Meander C) of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (15-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (50-88 cm depth below surface). Sediments were homogenized from the ~10 cm cores for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0151851. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 405 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled genomes from East River floodplain sediments near Crested Butte, CO, USA (June to September 2019)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken in 2019 in June (flooded conditions) and September (drained conditions) at two locations (MCB1 and MCB3) near the Meander C/Pumphouse floodplain sites of the East River. Sediment cores were collected from 2 depths, a near-surface, generally unsaturated depth (30-40 centimeter (cm) depth below surface) and a deeper depth influenced by flooding with redoximorphic features (70-80 cm depth below surface). Sediments were homogenized from the 10 cm core for microbial analyses. A total of 24 metagenomes were sequenced through the Joint genome institute (JGI) corresponding to 8 samples sequenced in triplicate. These metagenomes can be found under Genomes Online Database (GOLD) sequencing project: Gs0141020. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 436 MAG fasta files and a csv file with quality, taxonomic classification (Genome Taxonomy Database Release RS220), and metagenome accessions for MAGs. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

54 ENVIRONMENTAL SCIENCES↗

The Global Spectra-Trait Initiative: A database of paired leaf spectroscopy and functional traits associated with leaf photosynthetic capacity (v1.0.0)

The Global Spectra-Trait Initiative (GSTI) aims to generate generalizable spectra trait models using reflectance data to predict leaf traits associated with the photosynthesis capacity of leaves. It comprises a synthesized dataset of leaf trait data, input datasets and code. Leaf traits include the maximum carboxylation rate of rubisco (Vcmax), the maximum electron transport rate (Jmax), the dark respiration, as well as the prediction of leaf nitrogen, leaf mass per area (LMA), and leaf water content (LWC). The dataset comprises >7500 paired observations from around 400 species from a broad range of biomes. This dataset comprises a zip file of the GSTI GitHub repository (https://github.com/plantphys/gsti), the synthesized database (.csv) and database metadata files. This dataset was updated on 2025-12-12 with minor edits to mirror the accepted manuscript version and GitHub release (Version 1.0.0 (ESSD accepted version)). Edits included minor changes to the project documentation on GitHub and removal of 12 duplicate entries from the database.

54 ENVIRONMENTAL SCIENCES↗