Search NASA⌕ Search

SEARCH · Search NASA

Results for “Science Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

End-of-Winter Snow Depth, Temperature, Density, and SWE Measurements at Teller Road Site, Seward Peninsula, Alaska, 2023

Measurements of end-of-winter snow properties were collected at the NGEE Arctic Teller Road Site at mile marker 27 (TL_MM27) from March 31 to April 3, 2023. This dataset contains one *.pdf user guide and four *.csv data files of spatially distributed values of snow depth, snow water equivalent (SWE), snow temperature, and snow density. Snow temperature and snow density were collected throughout the snowpack at different depths. Data were collected toward the end of the winter season from late March to early April, when the snowpack is typically near its maximum. A Snow-Hydro MagnaProbe (http://www.snowhydro.com/products/column2.html) was used to improve collection efficiency and enhance spatial coverage. This dataset is a continuation of the previous end-of-winter snow surveys conducted at the Teller Road Site in 2016-2018 (Wilson et al., 2020; https://doi.org/10.5440/1592103), 2019 (Bennett et al., 2021; https://doi.org/10.5440/1798170), and 2022 (Bennett et al., 2022; https://doi.org/10.5440/1887250).The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic) was a 15-year research effort (2012-2027) to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy’s Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy’s Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Temporal Study 2022-2024: Sample-Based Surface Water Dissolved Inorganic Carbon, Dissolved Organic Carbon, Total Nitrogen, Stable Isotopes, and Total Suspended Solids from across Multiple Watersheds in the Yakima River Basin, Washington, USA

This dataset supports a broader study examining the drivers of temporal variability in sediment respiration rates in the Yakima River Basin. The dataset provides geochemistry data generated from samples collected at bi-weekly or monthly intervals at six sites across the Yakima River Basin in Washington, USA. Sample and sensor data from previous years (2021-2022) can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1898912 and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1892054, respectively. Related sensor data from 2022-2024 will be published separately. This dataset is comprised of one main data folder containing (1) file-level metadata; (2) data dictionary; (3) readme; (4) field metadata; (5) dissolved inorganic carbon (DIC) and averages; (6) dissolved organic carbon (DOC; reported as non-purgeable organic carbon; NPOC) and averages; (7) total dissolved nitrogen (TN) and averages; (8) total suspended solids (TSS); (9) stable isotopes; (10) surface water sampling protocol; (11) sensor protocol; (12) methods codes; and (13) international generic sample number (IGSN) mapping file. All files are .csv or .pdf. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. For data and scripts associated with "Shifts in rain-snow partitioning drive faster water transit times in the US Pacific Northwest" (Butler et al., 2026), go to https://data.ess-dive.lbl.gov/datasets/doi:10.15485/3025481

18-O↗

Terrestrial laser scanning data (Levels 0 and 1) for Pasoh, Malaysia, Sep 2024

This data package contains data from terrestrial laser scanning (TLS) at the Pasoh Forest Reserve, Malaysia. The Pasoh Forest Reserve is a facility of the Forest Research Institute Malaysia, and contains evergreen lowland dipterocarp forest. The Next-Generation Ecosystem Experiments Tropics (NGEE-Tropics) study areas at Pasoh were established to study how different species respond to climatic variation and soil water availability. Two study areas were chosen representing different topography and species. The TLS data archived here were collected to provide detailed, three-dimensional information about forest structure. Specifically, data were collected to allow tree-level characterization of woody structure and leaf area for 12 focal trees with FloraPulse and sap flux sensors, facilitating estimation of woody biomass and leaf area to allow upscaling of water content and transpiration data to the tree-level. Scan positions were not selected to provide consistent data for non-focal trees with the study areas. This data package contains the following data: - High-level files document further details of the campaign and data package: 1_CampaignSummary.csv provides details about the campaign and study site, 2_ScanAreasDetail.csv provides details about each separate scan area (groups of scans post-processed into a single point cloud), 3_TerrestrialLidarSensor.csv provides further technical details about the Riegl VZ-400i TLS sensor, TLS_CSV_dd.csv is a CSV Data Dictionary providing information about the fields in CSV files following the ESS-DIVE CSV File Formatting Guidelines Reporting Format, TLS_flmd.csv is a File Level Metadata file providing information about each file in the data package following the ESS-DIVE File Level Metadata Reporting Format, and README.txt is a text file describing the overall project and file structure. - Level 0 data are the raw data (.PROJ folders) as recorded by the Riegl VZ-400i TLS instrument before scan co-registration and post-processing with the Riegl's proprietary RiSCAN PRO software, which requires a license. - Level 1 data contain post-processed, co-registered data from each scan area. The "PointClouds" folder for each scan area contains a .las file with 1 cm resolution point cloud data exported from RiSCAN PRO. These are the main files likely to be of interest to most users and can be further processed with any software capable of manipulating .las files (e.g. Python, R CloudCompare). The "Project Information" folder contains log files from post-processing in RiSCAN PRO that may be of interest to users who want to see detailed records of post-processing, including all PDF reports generated by RiSCAN PRO. The "ScanPositions" folder contains information about the final position of all TLS scans, after post-processing, in multiple formats. The file ScanPositions_*.csv provides final geo-referenced scan positions, and the file SOP_backup_*.csv can be used in RiSCAN PRO to restore the co-registered scan positions if users wish to re-process raw data (Level 0 .PROJ folders) with RiSCAN PRO software (e.g., subsample to a different resolution, exclude a certain scan position, or apply different filters on reflectance or deviation values) without redoing time-consuming co-registration steps.

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with “Allometric scaling of hyporheic respiration across basins in the Pacific Northwest USA"

This data package is associated with the publication “Allometric scaling of hyporheic respiration across basins in the Pacific Northwest USA” submitted to JGR-Biogeosciences (Regier et al. 2025).This study used reach-scale modeled estimates of hyporheic aerobic respiration made by the River Corridor Model (Fang et al. 2020) and watershed characteristics across the Willamette and Yakima River basins to explore potential allometric scaling (i.e., power-law relationships between size and function) of cumulative hyporheic respiration across catchment-to-basin scales. Scaling was explored quantitatively via the R2, slope, and y-intercept of relationships between cumulative hyporheic respiration and watershed area, divided into hyporheic exchange flux (HEF) quantiles. We also explored relationships between allometric scaling and other watershed characteristics through linear regression, spatial patterns, and mutual information analyses. Our results also suggest variability of hyporheic respiration allometry for middle exchange flux quantiles, and in relation to land-cover. Our findings provide initial evidence that allometric scaling may be useful for predicting hyporheic biogeochemical dynamics across watersheds from reach to basin scales. This data package is associated with the GitHub repository found at https://github.com/peterregier/rc_wrb_yrb_scaling. The data package is organized into several key directories. The “data” folder contains multiple CSV files, including landscape heterogeneity, scaling analysis, and watershed boundary data. The “figures” folder has all figure files in both PDF and PNG formats. Core analysis scripts and figure generation scripts are in the “scripts” directory, systematically numbered for sequential execution. The root directory includes essential project files; please see the file ending in “flmd.csv” for a list and description of all files contained in this data package and the file ending in “dd.csv” for data dictionaries used to describe tabular column headers.

54 ENVIRONMENTAL SCIENCES↗

Data and scripts from: “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”

This data package includes data and scripts from the manuscript “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”.The study addressed common challenges faced in environmental sensing and modeling, including uncertain input data, missing sensor observations, and high-dimensional datasets with interrelated but redundant variables. Point-scaled meteorological and soil sensor observations were perturbed with noises and missing values, and denoising autoencoder (DAE) neural networks were developed to reconstruct the perturbed data and further predict evapotranspiration. This study concluded that (1) the reconstruction quality of each variable depends on its cross-correlation and alignment to the underlying data structure, (2) uncertainties from the models were overall stronger than those from the data corruption, and (3) there was a tradeoff between reducing bias and reducing variance when evaluating the uncertainty of the machine learning models.This package includes:(1) Four ipython scripts (.ipynb): “DAE_train.ipynb” trains and evaluates DAE neural networks, “DAE_predict.ipynb” makes predictions from the trained DAE models, “ET_train.ipynb” trains and evaluates ET prediction neural networks, and “ET_predict.ipynb” makes predictions from trained ET models.(2) One python file (.py): “methods.py” includes all user-defined functions and python codes used in the ipython scripts.(3) A “sub_models” folder that includes five trained DAE neural networks (in pytorch format, .pt), which could be used to ingest input data before being fed to the downstream ET models in ‘ET_train.ipynb” or ‘ET_predict.ipynb’.(4) Two data files (.csv). Daily meteorological, vegetation, and soil data is in “df_data.csv”, where “df_meta.csv” contains the location and time information of “df_data.csv”. Each row (index) in “df_meta.csv” corresponds to each row in “df_data.csv”. These data files are formatted to follow the data structure requirements and be directly used in the ipython scripts, and they have been shuffled chronologically to train machine learning models. The meteorological and soil data was collected using point sensors between 2019-2023 at(4.a) Three shrub-dominated field sites in East River, Colorado (named “ph1”, “ph2” and “sg5” in “df_meta.csv”, where “ph1” and “ph2” were located at PumpHouse Hillslopes, and “sg5” was at Snodgrass Mountain meadow) and(4.b) One outdoor, mesoscale, and herbaceous-dominated experiment in Berkeley, California (named “tb” in “df_meta.csv”, short for Smartsoils Testbed at Lawrence Berkeley National Lab).- See "df_data_dd.csv" and "df_meta_dd.csv" for variable descriptions and the Methods section for additional data processing steps. See "flmd.csv" and "README.txt" for brief file descriptions.- All ipython scripts and python files are written in and require PYTHON language software.

54 ENVIRONMENTAL SCIENCES↗

Data for Machado-Silva et al. (2024), "Short-Term Groundwater Level Fluctuations Drive Subsurface Redox Variability"

This dataset contains the analytical data reported in Machado-Silva et al. (2024) as part of the COMPASS-FME project, which seeks to advance a scalable, predictive understanding of the fundamental biogeochemical processes, ecological structure, and ecosystem dynamics that distinguish coastal terrestrial-aquatic interfaces from the purely terrestrial or aquatic systems to which they are coupled. The dataset consists of water quality parameters as well as redox potential, water content, and electrical conductivity. These data were collected in 2022 in Crane Creek (CRC), Portage River (PTR), and Old Woman Creek (OWC). Each of these sites included uplands (UP), transitions (TR), wetland-transition edge (WTE), and wetland (W) zones. The sites represent replicates of the Lake Erie terrestrial-aquatic interface under fluctuating water levels and are located in well-preserved areas with natural or restored marsh and forest cover.This dataset consists of a single data file (Machado_Silva_et_al_2024_EST_data.csv) that is in comma-separated value (CSV) format. No special software is required to read it.This dataset uses the ESS-DIVE Hydrologic Monitoring Reporting Format 1.0.

54 ENVIRONMENTAL SCIENCES↗

Data for Myers-Pigg et al. (2026), "Short-term coastal forest responses to a hurricane-scale freshwater and saltwater flooding experiment"

Coastal upland forests are exposed to intensifying precipitation regimes and sea level rise, increasing tree mortality and transforming these coastal forests into wetland ecosystems. Despite these well-known risks, the differing degrees to which hydrological, biogeochemical, and biological components of upland forests respond to novel salinity exposure is relatively unknown. The Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) experiment decouples two distinct disturbances associated with hydrological extremes: (1) flooding from heavy precipitation and (2) exposure to saline conditions from storm surge. This dataset includes data reported in Myers-Pigg et al. (2025), which analyzed data from the first TEMPEST flooding treatment in 2022. This includes: - Colored dissolved organic matter in porewaters - Soil temperature and oxygen - Groundwater temperature and chemistry - Dissolved organic carbon concentrations in porewaters - Soil-to-atmosphere CH4 and CO2 fluxes - Soil temperature, water content, and electrical conductivity - Root-influenced CH4 and CO2 flux - Tree sap flow velocity - The R analytical code and documentation about the computational environmental in which it was run (the "sessionInfo.txt" file) All data files are plain-text comma separated value (CSV) and no special software is required to read them.

54 ENVIRONMENTAL SCIENCES↗

EXCHANGE Campaign Degradation (ECD): Understanding Decomposition Dynamics Across Mid-Atlantic and Great Lakes Coastal Ecosystems

The EXploration of Coastal Hydrobiogeochemistry Across a Network of Gradients and Experiments (EXCHANGE) Degradation Experiment (EXCHANGE-D) is an in situ experiment designed to assess organic matter decomposition rates across coastal terrestrial-aquatic interfaces (TAIs), from coastal uplands through transition zones to wetlands. Through a network of partner scientists and coastal sites, we are testing how environmental gradients shape decomposition and carbon dynamics across terrestrial-aquatic interfaces. Using standardized tea bag substrates deployed across a network of diverse coastal sites, we compare decomposition rates at different fresh- and salt-water TAIs to develop transferable knowledge that improves the representation of organic matter degradation in coastal ecosystem models. For more information, please see https://compass.pnnl.gov/FME/EXCHANGE. This is Version 1 of the data package, which includes: ecd_README.pdf flmd.csv dd.csv ecd_soil_weom_L2.csv ecd_soil_ph_conductivity_L2.csv ecd_soil_gwc_L2.csv ecd_soil_teabag_degradation_L2.csv ecd_readme.pdf

coastal soils↗

Data from: "Reply to ‘The challenge of defining effectively-no-snow’"

This repository contains the data and code associated with the paper titled "Reply to ‘The challenge of defining effectively-no-snow’" published in Nature Reviews Earth and Environment, 2026. In this reply, we argue that the 10th percentile of peak SWE (Snow Water Equivalent), which we propose in the original article, can be used as intended given it's a standardized, impact-based benchmark for comparing snow conditions across regions, not as a literal measure of snow absence. We present new evidence with SNOwpack TELemetry (SNOTEL) data showing that years meeting the threshold are overwhelmingly associated with subsequent drought (given United States Drought Monitor conditions), supporting its hydrologic and societal relevance. We conclude that while the distinction between "effectively no snow" and "zero snow" should be clearly communicated, the original definition remains appropriate for assessing impacts on snow-dependent water systems. The file code_nree_ML_reply_2026.Rmd contains the main processing scripts which analyze the SNOTEL data. Data from the US Drought Monitor was downloaded at: https://usdmdataservices using the Get Drought Severity Statistics By Area Percent' option, saved to the *_HUC4_delineated.csv files (Hydrologic Unit Code), which are labeled accordingly. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

EARTH SCIENCE > TERRESTRIAL HYDROSPHERE > SNOW/ICE↗

iButton and Tinytag snow temperature measurements at Teller 27 and Kougarok 64, Seward Peninsula, Alaska, 2021-2022

Snow temperature measurements were collected at the NGEE Arctic Teller Road Site at mile marker 27 (TL_MM27) and at the Kougarok Road Site at mile marker 64 (KG_MM64) on the Seward Peninsula. Data were collected from October 1, 2021 to August 16, 2022 using iButton Link DS1926-F5# Thermochron miniature temperature sensors (https://www.ibuttonlink.com/products/ds1921g) and Tinytag TGP-4017 internal sensors (https://www.micronmeters.com/product/tgp-4017-internal-sensor-40-to-85-c-40-f-to-185-f) deployed across the Kougarok and Teller sites. These sensors are a cost-efficient way to collect snowpack temperatures at a higher spatial resolution than what is normally achieved. iButton data were collected every 4 hours beginning on October 1, 2021. Tinytag data collection began between October 9 and October 12, 2021 depending on sensor installation date. Tinytag data were collected every 30 minutes. In total, data were collected from 236 iButtons and 30 Tinytags. This dataset contains four *.csv files of near-ground surface temperatures at various locations throughout each study site and four *.shp files of sensor locations. Data were collected throughout the snow cover season so that snowpack characteristics could be derived using the temperature data. Sensors were placed both inside and outside of vegetation to better capture the spatial variability of snow properties across each domain. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Laboratory time series moisture manipulative experiment from sediment across San Antonio, Texas: time series aerobic respiration and geochemistry

This dataset supports a broader study examining the effects of wetting and drying on hyporheic zone respiration. The dataset provides data generated from a laboratory moisture manipulation experiment. The contents include time series aerobic respiration and moisture; dissolved oxygen; sediment geochemistry data; and field metadata (including qualitative information on instream and river corridor characteristics). Samples were collected as part of the WHONDRS Allison Veach collaboration (AV1). The data package associated with the AV1 study is available at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2529428. AV1 sampling occurred across 7 perennial and 7 intermittent streams in San Antonio, Texas. Each stream/site was visited both in summer during base flow (July-September 2023) and winter during peak flow (January-February 2024). This study uses subsamples from a subset of AV1 samples. The original field samples were labeled as AV1_###. Subsequent subsamples for this study were labeled as EV_###. The labels from the field samples and the EV subsamples can be mapped directly based on the digits following the prefix and underscore (i.e., EV_001 is a subsample from AV1_001). See the critical details section below for more details on sample naming. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) field protocol; and a (6) a subfolder with sediment sample data from the incubation experiment. The sample data subfolder contains (1) effect size; (2) iron (II); (3) gravimetric moisture; (4) respiration rates; (5) raw dissolved oxygen values and plots; (6) specific conductance; (7) pH; (8) temperature; (9) a summary containing mean, median, and standard deviation values of each data type for each treatment (wet and dry); and (10) methods codes. All files are .csv or.pdf.

54 ENVIRONMENTAL SCIENCES↗

WHONDRS Surface Water and Sediment Geochemistry and Organic Matter Characterization Data from Streams across HJ Andrews Experimental Forest, Oregon (v2)

This dataset supports a broader study developing conceptual models for river corridor critical zone processes across spatial scales and was generated in collaboration with the HJ Andrews River Corridor Critical Zone Workshop in 2025. The dataset provides surface water geochemistry (dissolved organic carbon, total dissolved nitrogen) from 48 sites across the HJ Andrews Experimental Forest, Oregon (https://andrewsforest.oregonstate.edu). Some of the sites have been impacted by the Holiday Farm Fire and the Lookout Fire in 2020 and 2023, respectively. Related data were collected as part of the workshop and will be published separately in collaboration with other workshop attendees and available at http://www.hydroshare.org/resource/b274c4a234bf4b12b7cb8a54a696c629. Related genomic data can be found on the National Center for Biotechnology Information (NCBI) under BioProject PRJNA1503030 (see critical details section below for more information). Additional related data collected in 2016 from a similar effort can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/3377027 and http://www.hydroshare.org/resource/ea6c0832885a46c3939e7bb22e48e754 and are described within https://doi.org/10.5194/essd-11-1-2019 (Ward et al., 2019). This data package was originally published in March 2026. It was updated in August 2026 (v2; new and modified files). See the change history section in the readme for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) a folder of field photos, (2) a folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data, (3) a data checks report, (4) a folder of sample data, (5) file-level metadata, (6) data dictionary, (7) field metadata, (8) readme, (9) international generic sample number (IGSN) mapping file; and (10) field protocol. The sample data subfolder contains surface water and sediment (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages, (2) total dissolved nitrogen data and averages, (3) methods codes, (4) FTICR-MS methods; and (5) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains the CoreMS processed data and seven subfolders, thee containing .xml files for each sample type (sediment, surface water and blank samples), three containing the sediment CoreMS output files for each sample type (sediment, surface water and blank samples), and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .Rmd, .py, .cal, .json, .jpg, or .jpeg.

Biogeochemistry↗

Metagenome-assembled genomes from soil samples in control and warming plots in Blodgett Forest, CA (2014-2021)

The pathways of carbon transport and loss through and from soils—soil organic matter (SOM) depolymerization to dissolved organic carbon and mineralization to carbon dioxide (CO2)—are fundamentally driven by microbial activity, which is strongly regulated by environmental conditions. As part of LBNL (Lawrence Berkeley National Laboratory) TES (Terrestrial Ecosystem Science) Belowground Biogeochemistry Science Focus Area (SFA), we have established a novel whole-soil long-term warming experiment at the University of California (UC) Blodgett Forest Research Station (Sierra Nevada) in 2014, where we study the role of biogeochemical, microbial and geochemical process interactions in SOM decomposition and stabilization.Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community from soil depth profiles collected from 2014 to 2021 from three paired control and warming plots. We collected soil samples across a range of depth profiles (spanning surface to 90 cm deep) from three paired control and warming plots from a temperate mixed forest in Northern California. Each paired plot had been subjected to experimental warming since June 2014 to simulate a predicted climate change scenario for northern California. 101 soil metagenomes were sequenced at JGI (Joint Genome Institute) and UCSF (University of California San Francisco) Center for Advanced Technology and can be found under the JGI (Joint Genome Institute) GOLD (Genomes Online Database) Sequencing project Gs0151586 and NCBI (National Center for Biotechnology Information) Projects PRJNA1225762 and PRJEB39497. Metagenomes were assembled using JGI (Joint Genome Institute) Metagenome Workflow (10.1128/mSystems.00804-20). For each metagenome, the assembled contigs were binned into genomes using 3 binning algorithms (cocacola, metabat, and maxbin) and the resulting bins were consolidated using dastool. The consolidated bins from all metagenomes were pooled, filtered by completeness (>50%) and contamination (<25%), and dereplicated at 99% ANI (average nucleotide identity) using dRep (https://github.com/MrOlm/drep).The dataset includes a zip file of 2321 MAG (Metagenome Assembled Genome) fasta files, the accession numbers for the underlying metagenomes, and a csv file with MAG (Metagenome Assembled Genome) quality metrics and taxonomic classification (GTDB -Genome Taxonomy Database-RS220). This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type. A sample metadata file (samples.csv) that contains site information has also been included.

54 ENVIRONMENTAL SCIENCES↗

FAIR Ecosystems for Science at Scale

High Performance Computing (HPC) centers provide resources to users who require greater scale to “get science done”. They deploy infrastructure with singular hardware architectures, cutting-edge software environments, and stricter security measures as compared with users’ own resources. As a result, users often create and configure digital artifacts in ways that are specialized for the unique infrastructure at a given HPC center. Each user of that center will face similar challenges as they develop specialized solutions to take full advantages of the center’s resources, potentially resulting in significant duplication of effort. Much duplicated effort could be avoided, however, if users of these centers found it easier to discover others’ solutions and artifacts as well as share their own. The FAIR principles address this problem by presenting guidelines focused around metadata practices to be implemented by vaguely defined “communities”; in practice, these tend to gather by domain (e.g. bioinformatics, geosciences, agriculture). Domain-based communities can unfortunately end up functioning as silos that tend both to inhibit sharing of solutions and best practices as well as to encourage fragile and unsustainable improvised solutions in the absence of best-practice guidance. We propose that these communities pursuing “science at scale” be nurtured both individually and collectively by HPC centers so that users can take advantage of shared challenges across disciplines and potentially across HPC centers. We describe an architecture based on the EOSC-Life FAIR Workflows Collaboratory, specialized for use with and inside HPC centers such as the Oak Ridge Leadership Computing Facility (OLCF), and we speculate on user incentives to encourage adoption. We note that a focus on FAIR workflow components rather than FAIR workflows is more likely to benefit the users of HPC centers.

Wilkinson, Sean [ORNL] (ORCID:0000000214437479)↗

Temporal Study 2022-2024: Sensor-Based Time Series of Surface Water Temperature, Specific Conductance, Total Dissolved Solids, Turbidity, Chlorophyll A, and Dissolved Oxygen from across Multiple Watersheds in the Yakima River Basin in Washington, USA

This dataset supports a broader study examining the drivers of temporal variability in sediment respiration rates in the Yakima River Basin. The dataset provides periodic (bi-weekly or monthly) in situ hydrological and water chemistry sensor data, handheld sensor water chemistry data, general environmental context photos, and field metadata collected at six sites across the Yakima River Basin in Washington, USA. Sample and sensor data from previous years (2021-2022) can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1898912 and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1892054, respectively. Related sample data from 2022-2024 are available at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2562910. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions This dataset contains a folder of environmental context photographs and videos and (1) file-level metadata; (2) data dictionary; (3) readme; (4) field metadata; (5) field protocols; (6) international generic sample number (IGSN) mapping file; (7) handheld sensor data; and (8) two sensor subfolders. Each sensor subfolder (BarotrollAtm and MantaRiverData) contains a subfolder containing sensor time series data and plots. The BarotrollAtm Data subfolder contains In Situ Rugged BaroTROLL sensor pressure and air temperature data. The MantaRiverData subfolder contains Eureka Manta+ 35B multisonde temperature, specific conductance, and chlorophyll A. All files are .csv, .pdf, .jpg, .jpeg, .mp4, .png, or .mov.

54 ENVIRONMENTAL SCIENCES↗

Shaping the Future of Self-Driving Autonomous Laboratories Workshop

The "Shaping the Future of Self-Driving Autonomous Laboratories" workshop, held in Denver on November 7-8, 2024, brought together leading experts from materials science and computing to address the growing need to revolutionize scientific research through AI-driven autonomous laboratories. The workshop identified critical challenges, including the integration of heterogeneous data, development of AI systems that understand fundamental physical principles, and comprehensive safety protocols. Key recommendations emerged around developing universal laboratory equipment interfaces, implementing automated metadata collection systems, and creating hybrid AI approaches that combine data-driven learning with scientific principles. The workshop emphasized maintaining human oversight while leveraging automation, transforming scientific education to prepare the next generation of researchers, and establishing a national consortium leveraging DOE facilities as anchors for broader collaboration with academia and industry. Participants stressed the urgency of addressing the growing disconnect between human decision-making timescales and modern instrumentation capabilities, highlighting the need for strategic automation while preserving essential human insight and oversight in the research process.

36 MATERIALS SCIENCE↗

Human Liver Epithelium Response to HCoV-229E Infection Epigenomics (ACS-DP4)

The purpose of this experiment was to evaluate how wild-type Human coronavirus strain 229E (HCoV-229E) infection alters chromatin accessibility in infected cells only. Sample data was obtained for mock and infected (standard and UV-inactivated) immortalized human liver cells (HuH-7) and collected 24 hrs. post infection. Samples were processed using assay for transposase-accessible chromatin using high-throughput sequencing (ATAC-Seq) and generated bar coded library samples were evaluated for RNA sequencing (RNA-Seq) expression analysis. Processed ATAC-Seq datasets are openly accessible from the download button and contain secondary processed RNA-Seq results files and supporting metadata materials. Data download includes a sample naming key, infection titer metadata, normalized counts, and relevant computational source code information supporting data transparency and reuse.

59 BASIC BIOLOGICAL SCIENCES↗

Redox Potential of Intermittently Wet Soil, Old Woman Creek National Estuarine Research Reserve, Huron, OH, 2023-05-10 to 2023-12-15

This dataset contains collected reduction-oxidation (redox) potential measurements of the underlying soil at various depths within a wetland, referred to as The Cove, at Old Woman Creek Estuarine Research Reserve in Huron, OH. Redox potential was measured in the sediments of a coastal wetland to assess how redox potential varies over time with changes in hydrological events. Measurements were collected by Campbell Scientific CR1000X dataloggers paired with a PaleoTerra redox probe and reference electrodes. Measurements were collected at 3 different locations within The Cove at multiple depths into the underlying soil of the wetland and were collected every 10 minutes. The Redox_Datafile.csv contains the recorded measurements of the redox probes, and the RefElecDataFile.csv contains the background-noise measurements collected by the reference electrodes. Redox_InstallMethods describes the installation methods of the datalogger and associated probes.

54 ENVIRONMENTAL SCIENCES↗