Search NASA⌕ Search

SEARCH · Search NASA

Results for “interface science metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Five Years of Dissolved Oxygen, Temperature, Salinity, Depth, Weather Data from a Transitioning Wetland at Beaver Creek, Washington, USA

Groundwater dissolved oxygen (DO) variability in coastal system remains poorly understood despite its importance for biogeochemical cycling and ecosystem modeling. Here we investigate the temporal variability in groundwater DO and its hydro-climatic drivers across hourly to seasonal timescales in a transitioning wetland at Beaver Creek, Washington, USA. The site is transitioning from a freshwater forest to a brackish tidal wetland following removal of a barrier in 2014 that prevented tides from accessing the freshwater creek. By utilizing novel optical dissolved oxygen instrumentation (Opti O2, LLC) we obtained continuous, high-frequency (5-minute), in-situ measurements of DO from the flood-plain from June 26th, 2019 through September 30th, 2024. This 63 month dataset is comprised of groundwater dissolved oxygen, temperature, water level and salinity timeseries from the floodplain. This dataset also includes rainfall, air pressure, air temperature, and solar radiation data collected with a co-located Campbell ClimaVUE50 weather sensor. All data is contained within a single csv (2019-06-26 to 2024-09-30 Beaver Creek DO, saln, BGS, temp, weather.csv) that can easily be viewed either using software such as Excel or using any text editor.

54 ENVIRONMENTAL SCIENCES↗

Shaping the Future of Self-Driving Autonomous Laboratories Workshop

The "Shaping the Future of Self-Driving Autonomous Laboratories" workshop, held in Denver on November 7-8, 2024, brought together leading experts from materials science and computing to address the growing need to revolutionize scientific research through AI-driven autonomous laboratories. The workshop identified critical challenges, including the integration of heterogeneous data, development of AI systems that understand fundamental physical principles, and comprehensive safety protocols. Key recommendations emerged around developing universal laboratory equipment interfaces, implementing automated metadata collection systems, and creating hybrid AI approaches that combine data-driven learning with scientific principles. The workshop emphasized maintaining human oversight while leveraging automation, transforming scientific education to prepare the next generation of researchers, and establishing a national consortium leveraging DOE facilities as anchors for broader collaboration with academia and industry. Participants stressed the urgency of addressing the growing disconnect between human decision-making timescales and modern instrumentation capabilities, highlighting the need for strategic automation while preserving essential human insight and oversight in the research process.

36 MATERIALS SCIENCE↗

SysCaps (Language Interfaces for Simulation Surrogates of Complex Systems) [SWR-24-97]

You've found the official code repository for the paper "SysCaps: Language Interfaces for Simulation Surrogates of Complex Systems," presented at the Foundation Models for Science: Progress, Opportunities, and Challenges workshop at NeurIPS 2024. Our paper conjectures that interfaces (both text templates as well as conversational) makes interacting with simulation surrogate models for complex systems more intuitive and accessible for both non-experts and experts. "System captions", or SysCaps, are text-based descriptions of systems based on information contained in simulation metadata. Our paper's goal is to train multimodal regression models that take text inputs (SysCaps) and timeseries inputs (exogenous system conditions such as hourly weather) and regress timeseries simulation outputs (e.g. hourly building energy consumption). The experiments in our paper with building and wind farm simulators, which can be reproduced using this codebase, aim to help us understand whether a) accurate regression in this setting is possible and b) if so, how well can we do it. Paper: https://arxiv.org/abs/2405.19653

Emami, Patrick↗

EXCHANGE Campaign Degradation (ECD): Understanding Decomposition Dynamics Across Mid-Atlantic and Great Lakes Coastal Ecosystems

The EXploration of Coastal Hydrobiogeochemistry Across a Network of Gradients and Experiments (EXCHANGE) Degradation Experiment (EXCHANGE-D) is an in situ experiment designed to assess organic matter decomposition rates across coastal terrestrial-aquatic interfaces (TAIs), from coastal uplands through transition zones to wetlands. Through a network of partner scientists and coastal sites, we are testing how environmental gradients shape decomposition and carbon dynamics across terrestrial-aquatic interfaces. Using standardized tea bag substrates deployed across a network of diverse coastal sites, we compare decomposition rates at different fresh- and salt-water TAIs to develop transferable knowledge that improves the representation of organic matter degradation in coastal ecosystem models. For more information, please see https://compass.pnnl.gov/FME/EXCHANGE. This is Version 1 of the data package, which includes: ecd_README.pdf flmd.csv dd.csv ecd_soil_weom_L2.csv ecd_soil_ph_conductivity_L2.csv ecd_soil_gwc_L2.csv ecd_soil_teabag_degradation_L2.csv ecd_readme.pdf

coastal soils↗

Gap-filled methane and carbon dioxide fluxes across two ecosystem states at the US-OWC AmeriFlux site (2015−2016, 2020−2022)

This dataset contains gap-filled measurements of methane flux (FCH4), net ecosystem CO2 exchange (NEE) partitioned into gross primary productivity (GPP) and ecosystem respiration (RE), as well as latent heat flux (LE) from a Great Lakes coastal freshwater wetland at the US-OWC AmeriFlux site. The dataset covers the peak growing seasons (June−September) of 2015−2016, dominated by Typha spp., and 2020−2022, characterized by floating-leaved species (lotus and water lily). These data were generated to investigate how rising water levels and vegetation shifts influence CH4 and CO2 fluxes across two distinct ecosystem states in this wetland. The dataset, provided in CSV format, includes half-hourly gap-filled flux data from June to September for 2015, 2016, 2020, 2021, and 2022. The gap-filled data refers to measurements where missing values due to instrument issues or quality control were filled using artificial neural networks (ANNs).

54 ENVIRONMENTAL SCIENCES↗

Label-based Virtual Directories In dCache

Traditional filesystems organize data in directories. These directories are typically a collection of files whose grouping is based on a single criterion, e.g., the starting date of an experiment, experiment name, beamline ID, measurement device, or instrument. However, each file in a directory can belong to several logical groups, such as a special event type, experiment condition, or a part of a selected dataset. dCache is a storage system developed to store large amounts of scientific data, used by many HEP and Photon Science experiments. With recent developments in dCache, we have introduced a concept of file tagging, which dynamically groups files with the same label into virtual directories. The file labels can be added, removed, renamed, and deleted through the admin interface or via REST API. The files in virtual directories are exposed through all protocols supported by dCache. This contribution will describe the details of the implementation for file tagging in dCache and present our future development plans on automatic metadata extractions, a feature that will significantly simplify data management. Additionally, we are exploring the future use of virtual directories as a way to translate scientific data catalogs into filesystem views for direct data analysis.

Sahakyan, Marina [DESY]↗

A Data Science and Machine Learning Platform Supporting Large Particle Accelerator Control and Diagnostics Applications Final Report: SBIR Initial Phase II DE-SC0022583

The Machine Learning Data Platform (MLDP) is a product providing full-stack support for data science, Machine Learning, and Artificial Intelligence (ML/AI) applications at particle accelerator and large experimental physics facilities. It supports ML/AI applications from front-end, high-speed acquisition of heterogeneous, time-series data, through data archiving and management, to back-end analysis. The MLDP embodies a “data-science ready” platform for data analysis and ML/AI applications in diagnosis, modelling, control, and optimization of these facilities. It provides data scientists and applications a consistent, datacentric interface to archive data standardizing implementation and deployment of ML/AI algorithms to different operations configurations within the same facility, or between facilities. Being an open-source, public-domain project, the MLDP is intended for broadest possible impact by increasing accessibility and minimizing the required expertise for installation and operation. The MLDP can also be deployed at user facilities for experimental data collection, archiving, and analysis. It is capable of acquisition and archiving of heterogeneous data from experimental equipment (e.g., images, arrays, structures, etc.) along with system hardware configurations (e.g., scalars, tables), control system process variables, and any metadata required for provenance. Thus, the MLDP can manage experimental data through its entire lifecycle, from acquisition and archiving, through analysis and investigation, to release and final publication.

43 PARTICLE ACCELERATORS↗

BGC Atlas: a web resource for exploring the global chemical diversity encoded in bacterial genomes

Secondary metabolites are compounds not essential for an organism’s development, but provide significant ecological and physiological benefits. These compounds have applications in medicine, biotechnology and agriculture. Their production is encoded in biosynthetic gene clusters (BGCs), groups of genes collectively directing their biosynthesis. The advent of metagenomics has allowed researchers to study BGCs directly from environmental samples, identifying numerous previously unknown BGCs encoding unprecedented chemistry. Here, we present the BGC Atlas (https://bgc-atlas.cs.uni-tuebingen.de), a web resource that facilitates the exploration and analysis of BGC diversity in metagenomes. The BGC Atlas identifies and clusters BGCs from publicly available datasets, offering a centralized database and a web interface for metadata-aware exploration of BGCs and gene cluster families (GCFs). We analyzed over 35 000 datasets from MGnify, identifying nearly 1.8 million BGCs, which were clustered into GCFs. The analysis showed that ribosomally synthesized and post-translationally modified peptides are the most abundant compound class, with most GCFs exhibiting high environmental specificity. We believe that our tool will enable researchers to easily explore and analyze the BGC diversity in environmental samples, significantly enhancing our understanding of bacterial secondary metabolites, and promote the identification of ecological and evolutionary factors shaping the biosynthetic potential of microbial communities.

59 BASIC BIOLOGICAL SCIENCES↗

Old Woman Creek Wetland Sediment and Electrochemical Sensor Microbial Community, 2023

We are developing a technique to monitor microbiological activities referred to as zero resistance ammetry, which entails the deployment of graphite electrodes in sediments. Measurement of current between electrodes of contrasting redox regimes and/or predominant terminal electron accepting processes can be used as an indicator of the extents of microbiological activity. We deployed an electrode array at depths of 2 mm, 4 mm, 76 mm, 78 mm, 152 mm, 154 mm, 227 mm, and 229 mm below the wetland sediment water interface in the Old Woman Creek National Estuarine Research Center, Huron, OH, USA (Lat. = 41.380833, Long. = -82.508889). A core was collected from adjacent sediment and subsamples were collected from depth intervals of 0 – 25 mm, 25 – 127 mm, 127 – 128 mm, and below 178 mm. To determine if the microbial communities attached to the electrodes were reflective of the adjacent sediment-associated microbial community, we conducted a 16S rRNA gene-based (V4 region) survey of these respective materials. This data package contains the results of these surveys, including metadata on the depths from which samples were collected (samples.csv), DNA extraction and sequencing information (OWC_DEPTH_AMPLICON_SEQUENCING_METADATA), sequence processing information (OWC_DEPTH_BIOINFORMATIC_METADATA.csv), an operational taxonomic unit (OTU) table (OWC_DEPTH_97OTUS_TABLE.csv), and nucleotide sequences of OTUs (OWC_DEPTH_97OTUS_SEQS.fasta). All files can be opened using a text-editing application. The fasta file is compatible with bioinformatics applications.

54 ENVIRONMENTAL SCIENCES↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

DOE Repository Metadata Profile (DRMP): A Metadata Framework for Advancing Interoperability and AI Readiness Across Scientific Repositories

The Department of Energy (DOE) funds a diverse and distributed ecosystem of repositories that steward scientific data, publications, and software across its research programs, user facilities, and national laboratories. While significant progress has been made in standardizing dataset-level metadata, the metadata describing repositories themselves (their identity, governance, access interfaces, policies, and technical capabilities) remains inconsistent and fragmented across DOE-funded systems. This variability limits discoverability, interoperability, automated validation, and AI-driven analysis, all of which are increasingly essential for modern scientific workflows. To address this gap, the DOE Data Curation Working Group (DCWG) developed the DOE Repository Metadata Profile (DRMP). The DRMP is a practical, community-driven framework that defines how repositories can describe themselves in a consistent, machine-actionable, and scalable manner. The DRMP is not a new metadata schema. Instead, it is a mapping profile and structured element set capturing the essential characteristics of DOE repositories. It harmonizes repository-level metadata across six widely adopted community schemas: RE3Data; DCAT-US v3; Schema.org; Dublin Core; DataCite 4.6; and PREMIS 3.0. This harmonization eliminates reinvention and enables interoperability within DOE and across the broader scientific ecosystem. A core objective of the DRMP is to reduce burden on repositories by allowing them to reuse their existing metadata through a Rosetta-style crosswalk rather than redesigning local implementations. The profile introduces a three-level conformance model that supports incremental adoption: • Level 1 – Minimum Viable Record (MVR): foundational identification elements required for workflows, project registration, and basic repository presence. • Level 2 – Interoperable: structured metadata enabling alignment with national and international discovery systems. • Level 3 – AI-Ready: enhanced provenance, policy transparency, fixity, semantic context, and capabilities that support automated reasoning, model training governance, and machine-assisted curation. To support implementation, the DRMP includes JSON Schema definitions, OpenAPI patterns, and MCP templates that allow repositories to publish machine-readable metadata directly within existing platforms. These resources are modular and lightweight, enabling adoption without major architectural change. Adopting the DRMP enables repositories to: • Enhance discoverability and interoperability by aligning identifiers, classifications, and descriptive elements across widely used schema standards. • Support federated discovery and cross-registration across DOE systems, Data.gov, and international catalogs. • Enable AI agents and workflow orchestration systems to interpret repository-level metadata within the American Science Cloud (AmSC) through Model Context Protocol (MCP)-based context publication. • Demonstrate alignment with DOE’s open science, stewardship, and FAIR data priorities. This guidance represents a community-driven step forward. Through voluntary adoption and continued feedback, the DRMP advances a cohesive, machine-actionable description of DOE repositories that supports FAIR data practices, preparing the infrastructure for AI-enabled research, and strengthening the discoverability and reuse of DOE’s scientific outputs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Soil microbial ecology and microbiome-metabolite linkages improve understanding of ecosystem states along terrestrial-aquatic interfaces

These data are from Bandopadhyay et al., "Soil microbial ecology and microbiome-metabolite linkages improve understanding of ecosystem states along terrestrial-aquatic interfaces". This study aims to understand the soil microbial ecology along terrestrial-aquatic interfaces of a freshwater and estuarine region and how it relates to organic matter. We analyzed soil microbial (16S rRNA gene) and organic matter (Fourier-transform ion cyclotron resonance mass spectrometry, FTICR-MS) composition from upland (forested), transition (stressed forest), and wetland positions at three sites in each of the Lake Erie (freshwater) and Chesapeake Bay (estuarine) regions. This dataset includes 16S rRNA gene amplicon data (only processed file types included here) and organic matter composition from FTICR-MS data (raw and processed files included here) from upland (forested), transition (stressed forest), and wetland positions at three sites in each of the Lake Erie and Chesapeake Bay regions. These sites are part of the COMPASS-FME project (https://compass.pnnl.gov/FME/COMPASSFME). File formats and software needed to access files: 16S rRNA gene amplicon data: These files follow the format reported here https://ess-dive.gitbook.io/amplicon-sequencing-reporting-format#updates-in-v1.0.1. As per this format, there are four file types reported: 1. Taxon tables (also called sequence-by-sample or OTU (operational taxonomic unit)/ESV (exact sequence variant) tables) : available in a .txt file format and accessible using TextEdit or MS Excel. 2. Representative sequences (also called consensus sequences) : available in a .fasta format and accessible using TextEdit. 3. Sequencing metadata : available in a MS Excel workbook file format and CSV file format 4. Bioinformatic metadata : available in a MS Excel workbook file format and CSV file format FTICR-MS data: 1. Raw data converted to a processed file with intensities of the peaks in the given samples : available in a MS Excel CSV file format 2. Processed file used in analyses and visualizations (appended as icr_long_) : available in a MS Excel CSV file format 3. Metadata file for ICR features (appended as icr_meta) : available in a MS Excel CSV file format

54 ENVIRONMENTAL SCIENCES↗

Data for Machado-Silva et al. (2024), "Short-Term Groundwater Level Fluctuations Drive Subsurface Redox Variability"

This dataset contains the analytical data reported in Machado-Silva et al. (2024) as part of the COMPASS-FME project, which seeks to advance a scalable, predictive understanding of the fundamental biogeochemical processes, ecological structure, and ecosystem dynamics that distinguish coastal terrestrial-aquatic interfaces from the purely terrestrial or aquatic systems to which they are coupled. The dataset consists of water quality parameters as well as redox potential, water content, and electrical conductivity. These data were collected in 2022 in Crane Creek (CRC), Portage River (PTR), and Old Woman Creek (OWC). Each of these sites included uplands (UP), transitions (TR), wetland-transition edge (WTE), and wetland (W) zones. The sites represent replicates of the Lake Erie terrestrial-aquatic interface under fluctuating water levels and are located in well-preserved areas with natural or restored marsh and forest cover.This dataset consists of a single data file (Machado_Silva_et_al_2024_EST_data.csv) that is in comma-separated value (CSV) format. No special software is required to read it.This dataset uses the ESS-DIVE Hydrologic Monitoring Reporting Format 1.0.

54 ENVIRONMENTAL SCIENCES↗

Continuous snow depth and temperature measurements from dense network of above-ground distributed temperature profiling systems from 2021-09-23 to 2024-08-23, Seward Peninsula, Alaska

The dataset contains temperature measurements from distributed temperature profiling (DTP) systems (Dafflon et al., 2022; Wielandt et al., 2022; Wang et al., 2024a; Fiolleau et al., 2024) deployed vertically above the ground surface at a large number of locations from 2021 to 2024. The research is designed to improve understanding of the local heterogeneity in snow depth and snow thermal insulation dynamics, as well as their interactions in a discontinuous permafrost region (Wang et al., 2025). The DTP systems were deployed at 96 locations in a watershed along the Nome-Teller road at mile marker 27 (T27) and at 54 locations on a hillslope along the Kougarok road at mile marker 64 (K64) in the Seward Peninsula, Alaska. The probe location information is stored in Probe_locations_T27.csv and Probe_locations_K64.csv. Temperature measurements were recorded at 15-minute intervals using high-precision digital sensors (accuracy: ±0.1°C, resolution: 0.0078°C). The temperature probes, either 1.4 m or 1.6 m long, contain sensors spaced every 5 cm or 10 cm along their length. The temperature data are stored in compressed files following the format: DTP_snow_air_temperature_(site)_(start)_(end).zip, where site is either T27 or K64, and start and end represent the time series period. Within each ZIP file, individual CSV files are named by probe ID and contain temperature records at different heights above the ground surface.This dataset also includes derived snow depth time series over three snow seasons, estimated from temperature measurements. Snow depth was estimated by identifying the consecutive sensor pair that exhibited the largest drop in high-frequency temperature fluctuations (detailed in the methods). These data are stored in: Snow_depths_flags_(site)_(start)_(end).csv, which includes snow depth time series and corresponding quality flags (defined in the methods) from different probes. Additionally, the dataset includes derived metrics and supporting measurements at selected locations over two snow seasons, contributing to the manuscript of Wang et al., 2025. These locations were chosen based on the availability of high-quality snow depth time series during both seasons. The additional data include: (1) Air temperature proxies measured from the top sensors on the pole when they were not buried by snow, stored in Air_temperature_proxies_(site)_(start)_(end).csv (2) Ground interface temperature, recorded at 3 cm above the ground, stored in Ground_interface_temperature_(site)_(start)_(end).csv (3) Site characteristics, including vegetation height, elevation, and the topographic position index (TPI) within a 50 m radius, stored in Selected_probe_locations_gps_vegheight_tpi_elevation_(site).csv. These metrics were derived from 1 m resolution summer LiDAR-based digital elevation models and digital surface models from Singhania et al., 2023, DOI:10.5440/1832016. Metadata files include data descriptions (_dd.csv) for tabular data. All included files are listed and described in xxxx_flmd.csv.This dataset is an updated version of a previous archive (Wang et al., 2024b, DOI: 10.15485/2475020), incorporating multiple seasons and improved snow depth estimation. Please note that due to large amount of information present in this dataset, many specificities associated with the acquisition of snow temperature, air temperature proxy and estimation of snow depth, and the future archiving of additional datasets on the soil temperature, thaw depth and soil characteristics at these locations, the author would welcome being contacted by people planning to use this dataset.The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

TEMPEST3 surface runoff water chemistry and organic matter composition

Coastal flooding, driven by storm surges and sea level rise, can mobilize organic matter (OM) via runoff, while introducing compositionally distinct OM (e.g., estuarine OM) into the system. To understand event-scale OM dynamics, we monitored source waters and surface runoff during an ecosystem-scale field manipulation experiment, TEMPEST (Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments), in June 2024. The TEMPEST experiment is part of the COMPASS-FME (Coastal Observations, Mechanisms, and Predictions Across Systems and Scales – Field, Measurements, and Experiments) project and designed to investigate biogeochemical and ecological impacts of freshwater and seawater flooding on coastal terrestrial-aquatic interface ecosystems by simulating freshwater and seawater storm events in two 2000m2 coastal upland forest plots (freshwater and brackish seawater plots). The temporal coverage of this dataset is during the TEMPESTⅢ event (June 11-13, 2024). This dataset contains: - Surface runoff discharge measured by flumes - Sensor data (specific conductivity, salinity, dissolved oxygen, and temperature) - Particle size distribution - Total suspended sediment concentrations (TSS), particulate and dissolved organic carbon (POC, DOC) concentrations, total nitrogen and total dissolved nitrogen (TN, TDN) concentrations - Bulk particulate and dissolved OM compositions (stable C and N isotopes of particulates and optical measurements of chromophoric dissolved OM) - High resolution mass spectrometry analysis data - Water isotope data All data files are plain-text CSV (comma-separated value), and no special software is required to read them.

COMPASS-FME↗

Data for Stetten et al. (2025), "Biogeochemical controls on iron speciation and cycling across upland to shoreline gradients in freshwater and estuarine coastal soils (Lake Erie and Chesapeake Bay, United States)"

Coastal environments are dynamic interfaces that mediate carbon and nutrient exchanges between terrestrial landscapes and open waters, but it is unclear how biogeochemical reactions, in particular iron (Fe) redox transformations, affect the understanding and prediction of coastal ecosystem functions. This dataset includes measurements from two freshwater sites in the Western and Central basins of Lake Erie (Ohio, United States) and two estuarine sites in the Chesapeake Bay (Maryland, United States); the analytical results were reported by Stetten et al. (2025) in Science of the Total Environment. It was produced as part of the COMPASS-FME project, which seeks to advance a scalable, predictive understanding of the fundamental biogeochemical processes, ecological structure, and ecosystem dynamics that distinguish coastal terrestrial-aquatic interfaces from the purely terrestrial or aquatic systems to which they are coupled. The sites were sampled in November 2022 (CRC), December 2022 (MSM), February 2023 (GCW), and March 2023 (OWC); site codes follow those used by Pennington et al. (2025).The dataset consists of the following soil data:- Solid data (Fe concentration, etc.)- Porewater data (sulfate, sulfide, etc.)- Linear combination fitting results of X-ray absorption near edge structure (XANES) spectra; i.e., quantitative results of the oxidation state of Fe, indicated as a proportion of pure Fe(III) and Fe(II) model compounds- Linear combination fitting results of EXAFS (extended X-ray absorption fine structure) spectra, indicated as proportion of of Fe-model compounds (illite, smectite, etc.)Each data type has a single file in comma-separated value (CSV) format. No special software is required to read it.

54 ENVIRONMENTAL SCIENCES↗

Ground surface temperature derived Snow Cover Properties, Seward Peninsula, Alaska, 2019-2023

Snow-ground interface temperatures have been collected at the Teller mile marker 27 and Kougarok mile marker 64 field sites on the Seward Peninsula, Alaska from 2019 through 2023 (with data missing from Fall 2020 through Summer 2021 due to COVID). Temperatures were measured using iButton Link DS1921G-F5# Thermochron miniature temperature sensors and Tinytag TGP-4017 internal sensors deployed across the Kougarok 64 and Teller 27 field sites. These sensors are a cost-efficient way to collect snow-ground interface temperatures at a high spatial resolution, and when paired with air temperature data these measurements can provide insight into fine-scale variability in snowpack characteristics across the study sites. From this data, snow process metrics were calculated at each sensor location based on the methods outlined in Staub and Delaloye, 2017. Metrics are calculated daily for each sensor as well as over the entire season. These metrics include ground surface temperature (°C), the number of days under snow cover (number of days), the insulation effect of snow (unitless), the length of the transitional snow periods (number of days), as well as intermediaries such as temperature variability. Calculating these snow processes relies on the assumption that when snow covers a temperature sensor, it is buffered from diurnal fluctuations in air temperature by the insulating snow layer. More information on the calculated metrics can be found in the User Guide of this dataset, as well as in Staub and Delaloye’s 2017 publication Using Near-Surface Ground Temperature Data to Derive Snow Insulation and Melt Indices for Mountain Permafrost Applications. This dataset includes one daily and one seasonal *.csv file of metrics for every year of data, a daily and a seasonal *.csv data dictionary, and one User Guide document (*.pdf) describing data collection and processing.The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

COMPASS-FME Synoptic Site Characterization

This dataset contains soil biogeochemical and physicochemical characterization data for the COMPASS-FME synoptic sites.This dataset also contains data for the paper Patel et al. 2025 "Transition zones at the changing coastal terrestrial-aquatic interface", https://doi.org/10.1029/2025JG008978.Coastal soils are a significant but highly uncertain component of global biogeochemical cycles. These systems experience unique spatial and temporal variability in biogeochemical processes, driven by wetland-to-upland gradients and hydrological fluctuations. We studied drivers of coastal soil variability (a) at regional scales and (b) across transects from upland forest to wetland, in two contrasting regions — Lake Erie, a freshwater lacustrine system, and Chesapeake Bay, a saltwater estuarine system. Salinity-related analytes were a key driver of soil variability, not just in the saltwater system, but surprisingly, also in the freshwater system. We had hypothesized linear trends in biogeochemical parameters along the TAI – however, contrary to expectations, transition soils were not consistently intermediate between upland and wetland endmembers; the non-monotonic trends of carbon, phosphorus, iron along our transects suggest that these are key analytes to study in our regions. Rapidly changing soil factors across coastal gradients provide insights into which soil processes may act as precursors to ecosystem shifts. Our comprehensive soil characterization across the coastal transects provides essential data for mechanistic modeling of ecosystem dynamics.The data are provided as processed, csv files. Raw data and processing scripts can be accessed on GitHub (https://github.com/COMPASS-DOE/cmps-soil_characterization).A note on the nomenclature: the experimental design represents three points along the coastal gradient -- upland, transition, and wetland. "wetland" is referred to as "marsh" in the corresponding paper. The two terms can be used interchangeably for the sites in this study.

54 ENVIRONMENTAL SCIENCES↗