Search NASA⌕ Search

SEARCH · Search NASA

Results for “Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

A Compilation of Global Bio-Optical in Situ Data for Ocean Colour Satellite Applications – Version Three

A global in situ data set for validation of ocean colour products from the ESA Ocean Colour Climate Change Initiative (OC-CCI) is presented. This version of the compilation, starting in 1997, now extends to 2021, which is important for the validation of the most recent satellite optical sensors such as Sentinel 3B OLCI and NOAA-20 VIIRS. The data set comprises in situ observations of the following variables: spectral remote-sensing reflectance, concentration of chlorophyll-a, spectral inherent optical properties, spectral diffuse attenuation coefficient, and total suspended matter. Data were obtained from multi-project archives acquired via open internet services or from individual projects acquired directly from data providers. Methodologies were implemented for homogenization, quality control, and merging of all data. Minimal changes were made on the original data, other than conversion to a standard format, elimination of some points, after quality control and averaging of observations that were close in time and space. The result is a merged table available in text format. Overall, the size of the data set grew with 148 432 rows, with each row representing a unique station in space and time (cf. 136 250 rows in previous version; Valente et al., 2019). Observations of remote-sensing reflectance increased to 68 641 (cf. 59 781 in previous version; Valente et al., 2019). There was also a near tenfold increase in chlorophyll data since 2016. Metadata of each in situ measurement (original source, cruise or experiment, principal investigator) are included in the final table. By making the metadata available, provenance is better documented and it is also possible to analyse each set of data separately.

ocean color↗

Open Science for Plants in Space: Data Sharing, Standards, and Informatics for Reuse and Knowledge Discovery

Upcoming deep space missions will rely on plants for crew and ecosystem health. Open access space biology data enables scientists to examine the biological responses of plants to ionizing radiation, altered gravity, low atmospheric pressure, elevated CO2, altered photoperiods and many other abiotic stressors. Open Science is the practice of making research available to all, while respecting diverse cultures, to foster collaborations with equity. NASA has declared 2023 as the ‘Year of Open Science’ and created a 5-year Transform to Open Science (TOPS) initiative designed to rapidly transform the agency toward an inclusive culture of open science. NASA’s Open Science Data Repository (OSDR) within the Biological and Physical Sciences Division provides access to data from space-relevant biological experiments. OSDR combines two databases, GeneLab and Ames Life Sciences Data Archive (ALSDA) to maximize access to standardized ‘omics (e.g., transcriptomics, proteomics) and phenotypic data (e.g., microscopy, biomass), respectively. GeneLab started in 2014 with the creation of the first space-relevant FAIR (Findable, Accessible, Interoperable, Reusable) biological ‘omics repository, providing detailed metadata on investigation, sample, and assay levels. The addition of ALSDA to OSDR expands plant data analysis capabilities across both phenotypic and ‘omics data. Today, OSDR hosts 62+ plant datasets and has enabled 58 peer-reviewed publications. Most of these publications were collaboration efforts under the OSDR Analysis Working Groups (AWGs). AWGs provide great opportunities for investigators to collaborate with community members and set new standards for space-relevant data and metadata. The AWGs welcome any ASGSR members interested in contributing plant expertise for space biology, and to serve as subject matter experts as we establish the framework for modern plant data archiving. Investigators are encouraged to submit their space-relevant plant datasets to OSDR and visit the site to learn about the tools OSDR has to offer (osdr.nasa.gov/bio).

FAIR↗

Capturing, Analyzing, Maintaining, and Disseminating Shape Memory Material Data Between Information Management Systems

With an increased demand on reducing the time, cost, and effort to develop new materials, Integrated Computational Materials Engineering (ICME) has received widespread attention in various engineering disciplines as a catalyst for significantly reducing experimental testing during the material design process. An ICME approach to design can enable ‘fit-for-purpose’ materials to be realized in engineering applications by incorporating well-understood process-property-performance relationships between the various length and time scales in a material’s structure, enabling material optimization. However, such an approach requires validated multiscale models at the various length scales for a material, which in turn requires a large amount of data, a robust means of storing the data, and the ability to link data to developed material models. The NASA Vision 2040 [1] has identified nine key elements to enabling ICME approaches in system level design, with one being “Data, Information, and Visualization”, thus outlining the importance of a robust information management system for ICME. As the relationship between microstructure, properties, and material performance become better understood and incorporated into multiscale models that can be leveraged in application design, the emergence of new materials with application-driven properties can be realized. One such new material class that has seen growing attention are shape memory materials (SMM), in which a material can transition between a deformed and undeformed state via a reversible phase transformation when subject to a thermal, mechanical, or magnetic load [2]. SMMs have been used widely in aerospace and biomedical industries, including applications such as actuators, low-shock mechanisms, medical staples, braces, and stents [3, 4]. These materials exhibit unique behavior due to their ability to transition between phases, and thus the mechanisms that enable this transition must be captured in a data information management system and incorporated into SMM material models. At NASA Glenn Research Center, the Shape Memory Materials Database (SMMD) Tool has been developed to capture the necessary information that governs SMM material behavior and provide users the ability to select and visualize various SMMs for a specific application [5]. The database contains point-wise data for published SMM materials, along with the pedigree metadata for traceability necessary for a robust information management system. The database is also capable of storing in-house test data performed at NASA GRC by interacting with the developed Shape Memory Alloy (SMA) Analytics tool to extract the necessary point-wise values and populate the database. Although the SMMD Tool offers its users a single, authoritative source for SMM material data that is critical for model development and material design, the full material pedigree of the in-house test data for SMMs is not currently captured and is out of the scope for the SMMD tool. In this work, the schema for capturing SMM test data within the larger NASA GRC ICME Schema [6, 7, 8, 9] will be developed and implemented for thermomechanical tests conducted at NASA GRC. The developed schema will not only store the relevant data needed for the SMMD tool, but also the material pedigree (i.e., production of the bulk material, bulk material analysis, sample cut-out diagrams, sample fabrication procedure, etc.), test pedigree (i.e., test equipment used, measurement systems used, raw test data), and analysis pedigree (i.e., how the data in the SMMD tool is calculated). Furthermore, a Python-based framework will be developed to seamlessly interact between the SMA Analytics and SMMD tools, which will write the full dataset and associated metadata to the GRC Information Management System before passing the required point-wise data to the SMMD tool. Data informatics is a key element of the NASA Vision 2040, which requires not only that data is stored and maintained throughout the material lifecycle, but that the data is also accessible and reusable such that material development efforts can be minimized. Therefore, for an ICME design approach to be realized, a centralized information management system that drives the ICME process must be able to communicate with other databases. The work that will be presented in this presentation will therefore not only demonstrate the ability of NASA GRC’s information management system to capture SMM data, but also its ability to interact with pre-existing tools specialized for such materials.

Data management↗

Beyond Fair: Engagement, Data Usability, and Open Community Productivity through the NASA Open Science Data Repository

The FAIR principle (findable, accessible, interoperable, and reusable) governs the storage and sharing of NASA space biology and health data[1]. These guiding principles maximize reuse of data and the reproducibility of scientific findings. The NASA Open Science Data Repository (OSDR; an expansion of NASA GeneLab) was built on the FAIR principles and houses over 500 studies and close to 1000 datasets from decades of space life sciences experiments. OSDR embodies the FAIR principles through data governance that includes mediated, embargoed, and fully open access data. The FAIR data governance principles were recently proposed to be expanded to encompass a FAIREST framework for assessing research data repositories (FAIR + Engagement, Social connections, and Trust)[2]. FAIREST emphasizes the importance of data repositories engaging with the scientific community and gaining the trust of researchers regarding data quality. Trust also refers to the TRUST principles developed for assessment of digital repositories: Transparency, Responsibility, User Focus, Sustainability, Technology[3]. We present the “Open Science for Life in Space” Analysis Working Groups (AWGs) as evidence regarding the power of engagement, social connections, and trust which has enhanced OSDR’s capabilities and productivity. AWG members engage in two main activities. One, members provide feedback on OSDR scientific standards for data ingestion, curation, and reuse (study, subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability). Two, AWG members collaborate to mine-reuse OSDR data to conduct scientific analysis. With nearly 800 active members, the AWGs have resulted in 32 publications re-using OSDR data and contributed many papers in two major special issues in Cell (2020) and Nature (2024). AWGs also serve as networking groups, facilitate social connections between researchers at all levels of experience, and also have a social online ‘Forum’ used to keep members informed on projects and opportunities. This community-centric, productive, and trustworthy data culture has resulted in a broader effect with international space agencies, academics, and the commercial space sector wanting to submit their data to OSDR. Ten studies of Inspiration 4 data were recently publicly released by OSDR, as were some JAXA human data. Coming up soon in OSDR are data submissions from the European Space Agency, Virgin Galactic PIs, and SpaceX Polaris Dawn. A major benefit of OSDR is the array of standardized and uniformly formatted data (which was developed through AWG member consensus), from which visualization tools, analysis tools, and machine learning models can be built or trained. This talk will cover the Multi-Study Visualization Tool, the Environmental Data Application, RadLab, and a UCSF-NSF funded knowledge graph biomedical health discovery tool ‘SPOKE’ currently being integrated with OSDR. OSDR also provides training programs in bioinformatics and machine learning to improve the scientific community’s awareness of data availability and to boost their ability to perform data analysis. The increasing engagement of the scientific community and the public with technologies powered by artificial intelligence (AI) heightens the need for data analysis to be transparent. The AI for Life in Space initiative leverages the data products provided in OSDR to train AI models, with an emphasis on explainable and trustworthy AI, which would not be possible without FAIR data and metadata. Overall, here we will demonstrate the importance for NASA life sciences data repositories to adhere to the FAIREST framework, by providing examples and success stories from different aspects of OSDR.

data↗

Governing Data Findability, Accessibility, Interoperability and Reusability (FAIR) Compliance

The most recent data strategy documents at both the federal and NASA levels stipulate that systems should strive for the data they manage to be Findable, Accessible, Interoperable, and Reusable (FAIR). The NASA Life Sciences Portal (NLSP) has already begun leading efforts in this area for HRP, initiating efforts to comply with the FAIR principles. The broad interpretation of the FAIR principles has led to a plethora of tools that use a splay of metrics specifically but variably developed to judge how compliant data and systems are with the principles. A recent review [3] identified and studied 1,180 metrics across 20 publicly available tools for checking FAIR compliance of data and systems. Because of their very recent development, many organizations and data systems managers and developers have not yet had adequate time or resources to understand these FAIR compliance tools and metrics, their variations in design, accuracy or ease of application to their specific data sets and systems. Thus, it would be best for larger organizations like NASA to approach formulating a strategy for governance of FAIR compliance that can be flexibly applied and is adaptable to an evolving awareness knowledge of FAIR compliance methods and tools. In September 2024, the NASA Science Mission Directorate(SMD) organized a workshop on NASA science data repositories, including the topics of implementing FAIR and governing FAIR compliance across SMD. The initial part of these FAIR discussions focused on developing consensus around required science metadata fields. This is challenging given the diverse nature of NASA’s scientific data portfolio, the variety of metadata models and vocabularies used, and variable level of resources available to curate these data. Later discussion focused on three possible approaches to governing FAIR compliance: distributed, in which various programs, projects or systems define their own methods for assessing FAIR compliance, reporting results up appropriate management lines; centralized, in which higher-level organization(s) specify compliance tools or methods for the various data systems; and multi-level, in which a group comprised of individuals with expertise from multiple levels with organizations is formed to provide guidance and/or specifications for governing FAIR compliance. We report on the recommendations this session yielded, and how these might be shaped specifically to help implement and govern the compliance with FAIR of Human Research Program data and systems.

governance↗

Evaluation of Best Practices in Mitigating Startup Costs on Leadership-Class Supercomputers

Supercomputers at Department of Energy (DOE) National Laboratories face a widening range of workloads, from traditional modeling and simulation to Artificial Intelligence model training or complex multi-stage workflows, and beyond. At DOE Leadership Computing Facilities like the Oak Ridge Leadership Computing Facility (OLCF), these workloads demand concurrent access to large portions of the supercomputer’s resources. Launching a job across massive supercomputers is challenging from the start; the file system struggles with a large backlog of metadata requests as tens of thousands of processes read thousands of the same files, and the compute job cannot start until this is completed. There are multiple existing approaches to calm this metadata storm, ranging from vendor-developed tools like sbcast to National Laboratory-developed tools like Spindle and Copper. In this paper, we benchmark and discuss three common approaches to improving compute job launch latencies on Frontier: Slurm’s sbcast tool, Spindle, and Copper. We evaluate these tools by measuring the launch latencies of four workloads: OSU Microbenchmark’s osu_init, Pynamic, Python import mpi4py, and Python import torch. We provide discussion of the results, highlighting data that meet expectations and that do not meet expectations.

Hagerty, Nick [ORNL] (ORCID:0000000330014414)↗

Time-lapse imagery in 2017 and 2018 at the Lower Montane site in the East River Watershed, Colorado

Time-lapse imagery was collected using an automated RGB camera mounted on a pole at the base of the northeast-facing hillslope at the Lower Montane site in the East River Watershed, Colorado. The imagery was intended to support a better understanding of plant dynamics and their controls during the growing season. The dataset includes RGB images archived in four zip files (containing imagery in JPEG format), corresponding to photos taken from the hillslope and the adjacent floodplain during 2017 and 2018. A fifth zip file contains a few AVI movies that compare imagery between the two years. The AVI files can be read with most media players applications. The archive contains a total of five *.zip files and three csv metadata files (flmd.csv, dd.csv, and locations.csv).This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Processed sap flow and fine-root trait data associated with summer drought responses in temperate trees in Lisle, Illinois, USA (2019–2021)

These data support the manuscript “Acquisitive root exploration strategies help maintain higher peak sap flux rates during summer drought, but more root biomass does not”. The dataset includes processed sap flow measurements and fine-root trait data collected between 2019 and 2021 from temperate monodominant tree plots established in the 1920s to 1930s ranging in size from 0.05 to 0.8 ha at The Morton Arboretum in Lisle, IL. Sap flow was measured with ICT sap flow sensors using the heat ratio method. Fine-root traits were measured from soil cores which includes specific root length (SRL), specific root area (SRA), diameter, biomass, and length for diameter classes ≤1 mm and ≤2 mm. The package contains comma separated value (CSV) data files and associated metadata that can be viewed and analyzed using common software such as spreadsheet programs, R, and Python. These data are used to investigate how variation in fine-root traits relate to tree water use and drought response during summer drought linking belowground root traits and aboveground physiological responses.

drought↗

Hosting downscaled decision-relevant community data products in ESGF2-US

As regionally-relevant high-resolution Earth system data is increasingly relied upon across scientific, policy, and practitioner communities, there is an urgent need for coordinated and federated infrastructure to store, manage, standardize, and distribute decision-relevant community data products. Substantial effort is required to ensure that these products, which are often critical for regional impact assessments and decision-making, are findable, accessible, interoperable, and reusable. The Earth System Grid Federation US project (ESGF2-US) is addressing this challenge by expanding its open-source, distributed platform to support the hosting and dissemination of downscaled Earth system datasets. This expansion includes aligning new downscaled datasets with developing community standards for metadata and file structure, consistent with existing ESGF archives. This includes ensuring CF-compliance, applying CMORization where appropriate, and developing tools to streamline user access. In this paper, we highlight the technical and coordination work required to bring downscaled data into ESGF2-US and aim to inform the broader Earth system data user community about the growing availability and utility of these curated resources.

ESGF↗

Groundwater table elevation and temperature from 2015 to 2024 at the Lower Montane site in the East River Watershed, Colorado.

This groundwater level elevation and temperature data package is aimed at improving the predictive understanding of hydro-biogeochemical processes at the lower montane site in the East River Watershed, Colorado. The dataset is obtained using pressure transducers placed in shallow wells in the floodplain. This dataset contains data from wells with Location ID's ER-DOW (alias DO1West), ER-DOE (alias DO2East), ER-MBA1 (alias M1Bend1), ER-MBA2 (alias M1Bend2), ER-UPW (alias UP1West), ER-UPM (alias UP2), ER-UPE (alias UP3East). Another dataset contains the data from wells with Location ID's ER-CPA1 to ER-CPA6. Each file contains the water level elevation and the water temperature. Water level elevation has been obtained using the barometric pressure from the pressure transducer (Hobos sensor) in the well, barometric pressure from a sensor in air located at the same site (lower montane), depth from top-of-casing (TOC) to sensor measurement point, and TOC elevation. Data have been checked with a few measurements of water table depths. A real-time kinematic (RTK) global positioning system (GPS) has been used to survey the TOC (data in file Well_Location.csv). The water level elevation is given in UTM13N Geoid2012AB. While depth to water level is not present in the data files, it can be easily calculated with the TOC and distance to ground provided in the GPS coordinate file. The dataset quality is discussed in Collection/Analysis section of the methods. Time-series of measurements were initially added to the archive for the period 2015 to 2019, and later updated with time-series until 2024 (end of data collection). The dataset contains 8 *.csv data files, and 3 *.csv metadata files. Feel free to contact the author with any questions or collaboration interests. The publication year was updated from "2020" to "2025" to reflect the revised version of this dataset.

54 ENVIRONMENTAL SCIENCES↗

Groundwater table elevation and temperature from 2015 to 2024 across Meander C at the Lower Montane site in the East River Watershed, Colorado.

This groundwater level elevation and temperature data package is aimed at improving the predictive understanding of hydro-biogeochemical processes at the lower montane site in the East River Watershed, Colorado. The dataset is obtained using pressure transducers placed in shallow wells in the floodplain. This dataset contains data from wells ER-CPA1 to ER-CPA6. Another dataset contains the data from wells at nearby Locations. Each file contains the water level elevation and the water temperature. Water level elevation has been obtained using the barometric pressure from the pressure transducer (Hobos sensor) in the well, barometric pressure from a sensor in air located at the same site (lower montane), depth from top-of-casing (TOC) to sensor measurement point, and TOC elevation. Data have been checked with a few measurements of water table depths. A real-time kinematic (RTK) global positioning system (GPS) has been used to survey the TOC (data in file Well_Location.csv). The water level elevation is given in UTM13N Geoid2012AB. While depth to water level is not present in the data files, it can be easily calculated with the TOC and distance to ground provided in the GPS coordinate file. The dataset quality is discussed in Collection/Analysis section of the methods. Time-series of measurements were initially added to the archive for the period 2015 to 2019, and later updated with time-series until 2024 (end of data collection). The dataset contains 7 *.csv data files, and 3 *.csv metadata files. Feel free to contact the author with any questions or collaboration interests. The publication year was updated from "2020" to "2025" to reflect the revised version of this dataset.

54 ENVIRONMENTAL SCIENCES↗

Electrical Resistivity Tomography data from 2016 to 2018 at the Lower Montane site in the East River Watershed, Colorado

This dataset contains time-lapse Electrical Resistivity Tomography (ERT) data along a transect located on the northeast-facing hillslope at the lower montane site (Pumphouse site) in the upper East River Watershed. The monitoring dataset covers the period from November 2, 2016, to August 6, 2018. In addition, the archive also contains a baseline dataset from October 9, 2016. The ERT transect consisted of 128 electrodes with an electrode spacing of 1.25 m. The acquisition system was located in the middle of the transect, about 50 m on one side, and included an MPT (Multi-Phase Technologies) ERT system, a mini computer, and batteries with solar panels. Acquisition occurred daily under normal circumstances. The first 16 electrodes (from the upper end of the transect) could not be used after the cable was damaged during the 2017–2018 winter. Also, due to multiple failures in the power system, the temporal resolution of the data is much lower in 2018 compared to 2016 and 2017. The data have been processed and used in Dafflon et al., 2023, and the baseline dataset was used in Falco et al., 2019 (see reference list). This archive contains the measurements (ER.zip containing csv files) for each of the 326 acquisition times and a filtered version where only electrodes 17 to 128 are included (ERT_sm.zip containing csv files). The archive also contains the baseline dataset and two acquisitions with full reciprocals (ERT_RB.zip containing csv files), as well as all the raw MPT files (ERT_raw_MTP.zip). The geometry (electrode position and elevation) is provided in Universal Transverse Mercator (UTM) 13N Geoid2012AB in the file named ERT_Location.csv. The archive contains 1 *.csv data files, four *.zip files, and three metadata *.csv files. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Multiple RGB ortho-mosaics and digital surface models in 2017 and 2018 across the Lower Montane site in the East River Watershed, Colorado

Aerial imagery was collected at the Lower Montane site (Pumphouse) in the East River Watershed, Colorado during the spring, summer, and fall seasons of 2017 and 2018 to improve the understanding of seasonal vegetation dynamics and their drivers. The datasets include Red-Green-Blue (RGB) ortho-mosaics and digital surface models (DSMs) inferred from the Unoccupied Aerial System (UAS) acquired aerial RGB imagery for June 3, June 19, July 7, and August 14, 2017, and for March 14, April 26, June 1, June 18, July 6, and August 7, 2018. Real-Time Kinematic Global Positioning System (RTK-GPS) surveyed Ground control points (GCPs) were used to increase the reconstruction accuracy. The reconstructed RGB mosaics and DSMs have been trimmed to cover a similar spatial domain. The accuracy of the RGB mosaics is considered high (~10 cm). DSM accuracy is highest (~10 cm) where sufficient GCPS are available, and more difficult to assess elsewhere (see reconstruction reports for uncertainty estimates). The dataset includes a total of 20 GeoTIFF (.tif) files, 10 PDF (.pdf) files, 3 data CSV (.csv) files, and 2 metadata CSV (.csv) files. Feel free to contact the authors with any questions or collaboration interests.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with the manuscript "Organic Molecules are Deterministically Assembled in River Sediments"

This data package is associated with the publication "Organic Molecules are Deterministically Assembled in River Sediments" submitted to Scientific Reports (Stegen et al., 2024). The study applies community ecology methods to dissolved organic matter (DOM) chemistry from variably inundated riverbed sediments to uncover principles governing DOM composition at a reach-scale. This data package documents the workflow used to process and generate the main findings in the manuscript. The R scripts reference the raw, unprocessed Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data from another data package, available on ESS-DIVE at https://data.ess-dive.lbl.gov/view/doi:10.15485/1834208. The scripts then process the raw FTICR-MS data and generate the findings and figures presented in the associated manuscript. In brief, this study demonstrates that DOM assemblages in variably inundated sediments are primarily governed by deterministic variable selection, including sediment moisture effecting the degree of deterministic assembly. See the manuscript for more details pertaining to interpretation and implications of the findings. This data package is associated with the GitHub repository found at https://github.com/WHONDRS-Hub/ECA_2020_Sed.This data package is comprised of 6 scripts and 7 folders. The file-level metadata file (file ending in "flmd.csv") lists all files contained in this data package and descriptions for each. The data dictionary (file ending in "dd.csv) describes all tabular data columns and their respective definitions and units. The FTICR_Processing_Scripts produce the outputs found in the "Processed_Data" folder. The remaining scripts (located in the parent directory) produce the outputs found in the following four folders: (1) "MCD_Dendrograms", "MCD_Randomizations", "MCD_bNTI_Outcomes", and "OM_Null_Modeling". The fifth script additionally takes the three comma-separated values (CSV) files found in the parent directory as input ("VGC_texture.csv", "merged_weights.csv", and "ECA2_FTICR_BetaDisp.csv"). The outputs of each of the five scripts serve as the input to the following script, with the final outputs stored in the folder "OM_Null_Modeling".

54 ENVIRONMENTAL SCIENCES↗

NGEE Arctic 2019 Alder Ground Truth Survey, Seward Peninsula AK

In July 2019 we made traveled the road system outside of Nome, AK and detailed the GPS coordinates of alder shrublands for the purpose of ground-truthing alder maps of the region. Both visual and ground-based observations were made for patches of alder shrublands greater 5x5m and larger, ideally 10x10m. Visual observations were made from the car and GPS coordinates are approximate, placed by dropping pins on georeferenced pdfs using the Avenza app. Visual observations included positive identified alder shrublands as well as thickets of non-alder shrubs. Ground Observations were made at a subset of locations where we were able to hike to alders shrubland areas. Ground observations include GPS points (made with Garmin InReach) as well as relevant features of a centrally located, representative alder shrub in the patch (max height, basal diameter of all ramets, soil depth). Aboveground biomass (weight dry mass) of the surveyed shrub was calculated based on alder-specific allometric equations in Berner et al 2015 which our team checked for accuracy for the Seward Peninsula as part of Salmon et al 2019. This dataset contains three data files, three data dictionaries, and one file-level metadata file all in*.csv format plus one *.txt README file. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Organic layer thickness and carbon concentration in burned and unburned sites, Seward Peninsula, AK, 2022

Measurements associated with organic layer samples collected from naturally burned (1971, 2002, 2015, 2019) and unburned sites at the Kougarok Fire Complex, Seward Peninsula, AK, 2022. Here, a discontinuous permafrost underlies an arctic tundra ecosystem. Measurements include elemental carbon and nitrogen concentrations and stocks, organic layer thickness, and thaw depth. There are five files in *.csv format with one data file and four data description files including data dictionary, methods, terminology, and file-level metadata. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with “Sequential Precipitation Input Tagging (SPIT) to Estimate Water Transit Times and Hydrologic Tracer Dynamics within Water-Tagging Enabled Hydrologic Models” (v3)

This data package is associated with the publication “Sequential Precipitation Input Tagging (SPIT) to Estimate Water Transit Times and Hydrologic Tracer Dynamics within Water-Tagging Enabled Hydrologic Models” submitted to Journal of Advances in Modeling Earth Systems (Butler et al. 2025). This study developed the Sequential Precipitation Input Tagging (SPIT) framework to tag input precipitation and estimate water transit times and hydrologic tracers. SPIT tags all precipitation events at regular intervals over an extended period (monthly tags over seven years) in a hydrologic model from 2016-2022. SPIT is applied at six National Ecological Observatory Network (NEON) sites across the continental United States to calculate transit time distributions (TTD) and derive from these mean transit times (MTT), fractions of young water (Fyw), and hydrologic tracer concentrations in stream water (δ18O) within a water-tagging enabled version of the Weather Research and Forecast (WT-WRF-Hydro) model with national water model (NWM) configurations. We go on to validate WT-WRF-Hydro estimates against Butler et al. (2023), who analyzed the same NEON sites using stable water isotope data to estimate water transit times. This new tracking method provides a detailed picture of water movement and helps improve predictions about water availability in the future. This data package was originally published in January 2025. It was updated May 2025 (v2; new and modified files) and October 2025 (v3; new and modified files). File and folder names were not revised to indicate changes. See the change history section in the readme for more details. This data package contains the data and scripts used to develop the SPIT framework WT-WRF-Hydro (Water Tagging Weather Research and Forecasting Hydrologic) model and is associated with the following GitHub repository: https://github.com/zbutler33/SPIT-Framework. This data package contains five parent folders: (1) “Manipulated_outputs”, (2) “Metadata”, (3) “Observed”, (4) “Outputs”, and (5) “Scripts”. Each of these parent folders contains additional subfolders and files. Please see the FLMD (“v*_Butler_2024_WT_WRF_Hydro_flmd.csv”) for a list of all the files contained in this data package and descriptions for each. See the data dictionary (“v*_Butler_2024_WT_WRF_Hydro_dd.csv”) for definitions and units of all of the tabular (files ending in “.csv” and ".tsv") column headers.

54 ENVIRONMENTAL SCIENCES↗

Daily water stable isotopes, transpiration, and matrix potential data for an aspen and engelmann stand in the East River Watershed (version 2)

We provide daily stable isotope (2H & 18O) ratios in soil water and xylem (plant stem) water, as well as the sap flow (transpiration) and the soil's matric potential at a forested site near Gothic, Colorado, in the East River catchment. We measured the stable isotopic composition of the transpiration and the daily transpiration flux sum of three aspen and three engelmann spruce. In both forest stands, we installed a soil profile and measured the soil matric potential at 15, 30, and 60 cm depth as well as the stable isotopes of soil pore water at 5, 10, 30, 60, and 90 cm depths. All isotope measurements were done in situ via vapor probes connected to a cavity ring down spectrometer (Picarro L1240i).We further report the daily meteorological data observed at billy barr near our study site. We also provide for each tree the relative share of root water uptake derived from the isotope measurements via a Bayesian mixing model (MixSIAR).The daily data is provided as a time series in "Iso_MP_Sap_DataDaily_ESSDiveUpload.csv" and the units are provided in "dd.csv"; the location of the instrumented trees and soil profiles are given as latitude and longitude coordinates saved as CSV and KMZ files; and a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata.The data was gathered to investigate the short-term changes of the water sources (i.e., variation of root water uptake from different soil depths) of the studied subalpine trees.Update 07/16/2025: The relative and absolute plant water uptake depths were grouped to ensure that the MixSIAR model was applied with endmembers that differed in their d2H value by at least 3 permill and at least 1 permill for d18O. Whenever the difference between observed d2H values for two or more probes at neighboring depths was less than 3 permill, we used the average value for the source water endmember. For days at which probes that were not next to each other measurements did not differ at least 3 permill, the average of all probes between these two depths was used as the water source endmember representing the depth range between these two probes.

54 ENVIRONMENTAL SCIENCES↗