Search NASA⌕ Search

SEARCH · Search NASA

Results for “metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Metagenome-assembled genomes measured at 3 depths during snowmelt period in East River, CO (March, May, and June, September 2017)

Snowmelt is a critical biogeochemical period that accounts for large nitrogen (N) export events from high-elevation watersheds. Soil microbial populations bloom and immobilize N during snowmelt, yet the population size crashes in spring, which releases a pulse of soil N. We sought to discover the N sources fueling this microbial bloom and determine the fate of N following microbial die-off. Here, focusing on the snowmelt period within a headwater catchment of the Upper Colorado River Basin (East River, CO), we deployed strain-resolved metagenomics to identify the metabolic pathways and processes that mobilize soil N during and after snowmelt. Soil metagenome samples were taken from 6 snowpits from 3 depths (0-5cm, 5-15cm, >15cm) at 4 time points during snowmelt period (March 2017, May 2017, and June 2017, September 2017) generating 48 metagenomes. We reconstructed 474 metagenome-assembled genomes (MAGs) across all metagenomes.All 48 metagenomes were sequenced at JGI and raw data can be found under JGI (Joint Genome Institute) GOLD Study Gs0135149. Metagenome assemblies from IMG under the same study were used for genome binning. This dataset (1) a zip file of 474 MAGs (as fasta files, Gs0135149_bins_tar.gz), (2) sample metadata file with sample IGSNs (International Generic Sample Numbers) (samples.csv), (3) bounding box coordinates for the sampled locations (Gs0135149.kml), (4) metagenome metadata file listing IMG/M (Integrated Microbial Genomes/Metagenomes) metagenome accessions linking samples to metagenomes (metagenomes.csv), (5) location metadata file (locations.csv), (6) file-level metadata file (flmd.csv) and (7) data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Metagenome-assembled genomes from topsoils along a hillslope water gradient across early snowmelt to late summer in East River, CO

Drought is changing the American Mountain West at unprecedented rates with unknown consequences to soil microbiome composition and function. As a part of LBNL Watershed Science Focus Area (SFA), we investigated shifts in microbial community and transcriptional activity on a subalpine conifer-meadow transition zone throughout the summer of 2023 as soil dried down. This work took place in Crested Butte, CO on Snodgrass mountain, using a proxy for drought conditions.Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal community at 0-10cm from three sites along a hillslope water gradient across five timepoints from early snowmelt to late summer. 42 metagenomes were sequenced at Joint Genome Institute (JGI) and can be found under the JGI GOLD (Genomes Online Database) sequencing project Gs0166660. Metagenomes were assembled through an inhouse pipeline (see methods), binned using four autobinners (concoct, maxbin2, metabat2, and vamb) and consolidated using dastool. The consolidated bins from all metagenomes were pooled, filtered by completeness (>70%) and contamination (<10%), and dereplicated at 95% ANI using drep. This dataset (1) a zip file of 157 MAGs (as fasta files, Gs0166660_bins_tar.gz), (2) sample metadata file with sample IGSNs (International Generic Sample Numbers) (samples.csv), (3) bounding box coordinates for the sampled locations (Gs0166660.kml), (4) metagenome assembly and coassembly metadata file listing IMG/M (Integrated Microbial Genomes/Metagenomes) metagenome accessions linking samples to metagenomes (EastRiver_Drought_ESSDive_Metadata.csv), (5) location metadata file (locations.csv), (6) file-level metadata file (flmd.csv) and (7) data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Montane Conifer, Aspen, Meadow, and Sagebrush Metagenome Resolved Genomes and Traits in East River Watershed, Colorado, USA

Climate change is driving vegetation shifts in mountain watersheds, with unknown impacts on biogeochemical cycles. We hypothesize that these shifts will reshape soil microbiomes and associated biogeochemical processes. As a part of Lawrence Berkeley National Laboratory (LBNL) Watershed Science Focus Area (SFA), we assessed microbiome and microbial functional trait differences between soils under conifer, aspen, forby meadows, and sagebrush across the East River Watershed, CO, controlling for elevation and aspect.Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from soils 0-20cm in depth across three locations in the watershed—Headwaters, Upper Reaches, and Lower Reaches from August 3-11th 2016. Each location was further subdivided into two blocks, with one block on a west facing aspect, and two on the east aspect of the valley. Within blocks, two samples per vegetation type were taken (one at each depth). This resulted in 66 samples, which were sequenced at JGI and can be found under the Joint Genome Institute (JGI) Genomes Online Database (GOLD) sequencing project Gs0118068. Metagenomes were assembled through an inhouse pipeline (see methods), binned using four autobinners (concoct, maxbin2, metabat2, and vamb) and consolidated using dastool. The consolidated bins from all metagenomes were pooled, filtered by completeness (>75%) and contamination (<25%), and dereplicated at 95% ANI using drep. The dataset includes a zip file of 687 genomes (Vegtype_MAGS.zip), the accession numbers for the underlying metagenomes, a csv file with MAG quality metrics and taxonomy from Genome Taxonomy Database (GTDB) and National Center for Biotechnology Information (NCBI) taxonomic representative genome proteins (EastRiver_Vegtype_drep_genome_info.csv), and a file containing MAG quality metrics and taxonomy (gtdb_drep_bin_taxonomy.csv). The dataset additionally includes a sample metadata file (EastRiver_Vegtype_sample_metadata.csv), a metadata file used to register associated samples with IGSNs (International Generic Sample Numbers) (samples.csv), a Google KML file for the sampled locations (sample_collection_sites.kml), a location metadata file (locations.csv), a file-level metadata file (flmd.csv), and a data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

WHONDRS laboratory time series moisture manipulative experiment from soil core layers across eastern contiguous US: time series aerobic respiration, geochemistry, and aggregates

This dataset supports a broader study examining the effects of wetting and drying on soil layers across the eastern contiguous United States (CONUS). The dataset provides data generated from a laboratory moisture manipulation experiment. The contents include time series aerobic respiration and moisture; dissolved oxygen; sediment geochemistry data; and field metadata. Samples were collected as part of a collaboration between WHONDRS (Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems; https://whondrs.pnnl.gov) and MONet (Molecular Observation Network; https://www.emsl.pnnl.gov/monet). The field samples (soil cores) were labeled as MEL_##_COR and subsequent subsamples begin with MEL_##. Additional subsamples were taken for the laboratory experiment and were labeled as EL_##. The labels from the MEL field samples and the EL subsamples can be mapped directly based on the digits following the prefix and underscore (i.e., EL_01 is a subsample from MEL_01). See the critical details section below for more details on sample naming and experimental design.For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.In addition to this readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions.This dataset is comprised of (1) a folder containing environmental context photos; (2) file-level metadata; (3) data dictionary; (4) field metadata; (5) readme; (6) international generic sample number (IGSN) mapping file; and (7) a subfolder with soil sample data from field samples and the incubation experiment. The sample data subfolder contains (1) effect size; (2) gravimetric moisture from field samples and incubation experiment; (3) respiration rates, raw dissolved oxygen values, and plots; (4) specific conductance, pH, and temperature from the incubation; (5) soil aggregates; (6) a summary containing median values of each data type for each treatment (wet and dry) in the incubation; (7) a summary containing averages for each data type of each soil layer; and (8) methods codes. All files are .csv, .pdf, .jpeg, or .jpg.

54 ENVIRONMENTAL SCIENCES↗

Data for "Depth of nutrient uptake by deep-rooted plants is regulated by water availability"

The data set consists of strontium (Sr) isotope ratios (87Sr/86Sr), water isotopes, soil cation concentrations, soil water potential sensor data, and results of 87Sr/86Sr mixing model. The plant canopy size files include the dataset of canopy dimension of sagebrush, lupine, and sunflower. The soil and plant ICPMS (Inductively Coupled Plasma Mass Spectrometry) data file includes both of 87Sr/86Sr, and cation concentration dataset from soil exchangeable pool, apatite pool, silicate extract, atmospheric rain deposition, and plant leaf and stem tissues. The plant dendrochronology file includes the dendrochronogical ring width of several sagebrush, and dendrochemical sample data includes the 87Sr/86Sr for each separated growth ring. The modeling result gives the proportion of nutrient sources of each plants (based on their 87Sr/86Sr in leaf tissues and growth rings) from atmospheric deposition and mineral weathering. Soil water potential data includes continuous collection of soil water potential dataset at 2 depths (30 cm and 60 cm, from Nov 24 - Jun 25) of the sampling site. All the samples were collected from 2 sampling campaign June and July 2023, and rain water is a separate sampling from Aug - Sept 2023, at north-facing hillslope near pumphouse site. The data showed that the depth of cation nutrient acquisition is thus tightly coupled with, and likely determined by, water availability in soil, saprolite and bedrock. The enhanced uptake of cations and water from regions of mineral weathering could confer plant and ecosystem resilience during low water years and may impact the rate of bedrock weathering and watershed chemistry during drought. This dataset includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type; a location metadata file (locations.csv); and a samples metadata file (samples.csv). All files are provided as comma-separated values (CSV) files (.csv). This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Leveraging Pre-Built Catalogs and Object-Level Scheduling to Eliminate I/O Bottlenecks in HPC Environments

Modern High-Performance Computing (HPC) environments face mounting challenges due to the shift from large to small file datasets, along with an increasing number of users and parallelized applications. As HPC systems rely on Parallel File Systems (PFS), such as Lustre for data processing, performance bottlenecks stemming from Object Storage Target (OST) contention have become a significant concern. Existing solutions, such as LADS with its object-level scheduling approach, fall short in large-scale HPC environments due to their inability to effectively address metadata I/O bottlenecks and the growing number of I/O processes. This study highlights the pressing need for a comprehensive solution that tackles both OST contention and metadata I/O challenges in diverse HPC workloads. To address these challenges, we propose SwiftLoad, an object-level I/O scheduling framework that leverages a metadata catalog to enhance the performance and efficiency of parallel HPC utilities. The adoption of the metadata catalog mitigates the metadata I/O bottlenecks that commonly occur in HPC utilities, a challenge that is particularly pronounced in object-level I/O scheduling. SwiftLoad addresses OST contention and the uneven distribution of I/O processes across different OSTs through mathematical modeling and incorporates a Loader Configuration Module to regulate the number of I/O processes. Evaluated with two representative utilities—data deduplication profiling and data augmentation—SwiftLoad achieved performance improvements of up to 5.63x and 11.0x, respectively, on a production supercomputer.

HPC↗

Videos, photos, and AI-derived grain size data associated with “High-throughput AI Video Surveys Enable Reproducible Multiscale Sediment Size Mapping, with Implications for Hydrobiogeochemical Parameterization”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “High-throughput AI Video Surveys Enable Reproducible Multiscale Sediment Size Mapping, with Implications for Hydrobiogeochemical Parameterization” under review. This data package includes five data types: 1) raw photos and videos from drone survey and walking smartphone surveys; 2) images derived from raw videos; 3) manual labeling of reference scales; 4) metadata for all images and photo resolution derived from artificial intelligence (AI) models or manual labels, 5) grain size data obtained from AI models for all photos, 6) metadata and grain size data after quality control, 7) summaries of sample efficiency for all data, and 8) computational fluid dynamics (CFD) data used to support hydro-biogeochemical (HBGC) parameter estimation. Such data is used to 1) demonstrate significant improvements in accuracy, efficiency, and quality control for grain size data collection with the help of AI models, 2) study the spatial heterogeneity of grain size and observation reproducibility based on tens of thousands of data points generated by the AI models, and 3) evaluate the impacts of grain size heterogeneity on key HBGC parameters across sediment-to-reach and hourly-to-yearly scales. In particular, the data package contains 116 folders and 179696 files. The files include 41 videos in .mov format, 64047 photos in .jpg format, 13541 video-derived photos in .png format, 12747 segmentation mask data in .tif format, 12747 segmentation data in .json format, 24771 .csv files that with metadata and grain size for each individual photo as well as water depth and velocity data from CFD and observation, 51791 .txt files of raw AI predicted labels, and 11 flight record data in .srt format. The summary for all metadata and grain size statistics information is included in “Scales_V3_NG.csv” and “Statistics_V3_NG.csv”. The summary for data that pass data quality control (QC) level 0-2 is included in “QCStatistics_V3_NG.csv”. The QC level 0 represents photos whose photo resolution is positive, excluding photos that miss reference scale. The QC level 1 means reference scale circularity uncertainty is less than 5% for smartphone images while representing photo resolution is larger than 0.44 mm/pixel for drone images. The QC level 2 means excluding photos whose grain number is less than 100, a minimum number of grains recommended by classic literature. The summary for each video’s name, length, frame rates, survey area, grain number, survey efficiency, etc. can be found in “QCSummary_V3_NG.csv”. The summary for site name, GPS coordinates, and number of images at each site can be found in “SitesSummary_V3_*.csv” files. Overall computational efficiency summary is reported in Table 4 of accompanying manuscript. Additionally, the nitrate concentration data used in this work was downloaded from an existing dataset published on ESS-DIVE (Boat-Dragged Sensor Hanford Reach.csv; Conner A. et al., 2020). We thank the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, Cowiche Canyon Conservatory, Port of Benton, and the Confederated Tribes and Bands of the Yakama Nation for access to field locations where the data were collected. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate data collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with “Point-scale organic-matter decomposition in streambeds is weakly associated with reach-scale respiration”

This data package is associated with “Point-scale organic-matter decomposition in streambeds is weakly associated with reach-scale respiration” published in EGU Biogeosciences (Stegen et al., 2026; https://doi.org/10.5194/bg-23-3981-2026). It contains cotton strip decomposition rates (Kcd and Kdd) collected across the Yakima River Basin (YRB), Washington, USA. These data were collected to support a broader study examining the drivers of spatial variability in sediment respiration rates in the Yakima River Basin. Associated data used in analysis, metadata, and field protocols can be accessed at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1923689, https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1969566, and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1987520. This data package is associated with the repository found at https://github.com/river-corridors-sfa/rcsfa-ST-2B-SSS-cotton-strip. A preliminary version of this data package was published in December 2025 at the time of manuscript submission. It was updated in June 2026, at the time of manuscript acceptance, to include additional metadata (this readme, data dictionary, and file level metadata). The data did not change. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This data package consists of (1) readme; (2) data dictionary (dd); (3) file level metadata (flmd); and (4) four folders: (1) R-scripts; (2) figures; (3) outputs from the scripts; and (4) published data. The published data folder contains a readme directing the user to download data in order to run the R-scripts. All files are .csv, .pdf, .R, .Rmd, and .txt. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

The Challenges of Interoperable Data Discovery

The Global Change Master Directory (GCMD) assists the oceanographic community in data discovery and access through its online metadata directory. The directory also offers data holders a means to post and search their oceanographic data through the GCMD portals, i.e. online customized subset metadata directories. The Gulf of Maine Ocean Data Partnership (GoMODP) has expressed interest in using the GCMD portals to increase the visibility of their data holding throughout the Gulf of Maine region and beyond. The purpose of the Gulf of Maine Ocean Data Partnership (GoMODP) is to "promote and coordinate the sharing, linking, electronic dissemination, and use of data on the Gulf of Maine region". The participants have decided that a "coordinated effort is needed to enable users throughout the Gulf of Maine region and beyond to discover and put to use the vast and growing quantities of data in their respective databases". GoMODP members have invited the GCMD to discuss further collaborations in view of this effort. This presentation. will focus on the GCMD GoMODP Portal - demonstrating its content and use for data discovery, and will discuss the challenges of interoperable data discovery. interoperability among metadata standards and vocabularies will be discussed. A short overview of the lessons learned at the Marine Metadata Interoperability (MMI) metadata workshop held in Boulder, Colorado on August 9-11, 2005 will be given.

Meaux, Melanie F.↗

Report on the Global Data Assembly Center (GDAC) to the 12th GHRSST Science Team Meeting

In 2010/2011 the Global Data Assembly Center (GDAC) at NASA's Physical Oceanography Distributed Active Archive Center (PO.DAAC) continued its role as the primary clearinghouse and access node for operational Group for High Resolution Sea Surface Temperature (GHRSST) datastreams, as well as its collaborative role with the NOAA Long Term Stewardship and Reanalysis Facility (LTSRF) for archiving. Here we report on our data management activities and infrastructure improvements since the last science team meeting in June 2010.These include the implementation of all GHRSST datastreams in the new PO.DAAC Data Management and Archive System (DMAS) for more reliable and timely data access. GHRSST dataset metadata are now stored in a new database that has made the maintenance and quality improvement of metadata fields more straightforward. A content management system for a revised suite of PO.DAAC web pages allows dynamic access to a subset of these metadata fields for enhanced dataset description as well as discovery through a faceted search mechanism from the perspective of the user. From the discovery and metadata standpoint the GDAC has also implemented the NASA version of the OpenSearch protocol for searching for GHRSST granules and developed a web service to generate ISO 19115-2 compliant metadata records. Furthermore, the GDAC has continued to implement a new suite of tools and services for GHRSST datastreams including a Level 2 subsetter known as Dataminer, a revised POET Level 3/4 subsetter and visualization tool, a Google Earth interface to selected daily global Level 2 and Level 4 data, and experimented with a THREDDS catalog of GHRSST data collections. Finally we will summarize the expanding user and data statistics, and other metrics that we have collected over the last year demonstrating the broad user community and applications that the GHRSST project continues to serve via the GDAC distribution mechanisms. This report also serves by extension to summarize the activities of the GHRSST Data Assembly and Systems Technical Advisory Group (DAS-TAG).

sea surface temperature (SST)↗

Sally Ride EarthKAM - Automated Image Geo-Referencing Using Google Earth Web Plug-In

Sally Ride EarthKAM is an educational program funded by NASA that aims to provide the public the ability to picture Earth from the perspective of the International Space Station (ISS). A computer-controlled camera is mounted on the ISS in a nadir-pointing window; however, timing limitations in the system cause inaccurate positional metadata. Manually correcting images within an orbit allows the positional metadata to be improved using mathematical regressions. The manual correction process is time-consuming and thus, unfeasible for a large number of images. The standard Google Earth program allows for the importing of KML (keyhole markup language) files that previously were created. These KML file-based overlays could then be manually manipulated as image overlays, saved, and then uploaded to the project server where they are parsed and the metadata in the database is updated. The new interface eliminates the need to save, download, open, re-save, and upload the KML files. Everything is processed on the Web, and all manipulations go directly into the database. Administrators also have the control to discard any single correction that was made and validate a correction. This program streamlines a process that previously required several critical steps and was probably too complex for the average user to complete successfully. The new process is theoretically simple enough for members of the public to make use of and contribute to the success of the Sally Ride EarthKAM project. Using the Google Earth Web plug-in, EarthKAM images, and associated metadata, this software allows users to interactively manipulate an EarthKAM image overlay, and update and improve the associated metadata. The Web interface uses the Google Earth JavaScript API along with PHP-PostgreSQL to present the user the same interface capabilities without leaving the Web. The simpler graphical user interface will allow the public to participate directly and meaningfully with EarthKAM. The use of similar techniques is being investigated to place ground-based observations in a Google Mars environment, allowing the MSL (Mars Science Laboratory) Science Team a means to visualize the rover and its environment.

Andres, Paul M.↗

Enabling Space Biological Knowledge Discovery Through Image and Video Data Sharing

Increased biomedical risks and challenges associated with deep space missions and experiments (cis-Lunar, Mars transit/surface) require new knowledge discovery and development of novel ecosystems. Supporting distant and long-duration missions and experiments requires biological data (from yeast, microbes, fruit flies, C. elegans, plants, crops, rodents, humans) be findable, accessible, interoperable, reusable (FAIR), and maximally open-access. As data-intensive, bioinformatic, meta-analytical, and computer-assisted approaches continue to be a centerpiece of modern research, the NASA Biological and Physical Sciences division is expanding its Open Science capabilities beyond NASA GeneLab. The NASA Ames Life Sciences Data Archive (ALSDA) is a repository which is responsible for collecting and access to space biological imagery and video, alongside tabular and environmental data. In this presentation, we will discuss strategies dealing with archiving, curating, and accessibility of images from very distinct imaging modalities (e.g., micro-computed tomography, magnetic resonance imaging, photographic images of plants, fluorescence microscopy, behavioral videos, etc.). There are two main challenges: 1. Open-source data storage and 2. Metadata related to the imagery-video. Both have been solved by leveraging two existing open-source systems. For data storage, ALSDA is utilizing components through the Open Microscopy Environment (OME), which can read most imaging proprietary formats and display on a web interface complex multidimensional images (Z stack, multi-channel, temporal, spectral). Most technical metadata from imaging modalities are captured seamlessly. For metadata capturing experimental details, ALSDA (like GeneLab) uses the ISA-Tab specification which relies on the ISA data model to order and classify metadata. The ISA data model uses a tree structure with three files to capture the metadata: The top layer is the Investigations file, the second layer is the Study file(s), and the last layer is the Assay file(s). We believe such an approach may be useful for other types of image research data from other investigators in the AGU community.

imaging↗

ICARTT File Format Enhancements: Supporting FAIRness and Data Discovery of Suborbital Campaign Data

Suborbital campaigns aim to accomplish a wide variety of goals and can include a variety of platforms, instruments, and parameters measured. In 2004, the ICARTT (International Consortium for Atmospheric Research on Transport and Transformation) standards were developed to fulfill data management needs for the ICARTT campaign. The ICARTT file format is text-based and composed of a header with important data description information and the data section. Built on the NASA Ames and GTE data formats, the ICARTT format was created to facilitate data exchange and promote collaborations among the science teams for achieving the ICARTT campaign goals. Due to its success and adaptation for use in many other field campaigns, the ICARTT file format became a NASA standard in 2010 and was amended in January 2017. These changes provided many enhancements, including the requirement for variable standard names. Primarily designed for airborne field studies, ICARTT has been further utilized for ground-based studies. NASA has made a commitment to build an inclusive open science community over the next decade. Open-source science strives to make publicly funded scientific research transparent, inclusive, accessible, and reproducible. The ICARTT format can host metadata that is critical for proper use of the data, particularly for in-situ measurements, and can enhance data discovery and accessibility. However, the required fields are often free text, meaning that the information is human readable, but not machine interpretable. Furthermore, the amount and type of information provided can vary significantly between principal investigators and campaigns. To support FAIR principles and interoperability, enhancements to the ICARTT standards are recommended. Possible recommendations include potential use of controlled and consistent vocabulary for variable standard name and certain common metadata elements; standardizing timestamps for easier data comparisons and analysis; and providing guidance on variable measurement units and how they are reported. Enhancing ICARTT metadata can further streamline the process to make suborbital data more readily available to the data user and improve variable-level metadata. Providing more variable-level metadata can enhance data searching and discovery, supporting NASA’s Open-Source Science Initiative (OSSI).

Megan Buzanowicz↗

Enabling Model Organism and Commercial Astronaut Data Access Through the NASA Open Science Data Repository

NASA’s Open Science Data Repository (OSDR) brings together omics data from NASA’s GeneLab project and non-omics data, including physiological, phenotypic, imaging, and behavioral data from NASA’s Ames Life Sciences Data Archive (ALSDA) collected from decades of space biology research, providing open and FAIR (findable, accessible, interoperable, and reusable) access of these precious data to scientists world-wide. This rich source of meticulously curated metadata and data from spaceflight and analog studies has been mined by the scientific community resulting in dozens of high impact scientific publications that reveals a complex network of molecular and physiological effects of spaceflight across living systems, from microbes to plants, to mammals. Understanding how these effects translate to the human condition is critical as we move deeper into the era of commercial space travel. However, the integration of data, specifically omics data, from astronauts is particularly challenging due to their sensitive nature. OSDR has risen to this challenge by developing a mechanism to control access to identifiable levels of omics data, such as raw sequence data, while enabling public access to processed, unidentifiable, data and associated metadata that will allow the scientific community to interrogate human astronaut data alongside data from model organisms to begin answering these critical questions. The 2021 SpaceX Inspiration4 (I4) mission collected a comprehensive atlas of biological measurements from four civilian astronauts, providing a wealth of data to characterize the effects of spaceflight on the human body. These data include both non-omics and omics assays such as direct RNA sequencing (RNA-seq), single nuclei ATAC-seq and RNA-seq, metagenomics, proteomics, and comprehensive metabolic and cytokine panels, all of which have been integrated into the OSDR system across no less than 9 studies. Each study has been carefully curated using community-backed OSDR standards for sample and assay level metadata ensuring these data are findable and accessible. In addition to hosting both raw and processed data from the principal investigator team for each assay type, the GeneLab team plans to re-process the I4 omics data using GeneLab’s standard processing pipelines. The GeneLab processed data outputs will allow for comparisons across studies on OSDR and enable visualization of these data through the OSDR data visualization platform thereby enabling data reusability and interoperability. Here we describe the robust privacy and security protocols implemented by OSDR to safeguard sensitive health data from astronauts while facilitating metadata and processed data sharing for research purposes. We further provide a road map for navigating the vast amount of data provided for each I4 study on the OSDR, including experimental design, associated experiments, payloads, and missions, data generation and analysis protocols, and associated scientific articles. Additionally, we illustrate how to interrogate the standardized metadata provided in the sample and assay tables as well as various means to download and access the data including programmatically through the GeneLab Open API (GLOpenAPI). The open access of datasets in NASA’s OSDR provides a unique opportunity for the scientific community, as well as citizen scientists and students, to continue using OSDR resources to further unlock profound insights into the consequences of space travel on the human body. Through implementation of security measures to protect sensitive human data, the OSDR seeks to strengthen the science exchange between the Biological and Physical Sciences Program and the Human Research Program, per recommendation 4-1 of the 2023-2032 Decadal Survey, and encourage further sharing and dissemination of astronaut data to provide the scientific community with the resources needed to lay the groundwork for developing targeted mitigation strategies to help withstand the rigors of long-duration spaceflight.

Amanda Marie Saravia-butler↗

Enabling Model Organism and Commercial Astronaut Data Access Through the NASA Open Science Data Repository

NASA’s Open Science Data Repository (OSDR) brings together omics data from NASA’s GeneLab project and non-omics data, including physiological, phenotypic, imaging, and behavioral data from NASA’s Ames Life Sciences Data Archive (ALSDA) collected from decades of space biology research, providing open and FAIR (findable, accessible, interoperable, and reusable) access of these precious data to scientists world-wide. This rich source of meticulously curated metadata and data from spaceflight and analog studies has been mined by the scientific community resulting in dozens of high impact scientific publications that reveals a complex network of molecular and physiological effects of spaceflight across living systems, from microbes to plants, to mammals. Understanding how these effects translate to the human condition is critical as we move deeper into the era of commercial space travel. However, the integration of data, specifically omics data, from astronauts is particularly challenging due to their sensitive nature. OSDR has risen to this challenge by developing a mechanism to control access to identifiable levels of omics data, such as raw sequence data, while enabling public access to processed, unidentifiable, data and associated metadata that will allow the scientific community to interrogate human astronaut data alongside data from model organisms to begin answering these critical questions. The 2021 SpaceX Inspiration4 (I4) mission collected a comprehensive atlas of biological measurements from four civilian astronauts, providing a wealth of data to characterize the effects of spaceflight on the human body. These data include both non-omics and omics assays such as direct RNA sequencing (RNA-seq), single nuclei ATAC-seq and RNA-seq, metagenomics, proteomics, and comprehensive metabolic and cytokine panels, all of which have been integrated into the OSDR system across no less than 9 studies. Each study has been carefully curated using community-backed OSDR standards for sample and assay level metadata ensuring these data are findable and accessible. In addition to hosting both raw and processed data from the principal investigator team for each assay type, the GeneLab team plans to re-process the I4 omics data using GeneLab’s standard processing pipelines. The GeneLab processed data outputs will allow for comparisons across studies on OSDR and enable visualization of these data through the OSDR data visualization platform thereby enabling data reusability and interoperability. Here we describe the robust privacy and security protocols implemented by OSDR to safeguard sensitive health data from astronauts while facilitating metadata and processed data sharing for research purposes. We further provide a road map for navigating the vast amount of data provided for each I4 study on the OSDR, including experimental design, associated experiments, payloads, and missions, data generation and analysis protocols, and associated scientific articles. Additionally, we illustrate how to interrogate the standardized metadata provided in the sample and assay tables as well as instructions for how to download and access the data. The I4 datasets described here re present the first ever comprehensive collection of commercial astronaut data.

Amanda M Saravia-Butler↗

Temporal Study 2022-2024: Sensor-Based Time Series of Surface Water Temperature, Specific Conductance, Total Dissolved Solids, Turbidity, Chlorophyll A, and Dissolved Oxygen from across Multiple Watersheds in the Yakima River Basin in Washington, USA

This dataset supports a broader study examining the drivers of temporal variability in sediment respiration rates in the Yakima River Basin. The dataset provides periodic (bi-weekly or monthly) in situ hydrological and water chemistry sensor data, handheld sensor water chemistry data, general environmental context photos, and field metadata collected at six sites across the Yakima River Basin in Washington, USA. Sample and sensor data from previous years (2021-2022) can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1898912 and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1892054, respectively. Related sample data from 2022-2024 are available at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2562910. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions This dataset contains a folder of environmental context photographs and videos and (1) file-level metadata; (2) data dictionary; (3) readme; (4) field metadata; (5) field protocols; (6) international generic sample number (IGSN) mapping file; (7) handheld sensor data; and (8) two sensor subfolders. Each sensor subfolder (BarotrollAtm and MantaRiverData) contains a subfolder containing sensor time series data and plots. The BarotrollAtm Data subfolder contains In Situ Rugged BaroTROLL sensor pressure and air temperature data. The MantaRiverData subfolder contains Eureka Manta+ 35B multisonde temperature, specific conductance, and chlorophyll A. All files are .csv, .pdf, .jpg, .jpeg, .mp4, .png, or .mov.

54 ENVIRONMENTAL SCIENCES↗

Terrestrial laser scanning data (Levels 0 and 1) for Pasoh, Malaysia, Sep 2024

This data package contains data from terrestrial laser scanning (TLS) at the Pasoh Forest Reserve, Malaysia. The Pasoh Forest Reserve is a facility of the Forest Research Institute Malaysia, and contains evergreen lowland dipterocarp forest. The Next-Generation Ecosystem Experiments Tropics (NGEE-Tropics) study areas at Pasoh were established to study how different species respond to climatic variation and soil water availability. Two study areas were chosen representing different topography and species. The TLS data archived here were collected to provide detailed, three-dimensional information about forest structure. Specifically, data were collected to allow tree-level characterization of woody structure and leaf area for 12 focal trees with FloraPulse and sap flux sensors, facilitating estimation of woody biomass and leaf area to allow upscaling of water content and transpiration data to the tree-level. Scan positions were not selected to provide consistent data for non-focal trees with the study areas. This data package contains the following data: - High-level files document further details of the campaign and data package: 1_CampaignSummary.csv provides details about the campaign and study site, 2_ScanAreasDetail.csv provides details about each separate scan area (groups of scans post-processed into a single point cloud), 3_TerrestrialLidarSensor.csv provides further technical details about the Riegl VZ-400i TLS sensor, TLS_CSV_dd.csv is a CSV Data Dictionary providing information about the fields in CSV files following the ESS-DIVE CSV File Formatting Guidelines Reporting Format, TLS_flmd.csv is a File Level Metadata file providing information about each file in the data package following the ESS-DIVE File Level Metadata Reporting Format, and README.txt is a text file describing the overall project and file structure. - Level 0 data are the raw data (.PROJ folders) as recorded by the Riegl VZ-400i TLS instrument before scan co-registration and post-processing with the Riegl's proprietary RiSCAN PRO software, which requires a license. - Level 1 data contain post-processed, co-registered data from each scan area. The "PointClouds" folder for each scan area contains a .las file with 1 cm resolution point cloud data exported from RiSCAN PRO. These are the main files likely to be of interest to most users and can be further processed with any software capable of manipulating .las files (e.g. Python, R CloudCompare). The "Project Information" folder contains log files from post-processing in RiSCAN PRO that may be of interest to users who want to see detailed records of post-processing, including all PDF reports generated by RiSCAN PRO. The "ScanPositions" folder contains information about the final position of all TLS scans, after post-processing, in multiple formats. The file ScanPositions_*.csv provides final geo-referenced scan positions, and the file SOP_backup_*.csv can be used in RiSCAN PRO to restore the co-registered scan positions if users wish to re-process raw data (Level 0 .PROJ folders) with RiSCAN PRO software (e.g., subsample to a different resolution, exclude a certain scan position, or apply different filters on reflectance or deviation values) without redoing time-consuming co-registration steps.

54 ENVIRONMENTAL SCIENCES↗

Terrestrial laser scanning data (Levels 0 and 1) from Urban Biogeochemistry Pilot Project sites, Knoxville, Tennessee, Jul 2024 - Jul 2025

This data package contains data from terrestrial laser scanning (TLS) at five urban park sites in Knoxville, Tennessee, USA. All parks include open-grown and/or closed-canopy trees and mixed nearby land use. These study sites were established as part of the Urban Biogeochemistry Pilot Project, which has an overall goal of better understanding how hydrobiogeochemical cycling is altered within the human environment. These five sites represent a gradient of urbanization, and were instrumented to understand hydrological and biogeochemical cycling (e.g., soil moisture, soil physical properties and biogeochemistry, tree transpiration, species type). The TLS data archived here were collected to provide detailed, three-dimensional information about forest structure. Specifically, data were collected to allow tree- and stand-level characterization of woody structure and leaf area. TLS scans were placed to capture the area around trees with sap flow sensors, and as much of a 50 m radius area around the meteorological station as possible given site property limits. Derived products will allow upscaling of water content and transpiration data. This data package contains the following data: - High-level files document further details of the campaign and data package: 1_CampaignSummary.csv provides details about the campaign and study site, 2_ScanAreasDetail.csv provides details about each separate scan area (groups of scans post-processed into a single point cloud), 3_TerrestrialLidarSensor.csv provides further technical details about the Riegl VZ-400i TLS sensor, TLS_CSV_dd.csv is a CSV Data Dictionary providing information about the fields in CSV files following the ESS-DIVE CSV File Formatting Guidelines Reporting Format, TLS_flmd.csv is a File Level Metadata file providing information about each file in the data package following the ESS-DIVE File Level Metadata Reporting Format, and README.txt is a text file describing the overall project and file structure. - Level 0 data are the raw data (.PROJ folders) as recorded by the Riegl VZ-400i TLS instrument before scan co-registration and post-processing with the Riegl's proprietary RiSCAN PRO software, which requires a license. - Level 1 data contain post-processed, co-registered data from each scan area. The "PointClouds" folder for each scan area contains a .las file with 1 cm resolution point cloud data exported from RiSCAN PRO. These are the main files likely to be of interest to most users and can be further processed with any software capable of manipulating .las files (e.g. Python, R CloudCompare). The "Project Information" folder contains log files from post-processing in RiSCAN PRO that may be of interest to users who want to see detailed records of post-processing, including all PDF reports generated by RiSCAN PRO. The "ScanPositions" folder contains information about the final position of all TLS scans, after post-processing, in multiple formats. The file ScanPositions_*.csv provides final geo-referenced scan positions, and the file SOP_backup_*.csv can be used in RiSCAN PRO to restore the co-registered scan positions if users wish to re-process raw data (Level 0 .PROJ folders) with RiSCAN PRO software (e.g., subsample to a different resolution, exclude a certain scan position, or apply different filters on reflectance or deviation values) without redoing time-consuming co-registration steps.

54 ENVIRONMENTAL SCIENCES↗