Search NASA⌕ Search

SEARCH · Search NASA

Results for “Science Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Walker Branch Watershed: Daily Stream Metabolism and Organic Carbon Spiraling Metrics in the West Fork of Walker Branch, Tennessee, USA, 2004-2010

This dataset contains daily metabolism estimates of gross primary production (GPP), ecosystem respiration (ER), and net ecosystem production (NEP), in addition to organic carbon spiraling length (SOC) and mineralization velocity (VfOC) estimates at the West Fork of Walker Branch, a small headwater stream, in the Walker Branch Watershed, Tennessee, USA. Observations were made from 2004-2010 (2004-01-01 to 2010-12-31). These data were generated to assess seasonal and interannual variability in metabolism and organic carbon spiraling and to explore potential driver variables, as analyses of intra- and interannual variability in metabolism and organic carbon spiraling are currently limited, leaving knowledge gaps in the driving mechanisms of and future changes to stream metabolism and carbon processing under climate change. Additionally, measurements of discharge (Q), stream width, stream- and canopy-level photosynthetically active radiation (PAR), water temperature, and precipitation from this time frame are included. This dataset contains one data file in comma separated (*.csv) format.

54 ENVIRONMENTAL SCIENCES↗

Quality-Controlled Meteorological Data from the Flood Control District of Maricopa County (FCDMC) Network, Phoenix, Arizona (1987-2024)

This dataset contains 15- or 30-minute interval meteorological data from the Flood Control District of Maricopa County (FCDMC), Arizona, USA, covering eight key variables across multiple sensor stations between 1987 and 2024. Each variable is stored as a separate CSV file, containing time-series data that have undergone rigorous quality control (QC) procedures and, where appropriate, short-gap interpolation for consistency. The quality control (QC) pipeline consisted of four sequential tests: (1) a range test to ensure all values fall within physically realistic limits, (2) a step test to identify abrupt and implausible changes between consecutive records, (3) a proximity test that validates flagged values from step test using data from nearby stations and exceedance probability thresholds, and (4) a persistence test to detect and remove periods of unrealistically constant readings. These thresholds were calibrated to Arizona’s environmental conditions and sensor specifications. After QC, short gaps (≤2 hours) were linearly interpolated to ensure consistent temporal resolution, except for wind variables. Due to a major upgrade in FCDMC’s data transmission system, only ALERT-2 protocol data (2016–2024) for wind variables are included; earlier ALERT-1 data were excluded because of irregular sampling and high missing rates. This dataset supports regional climate and infrastructure resilience studies by providing standardized, high-resolution meteorological data for the greater Phoenix metropolitan area.

54 ENVIRONMENTAL SCIENCES↗

WHONDRS laboratory time series moisture manipulative experiment from soil core layers across eastern contiguous US: time series aerobic respiration, geochemistry, and aggregates

This dataset supports a broader study examining the effects of wetting and drying on soil layers across the eastern contiguous United States (CONUS). The dataset provides data generated from a laboratory moisture manipulation experiment. The contents include time series aerobic respiration and moisture; dissolved oxygen; sediment geochemistry data; and field metadata. Samples were collected as part of a collaboration between WHONDRS (Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems; https://whondrs.pnnl.gov) and MONet (Molecular Observation Network; https://www.emsl.pnnl.gov/monet). The field samples (soil cores) were labeled as MEL_##_COR and subsequent subsamples begin with MEL_##. Additional subsamples were taken for the laboratory experiment and were labeled as EL_##. The labels from the MEL field samples and the EL subsamples can be mapped directly based on the digits following the prefix and underscore (i.e., EL_01 is a subsample from MEL_01). See the critical details section below for more details on sample naming and experimental design.For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.In addition to this readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions.This dataset is comprised of (1) a folder containing environmental context photos; (2) file-level metadata; (3) data dictionary; (4) field metadata; (5) readme; (6) international generic sample number (IGSN) mapping file; and (7) a subfolder with soil sample data from field samples and the incubation experiment. The sample data subfolder contains (1) effect size; (2) gravimetric moisture from field samples and incubation experiment; (3) respiration rates, raw dissolved oxygen values, and plots; (4) specific conductance, pH, and temperature from the incubation; (5) soil aggregates; (6) a summary containing median values of each data type for each treatment (wet and dry) in the incubation; (7) a summary containing averages for each data type of each soil layer; and (8) methods codes. All files are .csv, .pdf, .jpeg, or .jpg.

54 ENVIRONMENTAL SCIENCES↗

Laboratory time series moisture manipulative experiment from sediment across the contiguous US: time series aerobic respiration and geochemistry (v2)

This dataset supports a broader study examining the effects of wetting and drying on hyporheic zone respiration across the contiguous United States (CONUS). The dataset provides data generated from a laboratory moisture manipulation experiment. The contents include time series aerobic respiration and moisture; dissolved oxygen; sediment geochemistry data; and field metadata (including qualitative information on instream and river corridor characteristics). Samples were collected as part of the WHONDRS CONUS-Scale Model-Sample Study (CM). This study was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the CONUS. The data package associated with the CM study is available at https://data.ess-dive.lbl.gov/view/doi:10.15485/1923689. CM sampling began in April 2022 and ended in October 2023. This study uses subsamples from a subset of CM samples collected between June 2022 and June 2023. The original field samples were labeled as CM_###. Subsequent subsamples for this study were labeled as EC_###. The labels from the field samples and the EC subsamples can be mapped directly based on the digits following the prefix and underscore (i.e., EC_001 is a subsample from CM_001). See the critical details section below for more details on sample naming. This data package was originally published in August 2024. It was updated in February 2026 (v2; new and modified files). See the change history section in the readme for more details. For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) field protocol; and a (6) a subfolder with sediment sample data from the incubation experiment. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC); (2) total nitrogen (TN); (3) adenosine triphosphate (ATP); (4) percent carbon and nitrogen; (5) effect size; (6) iron (II); (7) gravimetric moisture; (8) respiration rates and raw dissolved oxygen values; (9) specific conductance; (10) pH; (11) temperature; (12) a summary containing median values of each data type for each treatment (wet and dry); (13) methods codes; (14) FTICR-MS methods; and (15) a subfolder of 9.4 Tesla FTICR-MS data. This folder contains three subfolders, one containing the sediment .xml data files, one containing the sediment CoreMS output files, the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .ref, or .xml.

54 ENVIRONMENTAL SCIENCES↗

Pyrogenic Organic Matter Laboratory Experiment: Aerobic Respiration and Geochemistry from Variably Inundated Stream Sediments (v3)

This dataset supports a broader study examining the effects of variable inundation and pyrogenic organic matter on ecosystem respiration. The dataset provides data generated from a laboratory batch experiment investigating the interaction between variable inundation conditions (wet and dry sediment) and pyrogenic organic matter (burned and unburned treatments). The contents include time series dissolved oxygen, sediment geochemistry data, and field metadata (including qualitative information on instream and river corridor characteristics). This data package was originally published in November 2025. It was updated in April 2026 (v2; new and modified files) and May 2026 (v3; modified files). See the change history section in the readme for more details For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) international generic sample number (IGSN) mapping file; (5) readme; (6) field protocol; (7) sample name metadata; (8) an environmental context picture for the dry and inundated sampling locations; and (9) a subfolder with sample data from the sediment incubation experiment. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC); (2) total nitrogen (TN); (3) gravimetric moisture; (4) partial pressure and production rates of carbon dioxide, methane, and nitrous oxide; (5) field wet sediment mass, dry sediment mass, water mass, and field wet sediment volume in incubation and sediment NPOC/TN vials; (6) methods codes; (7) respiration rates, pH, and temperature from after the incubation, raw time series dissolved oxygen and temperature, and a subfolder containing associated plots and scripts; (8) ions; (9) FTICR-MS methods; and (10) a subfolder of 12 Tesla (12T) FTICR-MS data. This folder contains the CoreMS processed data and three subfolders, one containing the .xml files, one containing the CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .html, .Rmd, .py, .cal, .json, or .jpg.

54 ENVIRONMENTAL SCIENCES↗

WHONDRS River Corridor Sediment and Water Geochemistry and In Situ Sensor Data from 7 Perennial and 7 Intermittent Streams across San Antonio, Texas (v3)

This dataset supports a broader study examining the effects of intermittency on sediment respiration. The dataset provides sediment and surface water geochemistry and in situ sensor data from 7 perennial and 7 intermittent streams in San Antonio, Texas. Each stream/site was visited both in summer during base flow (July-September 2023) and winter during peak flow (January-February 2024). Related data were collected and will be published separately in collaboration with A. Veach. The data package was originally published in April 2025. It was updated in June 2025 (v2; modified and new files) and September 2025 (v3; modified files). See the change history section in the readme for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of two folders of field photos and videos, one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) international generic sample number (IGSN) mapping file; (6) field protocol; (7) a subfolder with sample data; and (8) a subfolder with sensor data. The sample data subfolder contains (1) surface water and sediment dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) surface water and sediment total nitrogen data and averages; (3) sediment grain size data; (4) sediment iron (II) data and averages; (5) wet sediment mass, dry sediment mass, water mass, and wet sediment volume in incubation and sediment ICR vials; (7) sediment incubation respiration rate data and averages; (8) normalized respiration rate data and averages; (9) methods codes; (10) sediment percent carbon and nitrogen; (11) sediment X-ray diffraction (XRD) data; (12) gravimetric moisture and averages; (13) a subfolder with sediment incubation respiration data, scripts, and plots; (14) surface water and sediment FTICR methods; and (15) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains five subfolders, one containing the sediment .xml data files, one containing the water .xml files, one containing the sediment CoreMS output files, one containing the water CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). The sensor data subfolder contains (1) a subfolder with miniDOT dissolved oxygen and temperature data and plots; (2) miniDOT dissolved oxygen and temperature summary data; and (3) miniDOT installation methods. All files are .csv, .pdf, .R, .xml, .d, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, .png, .mov, or .mp4. CORRECTION: The data processing methods for FTICR described in “v3_WHONDRS_AV1_Methods_Codes.csv” mistakenly indicate that users should process the data in Formultitude. The corrected description should read: “Both unprocessed and processed data are provided to allow users flexibility in data processing. Instructions and scripts for processing the data using CoreMS are included.” CORRECTION: Carbon and nitrogen content are reported as percentages. The current column headers "01395_C_percent_per_mg" and "01397_N_percent_per_mg" are incorrect. These should read "01395_C_percent" and "01397_N_percent" and will be corrected in the next version of this data package.

54 ENVIRONMENTAL SCIENCES↗

Data and Scripts associated with a manuscript on ecosystem responses to wildfires in the Columbia River Basin

This data package is associated with the publication “Ecosystem leaf area, gross primary production, and evapotranspiration responses to wildfire in the Columbia River Basin” submitted to Biogeosciences (Shi et al., 2024; doi: 10.22541/au.171053013.30286044/v1). In this research, data products, leaf area index (LAI), gross primary production (GPP), and evapotranspiration (ET), from the Moderate Resolution Imaging Spectroradiometer (MODIS) are used to quantify the resistance and resilience of different ecosystem types in the Columbia River Basin (CRB). A machine learning algorithm, random forest (RF), was used to examine the impacts of precipitation, vapor pressure deficit (VPD), and burn severity from Monitoring Trends in Burn Severity (MTBS) on ecosystem resilience. The data package includes the processed MODIS data products, precipitation, VPD, and burn severity in 138 fire regions in CRB and the input files for RF model training. This data package includes six folders. The MODIS products are included in three MODIS_* folders with shell scripts for data clipping and *ncl files for data processing: (1) “/MODIS_LAI_CRB”; (2) “/MODIS_GPP_CRB”; and (3) “/MODIS_ET_CRB”. All the processed data for each fire event are NetCDF formatted. The MTBS burn severity data and the shell and *ncl scripts used for data processing are in the folder named (4) “MTBS_fire”. The ERA meteorological fields and the data processing scritps are in (5) “ERA_Var_CR”. All the scripts for figure development are in the format of *ncl and in the folder (6) “paper_scripts”. See the file ending in “flmd.csv” for a list of all files contained in this data package and descriptions for each. Tabular column headers and units are described in the data dictionary file ending in “dd.csv”.

54 ENVIRONMENTAL SCIENCES↗

Soil Core Chemistry of Wetland, Old Woman Creek National Estuarine Research Reserve, Huron, OH, 2023-10-24 to 2024-05-20

This dataset contains the chemistry data of soil samples collected from a wetland, referred to as The Cove, at Old Woman Creek Estuarine Research Reserve in Huron, OH. Soil core extractions were performed to analyze what nutrient and/or metal constituents were present at different depths and what biogeochemical activity this could indicate. Three soil cores were collected at three locations within The Cove. The soil cores were removed from their core tubing and were cut into 4 segments down the length (or depth) of the core: top to 1-inch deep, from 1 inch to 5 inches, 5 inches to 7 inches, and 7 inches to the bottom of the core (approximately 10 inches). These soils segments were each homogenized and sub-sampled for various chemical analyses. Soil chemistry measurements are reported in the SoilChem_DataTable.csv file. Collection information about the samples can be found within the SoilChem_SampleMetdata.csv file. All files associated with this dataset are listed in the SoilChem_FLMD.csv file.

54 ENVIRONMENTAL SCIENCES↗

Data Files for Runoff Evaluation in an Earth System Land Model for Permafrost Regions

Modeling of hydrological runoff is essential for accurately capturing spatiotemporal feedbacks within the land–atmosphere system, particularly in sensitive regions such as permafrost landscapes. However, substantial uncertainties persist in the terrestrial runoff parameterization schemes used in Earth system and land surface models. This is particularly true in permafrost regions, where landscape heterogeneity is high and reliable observational data are scarce.This data set includes all files that were produced and applied in the paper Runoff Evaluation in an Earth System Land Model for Permafrost Regions [Xiang et al. in review]. The paper is in review as of July 1 2025 in Geoscientific Model Development (GMD). In this study, we evaluate the performance of runoff parameterization schemes in the Energy Exascale Earth System Model (E3SM) land model (ELM). Our proposed framework leverages simulation results from the Advanced Terrestrial Simulator (ATS), which is a physics-rich integrated surface/subsurface hydrologic model that has been successfully evaluated previously in Arctic tundra regions. We used ATS to simulate runoff from 22 representative hillslopes in the Sagavanirktok River basin, located on the North Slope of Alaska, then compared the output with ELM’s parameterized representation of total runoff. This dataset contains 2 figure image files (*.png, *jpg) that describe the study site and methods, as well as folders (Figure*.zip) that contain the associated data files (*.csv, *.dat) and python code notebooks (*.ipynb) for figures 3-7 in the paper. Jupyter notebook (*.ipynb) files that produce the figure files using the associated data files will run within a python environment configured with Jupyter Lab or Notebook packages.

54 ENVIRONMENTAL SCIENCES↗

Machine learning model inputs, outputs, and scripts associated with “Artificial intelligence-guided iterations between observations and modeling significantly improve environmental predictions”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “Artificial intelligence-guided iterations between observations and modeling significantly improve environmental predictions” (Malhotra et al., in prep). This effort was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the contiguous United States (CONUS). New machine learning models were created every month to guide sampling locations. Data from the resulting samples were used to test and rebuild the machine learning models for the next round of sampling guidance. Associated sediment and water geochemistry and in situ sensor data can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1923689, https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1729719, and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1603775. This data package is associated with two GitHub repositories found at https://github.com/parallelworks/dynamic-learning-rivers and https://github.com/WHONDRS-Hub/ICON-ModEx_Open_Manuscript. In addition to this readme, this data package also includes two file-level metadata (FLMD) files that describes each file and two data dictionaries (DD) that describe all column/row headers and variable definitions. This data package consists of two main folders (1) dynamic-learning-rivers and (2) ICON-ModEx_Open_Manuscript which contain snapshots of the associated GitHub repositories. The input data, output data, and machine learning models used to guide sampling locations are within dynamic-learning-rivers. The folder is organized into five top-level directories: (1) “input_data” holds the training data for the ML models; (2) “ml_models” holds machine learning (ML) models trained on the data in “input_data”; (3) “examples” contains files for direct experimentation with the machine learning model, including scripts for setting up “hindcast” run; (4) “scripts” contains data preprocessing and postprocessing scripts and intermediate results specific to this data set that bookend the ML workflow; and (5) “output_data” holds the overall results of the ML model on that branch. Each trained ML model resides on its own branch in the repository; this means that inputs and outputs can be different branch-to-branch. There is also one hidden directory “.github/workflows”. This hidden directory contains information for how to run the ML workflow as an end-to-end automated GitHub Action but it is not needed for reusing the ML models archived here. Please see the top-level README.md in the GitHub repository for more details on the automation. The scripts and data used to create figures in the manuscript are within ICON-ModEx_Open_Manuscript. The folder is organized into four folders which contain the scripts, data, and pdf for each figure. Within the “fig-model-score-evolution” folder, there is a folder called “intermediate_branch_data” which contains some intermediate files pulled from dynamic-learning-rivers and reorganized to easily integrate into the workflows. NOTE: THIS FOLDER INCLUDES THE FILES AT THE POINT OF PAPER SUBMISSION. IT WILL BE UPDATED ONCE THE PAPER IS ACCEPTED WITH ANY REVISIONS AND WILL INCLUDE A DD/FLMD AT THAT POINT. We thank the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, Cowiche Canyon Conservatory, Washington State Parks and Recreation Commission (Scientific Research Permit #210901), and the Confederated Tribes and Bands of the Yakama Nation for access to field locations where the samples labeled “SSS” were collected. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview. WHONDRS consortium members were asked to provide any acknowledgments for the collection of samples labeled “CM” and the following is a list of acknowledgments that were submitted with their corresponding Site IDs: (MART) Research activities were conducted in part on the Wind River Experimental Forest within the Gifford Pinchot National Forest; (MP- 100379) Philadelphia is part of Lenapehoking, the ancestral homelands of the Lenape peoples; (MP-102398) Land surveyed is the ancestral homelands of the Nookhose'iinenno (Arapaho), Tsis tsis'tas (Cheyenne), and Nuuchu (Ute); (MP-100749 and MP- 100747) Georgia Coastal Ecosystem LTER, OCE-1832178; (SP-70 and SP-72) Eastern Shoshone, Shoshone-Bannock; (MP- 102944) Funded by Oregon Watershed Enhancement Board. On the traditional lands of the Confederated Tribes of the Siletz, Confederated Tribes of the Grand Rhonde, and the Clatsop-Nehalem Confederated Tribe; (MP- 100607) Holiday Creek is located on the traditional territory of the Monacan Indian Nation; (SP-45) Lafayette Blue Springs State Park; (MP-102420) NSF DEB-2016749; (MP-100019) New Hampshire Agriculture Experiment Station; (SP-35) Rayonier (land owner; https://www.rayonier.com/); (MP- 101276) US Department of Energy, Office of Science, Biological and Environmental Research, Subsurface Biogeochemical Research, Watershed Dynamics and Evolution SFA at ORNL; (MP- 103224) Watershed Dynamics and Evolution SFA at ORNL; (MP- 101584) Traditional lands of the Oceti Sakowin (Dakota, Lakota, Nakoda) and Anishinaabe Peoples.

54 ENVIRONMENTAL SCIENCES↗

Digital camera imagery for vegetation phenology, Seward Peninsula, Alaska, 2022-2023

Timelapse camera images from Council Mile Marker (MM) 71, Kougarok MM 64, Kougarok Fire Complex (KFC), and Teller MM 27 NGEE-Arctic field sites on the Seward Peninsula, Alaska, captured from July 2022 to July 2023. Eight Wingscape Timelapse Pro cameras, and thirty-one Power-interval Camera Automation Modules (PiCAMs) designed by Brookhaven National Laboratory?s Terrestrial Ecosystem Science and Technology (TEST) group were deployed targeting patches of low and tall shrubs (including Alnus sp. and Salix sp.) and general vegetation and landscape views. Images from Wingscape cameras were recorded at hourly intervals from 11 AM to 2 PM, and images from PiCAMs were recorded at 5 hourly intervals from 12 AM to 8 PM, continuously for 12 months and capture vegetation phenology, snow accumulation and snow melt events. This data package includes images (*.jpg), organized by site and camera ID, and metadata with details of the cameras used, number of images recorded, start and end dates, GPS locations and example fields of view. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Mondo: integrating disease terminology across communities

Precision medicine aims to enhance diagnosis, treatment, and prognosis by integrating multimodal data at the point of care. However, challenges arise due to the vast number of diseases, differing methods of classification, and conflicting terminological coding systems and practices used to represent molecular definitions of disease. This lack of interoperability artificially constrains the potential for diagnosis, clinical decision support, care outcome analysis, as well as data linkage across research domains to support the development or repurposing of therapeutics. There is a clear and pressing need for a unified system for managing disease entities⁠—including identifiers, synonyms, and definitions. To address these issues, we created the Mondo disease ontology—a community-driven, open-source, unified disease classification system that harmonizes diverse terminologies into a consistent, computable framework. Mondo integrates key medical and biomedical terminologies, including Online Mendelian Inheritance in Man (OMIM), Orphanet, Medical Subject Headings (MeSH), National Cancer Institute Thesaurus (NCIt), and more, to provide a comprehensive and accurate representation of disease concepts with fully provenanced and attributed links back to the sources. Mondo can be used as the handle for curation of gene–disease associations utilized in diagnostic applications, research applications such as computational phenotyping, and in clinical coding systems in clinical decision support by pointing the clinician to the numerous knowledge resources linked to the Mondo identifier. Mondo's community-centric approach, stewarded by the Monarch Initiative's expertise in ontologies, ensures that the ontology remains adaptable to the evolving needs of biomedical research and clinical communities, as well as the knowledge providers.

biomedical informatics↗

Vegetation classification map and covariates associated with NEON AOP survey, East River, CO 2018

This package includes geospatial data layers developed to investigate how environmental gradients—specifically topography and near-surface soil properties—drive the spatial arrangement of dominant plant communities in mountainous watersheds. The geospatial products, which support the analysis of these ecological relationships, are derived from airborne hyperspectral and LiDAR datasets acquired by the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP), in conjunction with an extensive ground field campaign conducted in summer 2018. This work is part of the DOE Watershed Function Science Focus Area (SFA) and features geospatial datasets developed based on observations and ground data collected at East River, Colorado, in collaboration with the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP) survey in June 2018. Classification Map: - Classification Map (PNG, GeoTIFF): Derived from hyperspectral and LiDAR airborne data using a machine learning approach. - Class Code Mapper (CSV): Associates pixel values with corresponding vegetation/non-vegetation classes. - Classification Reference Data (CSV): Reference data used in the machine learning procedure. LiDAR-Derived Products: - Topographical Metrics (GeoTIFFs): Elevation, slope, curvature, TWI, TPI, solar insolation, and canopy height model (CHM), smoothed with a 5x5 pixel window. Vegetation Indices: - GeoTIFFs of NDVI, NDNI, NDWI: Vegetation indices derived from hyperspectral data. Urban Masks: - Urban Mask (GeoTIFF): Applied to the mapping to convert bare soil classes to urban classes. Software Compatibility: GeoTIFFs: Can be visualized with GIS software or libraries that support GeoTIFF images. CSV Files: Can be opened with any software that handles comma-separated values. The FLMD file provides details and links to the source datasets used to derive the products. The manuscript (in the Method session) provides details on how each product was derived. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Update on 2026-03-25: Since the original dataset publication date of 02/28/2020, this package has a new classification map derived by an improved methodology. This update also includes additional ground data that improved the representation of some of the communities. See the methods for further details on what has changed between versions.

2018 NEON and 2025 CHESS Campaigns↗

Aqueous Organic Matter from Kougarok Fire Complex, Alaska, 2023

Chemical analyses of aqueous organic matter extracted by filtration from a small set of organic layer samples collected from burned and unburned tussock tundra sites in the Kougarok Fire Complex, near Nome, Alaska. There are five files in *.csv format with one data file and four data description files including data dictionary, methods, terminology, and file-level metadata.The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

PreSens Dissolved Oxygen and Temperature Data, Old Woman Creek National Estuarine Research Reserve, Huron, OH, 2022-07-06 to 2023-12-15

This dataset contains collected dissolved oxygen (DO) and temperature measurements of the air, surface water, and underlying soil of a wetland at Old Woman Creek Estuarine Research Reserve in Huron, OH. Dissolved oxygen and temperature measurements were made in the first few centimeters of wetland sediment to assess how oxygen profiles changed over time with hydrological events. Measurements taken outside of the sediment were taken as comparison. Measurements were taken by hand with a PreSens Fibox 4 transmitter, paired with an oxygen dipping probe and temperature sensor, at each field outing. Measurements were taken approximately every 2 weeks around 10 AM EST. The PreSens_DO_Temp_Datafile.csv contains the recorded measurements, and the PreSens_InstallationMethods.csv describes the deployment of the sensor.

54 ENVIRONMENTAL SCIENCES↗

miniDOT Logger Dissolved Oxygen and Temperature Data of Wetland Surface Water, Old Woman Creek NERR, Huron, OH, 2022-06-15 to 2023-12-15

This dataset contains timeseries data of dissolved oxygen (DO) and temperature measurements of the surface water in a wetland at Old Woman Creek Estuarine Research Reserve in Huron, OH. Dissolved oxygen and temperature measurements were made in the overlying water column of wetland to assess how oxygen changed over time with hydrological events. Measurements were collected by a miniDOT Logger. Data was collected over 1.5 years (June 2022 to December 2023). The miniDOT_DO_Temp_DataFile.csv contains the DO and temperature measurements that were collected every 10 minutes. The water levels of the site varied over-time as the wetland flooded and dried. So, sometimes the miniDOT logger would be out of the water column, resulting in high oxygen levels. Depth of the water column was recorded at every in-person site visit. The miniDOT_Depth_DataFile.csv contains the hand-measured surface water depth measurements for comparison to the logger-collected data. Information on the deployment and measurement methods can be found in the miniDOT_InstallationMethods.csv file.

54 ENVIRONMENTAL SCIENCES↗

Plant Characteristics, Porewater, Gas Flux, and Soil Biogeochemistry at Council Road Site Mile Marker 71, Seward Peninsula, Alaska, 2023

Data collected at Council, AK (64°51’35.0”N 163°41’59.1”W) during a summer campaign in 2023. Water data consists of soil porewater collected by centrifuging soil cores and also by field collection with porewater samplers (rhizons). Gas data consists of CO2, CH4 and N2O surface soil fluxes measured with a portable FTIR analyzer. Plant and root data consists of biomass, root length, diameter and mass. Soil data consists of total C and N. Air, water and soil samples span two main locations: a thermokarst wetland and an upland tussock. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).This dataset was generated to broadly address the following research question: how will climate change (i.e., thawing permafrost, landscape change) alter the ecosystem flux (sink versus source) of important greenhouse gases such as CO2, CH4 and N2O?Description of the contents of this data package: This dataset contains 5 different individual .csv files containing plant, soil, water and gas data. No software is needed to utilize them. PFTCover: Plant functional type ground cover in 1x1 meter plots. SoilCores: Solidphase and porewater phase soil biogeochemical variablesPlantData: Above and belowground plant traits. GasFlux: Surface plant-soil gas measurementsFieldPorewater: Field collected porewater biogeochemical variables

54 ENVIRONMENTAL SCIENCES↗

Data for Roebuck et al. (2025), "Differences in dissolved organic matter composition between rivers and estuaries is conserved across freshwater and saltwater coastal regions"

Dissolved organic matter (DOM) in coastal surface waters influences local water quality and is an important component of biogeochemical cycling in coastal systems, but the processes that alter DOM composition along lower reaches of rivers and estuarine waters are poorly understood. Roebuck et al. (2025) leveraged a spatially distributed community sampling effort in coastal ecosystems across two regions to identify broad spatial drivers of surface water DOM composition and identify transferable trends between saltwater and freshwater coastal systems. Samples were collected by community members from 47 locations within the mid-Atlantic and Great Lakes coastal regions.This dataset includes:* A selection of commonly reported absorbance and fluorescence peaks normalized to dissolved organic carbon concentrations* Parallel factor output from the EC1 fluorescence datasets* A selection of commonly reported absorbance and fluorescence peaks * Spectral indices output from matlab script for absorbance and fluorescence datasets* CO2sys calculations of pH changes under varying temperatures and a constant salinity, DIC, and alkalinity concentrationAll data files are plain-text CSV (comma separated value) and no special software is required to read them.

54 ENVIRONMENTAL SCIENCES↗