Search NASA⌕ Search

SEARCH · Search NASA

Results for “Science Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Dated soil C–N–P profiles, water quality, and chamber fluxes across Ohio and Michigan wetlands (2024–2025)

This dataset includes dated soil core chemistry (bulk density, phosphorus, nitrogen and carbon concentrations), water quality, and chamber flux measurements collected from wetlands in the Midwest United States—12 sites in Ohio, one site in Indiana, one site in Michigan—collected in the spring or summer of 2024 or 2025, all in (.csv) format. These data were generated to examine how wetland restoration, management activities, and time since restoration affect biogeochemical processes, carbon sequestration, nutrient accumulation, water quality, and greenhouse gas emissions. Specifically, these data aim to investigate how restored wetlands differ from natural wetlands in terms of carbon, nitrogen, phosphorus dynamics, as well as carbon dioxide and methane fluxes. Also included are surface and porewater quality parameters and chamber flux measurements across these different wetlands. Sampling was conducted at various sites representing a range of restoration stages, from about 4 years post-restoration up to 105 years post-restoration, and also includes a natural wetland used as a reference in Michigan. These data can be used to determine carbon sequestration rates, nutrient cycling, and to enhance our understanding of biogeochemical responses to wetland restoration in temperate ecosystems. This data package contains (1) a csv file (Water_Quality.csv) containing water quality data (dissolved organic carbon, total dissolved nitrogen, and temperature) organized by location; (2) a csv file (Soil_C_N_P_Seq.csv) containing carbon, nitrogen, and phosphorus concentrations at each soil level and time of each soil level, as well as their sequestration rates; (3) a csv file (CH4_CO2_Flux.csv) including methane and carbon dioxide fluxes that were measured with a chamber; (4) a file-level metadata (FLMD.csv) file that lists each file contained in the dataset with associated metadata; (5) a data dictionary (DD.csv) file that contains terms/column headers used throughout the files along with a definition, units, and data type; and (6) a locations metadata file (Location_metadata.csv).

Earth Science > Atmosphere > Atmospheric Chemistry↗

Topography, surface water distribution and subsurface structure in 2023 across an Arctic coastal tundra site near Utqiagvik, Alaska

Subsurface electrical resistivity tomography (ERT), active layer thickness measurements, photogrammetry, and topographic data were collected in September 2023 along a 475 m long, 20 m wide corridor that traverses various polygon types within the Barrow Environmental Observatory (BEO) on the Alaskan Arctic Coastal Plain, approximately 4 miles from the Beaufort Sea near Utqiaġvik, Alaska. These measurements were designed to assess decadal changes in surface water distribution, topography, and subsurface structure across this dynamic landscape. This archive contains the datasets acquired in 2023 and references to the datasets acquired previously at the same location. The ERT survey was conducted along the 475 m transect using 0.5 m electrode spacing and a roll-along acquisition strategy. Thaw layer thicknesses were measured with a tile probe along the same transect. Photogrammetry data were acquired using an unoccupied aerial vehicle (UAV) and were used to generate a digital elevation model and an RGB mosaic. A real-time kinematic (RTK) GPS was used to survey the ERT electrodes and the ground control points for the aerial imagery. The dataset contains 5 *.csv data files, 6 *.csv metadata files, and 6 *.tif files.The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Topsoil bulk geochemical compositions - An updated harmonized global dataset

Mineral weathering is a key biogeochemical process because of the capacity of minerals to stabilize organic matter. However, predicting soil weathering status across large spatial areas still isn’t possible due to a lack of global data and theoretical frameworks. To address this knowledge gap, multiple global datasets of bulk topsoil geochemical compositions have been harmonized using R. These datasets document topsoil bulk geochemical compositions across five continents (n = ~16,000 observations). Source data for these observations include the EuroGEOSurveys Geochemical Baseline Database (FOREGS), the US Geological Survey National Geochemical Database (NASGLP), the Geochemical Atlas of Australia (GAA), the US Geological Survey Alaska Geochemical Database (AGD84), the National Cooperative Soil Survey (NCSS), the European Geochemical Mapping of Agricultural Soil (GEMAS), Ecorespira-Amazon (ERA), the New Zealand Geochemical Baseline Survey (NZ_GBS), and the African Soil Information Service (AFSIS). Major elements observed include Aluminum (Al), Calcium (Ca), Iron (Fe), Potassium (K), Magnesium (Mg), Sodium (Na), Titanium (Ti), Manganese (Mn), Phosphorus (P), Carbon (C), and Sulfur (S). This data package includes the harmonized dataset itself, and the R scripts necessary to harmonize these datasets, in addition to metadata that describes all columns, files, and databases used in this project. Methods & Sampling Step 1 – Databases of geochemical data identified This study aimed to leverage existing measurements of topsoil geochemical data. Databases were first identified and deemed appropriate for inclusion if they were measuring soils and performed these measurements on the <2mm soil fraction. Databases such as NCSS and AGD84 needed more post processing to include in the database and this was done using the NCSS_datamerge_031626 R file and Alaska_USGSmerge_031626 R file, respectively. Step 2 – Database harmonization Once appropriate databases were identified, they were harmonized for ease of analysis using the R script Database_Harmonization_031826. This included removing columns from original datasets that would not be used in analysis (removed columns are noted in the code). Then, data cleaning procedures specific to each dataset were undertaken. This includes standardizing columns to include units and adding metadata columns regarding procedures for analyzing specific elements. Functions for standardizing measurements and units are outline in R files: calculate element_mg_kg_031626, calculate_oxide_wt_perc_031626, change_oxide_caps_031626, and conv_2_numeric_031626. This also included adding a unique identifier for each sample to identify it with its respective database (see CD_ID in data dictionary). Geographic information: Data reflect a compilation of datasets collected globally. Geographic areas covered by each of the datasets include: - EuroGEOSurveys Geochemical Baseline Database (FOREGS) - European continent - North American Soil Geochemical Landscapes (NASGLP) - continental United States and limited parts of Canada (see database key for more details) - National Geochemical Survey of Australia (GAA) - Australia - Alaska geochemical database (AGDB4) - Alaska - National Cooperative Soil Survey (NCSS) - Global measurements, but concentrated in the continental United States - Geochemical data for arable land and land under permanent grass cover in continental Europe (GEMAS) - continental Europe - Ecorespira-Amazon (ERA) - Geochemical data from the Amazon basin - Geochemical baseline data for New Zealand (NZGBS) - New Zealand - Geochemical data collected across continental Africa (AfSIS) - Measurements across Africa

EARTH SCIENCE > LAND SURFACE > SOILS↗

A Relevancy Algorithm for Curating Earth Science Data Around Phenomenon

Earth science data are being collected for various science needs and applications, processed using different algorithms at multiple resolutions and coverages, and then archived at different archiving centers for distribution and stewardship causing difficulty in data discovery. Curation, which typically occurs in museums, art galleries, and libraries, is traditionally defined as the process of collecting and organizing information around a common subject matter or a topic of interest. Curating data sets around topics or areas of interest addresses some of the data discovery needs in the field of Earth science, especially for unanticipated users of data. This paper describes a methodology to automate search and selection of data around specific phenomena. Different components of the methodology including the assumptions, the process, and the relevancy ranking algorithm are described. The paper makes two unique contributions to improving data search and discovery capabilities. First, the paper describes a novel methodology developed for automatically curating data around a topic using Earthscience metadata records. Second, the methodology has been implemented as a standalone web service that is utilized to augment search and usability of data in a variety of tools.

earth science phenomena↗

Daily evapotranspiration changes during heatwaves at 32 NEON sites, 2019-2021

This dataset provides partitioned evapotranspiration (ET, the combined loss of water from soil and plant surfaces) anomalies during heatwave events—soil evaporation (E) and transpiration (T)—for 268 heatwave events across 32 National Ecological Observatory Network (NEON) flux sites in the contiguous United States from 2019–2021. Using an ensemble of four high-frequency turbulence methods (Flux-variance Similarity, Conditional Eddy Covariance [CEC], CEC with Water-Use Efficiency, and Conditional Eddy Accumulation; see Zahn and Bou-Zeid 2024), half-hourly transpiration-to-evapotranspiration (T/ET) ratios were derived from 20 hertz (Hz, cycles per second) eddy covariance measurements of carbon dioxide (CO₂) and water vapor (H₂O) concentrations. The dataset spans six vegetation types including evergreen and deciduous forests, grasslands, cultivated crops, shrublands, and emergent herbaceous wetlands. Data Package Contents: The dataset includes a single CSV (comma-separated values) file containing daily anomalies (deviations from baseline conditions) for transpiration (Delta_T), evaporation (Delta_E), total evapotranspiration (Delta_ET), and T/ET ratio (Delta_T_ET) during each day of identified heatwave events. The file also includes site codes, dates, heatwave event identifiers, and day-of-heatwave indicators. The CSV file can be opened with spreadsheet software (Microsoft Excel, Google Sheets) or programming environments (Python, R, MATLAB). This resource enables researchers to investigate ecosystem-specific responses to thermal extremes, validate land surface model partitioning of ET fluxes, and examine feedbacks between water cycling and surface energy balance during heatwaves. The dataset is particularly valuable for studies linking vegetation hydraulic strategies to climate resilience, as it captures the divergent responses of shallow-rooted versus deep-rooted ecosystems. Potential applications include improving drought early warning systems, informing irrigation management strategies, and advancing our mechanistic understanding of land-atmosphere interactions under extreme heat conditions.

Day of Heatwave↗

Patch-level CO2 and CH4 fluxes and porewater concentrations in experimental wetlands, 5 and 10 PPT saltwater intrusion simulations, Louisiana 2023-2024

This dataset containes carbon dioxide (CO2) and methane (CH4) flux measurements collected from wetland vegetation patches dominated by Typha domingensis and Panicum hemitomon to assess greenhouse gas flux responses to experimental saltwater intrusion (SWI) pulses. Measurements were conducted before, during, and after simulated SWI events at target salinities of approximately 5 parts per thousand (ppt) with durations of 6, 10, and 17 days and 10 ppt with a duration of 48 days, alongside a control wetland (with no salinity added, flood manipulation only). These data were generated to evaluate how the magnitude and duration of SWI alter wetland carbon exchange and related biogeochemical and plant responses. This data package includes flux measurements from the wetland surface (i.e, soil/water surface and enclosed vegetation) and from the soil/water surface only; porewater and surface water concentrations of CO2 and CH4; salinity, pH, electrical conductivity collected in porewater (at 5, 10, and 20 cm soil depths) and in surface water; soil redox potential; leaf spectral indices, leaf vapor pressure deficit, stomatal conductance; water level, salinity, and photosynthetically active radiation; and aboveground biomass.

EARTH SCIENCE > AGRICULTURE > SOILS > SOIL RESPIRA↗

Interoperability and Other Aspects of Guiding Data Producers for the Benefit of End Users

The purpose of this paper is to discuss how the Climate and Forecast (CF) Metadata Conventions and netCDF standard have influenced the recommendations and guidance provided to producers of data products based on NASA’s Earth observations. It has been long-recognized that interoperable datasets and use of standards and conventions are beneficial to the users of these datasets, especially those who make use of multiple datasets for their research and applications. The Dataset Interoperability Working Group (DIWG), one of NASA’s Earth Science Data System Working Groups (ESDSWGs), was established in 2013, and has developed and published many recommendations. The Data Product Development Guide (DPDG) Working Group, established in 2018 as another of the ESDSWGs, has published a DPDG for Data Producers and a Quick Start Guide, incorporating guidance from many sources, including the recommendations from the DIWG. The DPDG includes recommendations regarding data formats (prominently netCDF-4) and metadata based primarily on the CF Metadata Conventions and the Attribute Convention for Data Discovery (ACDD). In early 2023, it was decided that the Resource Center for Data Producers (RCDP) Working Group be established as another ESDSWG, with the goals of providing all the information relevant and helpful for data producers via an easily accessible website, and of recommending how the DPDG and QSG could be maintained as living documents, given the rapidly changing technologies, and the need for incorporating the experience and feedback from the users of these documents.

Data product development↗

Data and scripts associated with the manuscript "Organic Molecules are Deterministically Assembled in River Sediments"

This data package is associated with the publication "Organic Molecules are Deterministically Assembled in River Sediments" submitted to Scientific Reports (Stegen et al., 2024). The study applies community ecology methods to dissolved organic matter (DOM) chemistry from variably inundated riverbed sediments to uncover principles governing DOM composition at a reach-scale. This data package documents the workflow used to process and generate the main findings in the manuscript. The R scripts reference the raw, unprocessed Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data from another data package, available on ESS-DIVE at https://data.ess-dive.lbl.gov/view/doi:10.15485/1834208. The scripts then process the raw FTICR-MS data and generate the findings and figures presented in the associated manuscript. In brief, this study demonstrates that DOM assemblages in variably inundated sediments are primarily governed by deterministic variable selection, including sediment moisture effecting the degree of deterministic assembly. See the manuscript for more details pertaining to interpretation and implications of the findings. This data package is associated with the GitHub repository found at https://github.com/WHONDRS-Hub/ECA_2020_Sed.This data package is comprised of 6 scripts and 7 folders. The file-level metadata file (file ending in "flmd.csv") lists all files contained in this data package and descriptions for each. The data dictionary (file ending in "dd.csv) describes all tabular data columns and their respective definitions and units. The FTICR_Processing_Scripts produce the outputs found in the "Processed_Data" folder. The remaining scripts (located in the parent directory) produce the outputs found in the following four folders: (1) "MCD_Dendrograms", "MCD_Randomizations", "MCD_bNTI_Outcomes", and "OM_Null_Modeling". The fifth script additionally takes the three comma-separated values (CSV) files found in the parent directory as input ("VGC_texture.csv", "merged_weights.csv", and "ECA2_FTICR_BetaDisp.csv"). The outputs of each of the five scripts serve as the input to the following script, with the final outputs stored in the folder "OM_Null_Modeling".

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with “Sequential Precipitation Input Tagging (SPIT) to Estimate Water Transit Times and Hydrologic Tracer Dynamics within Water-Tagging Enabled Hydrologic Models” (v3)

This data package is associated with the publication “Sequential Precipitation Input Tagging (SPIT) to Estimate Water Transit Times and Hydrologic Tracer Dynamics within Water-Tagging Enabled Hydrologic Models” submitted to Journal of Advances in Modeling Earth Systems (Butler et al. 2025). This study developed the Sequential Precipitation Input Tagging (SPIT) framework to tag input precipitation and estimate water transit times and hydrologic tracers. SPIT tags all precipitation events at regular intervals over an extended period (monthly tags over seven years) in a hydrologic model from 2016-2022. SPIT is applied at six National Ecological Observatory Network (NEON) sites across the continental United States to calculate transit time distributions (TTD) and derive from these mean transit times (MTT), fractions of young water (Fyw), and hydrologic tracer concentrations in stream water (δ18O) within a water-tagging enabled version of the Weather Research and Forecast (WT-WRF-Hydro) model with national water model (NWM) configurations. We go on to validate WT-WRF-Hydro estimates against Butler et al. (2023), who analyzed the same NEON sites using stable water isotope data to estimate water transit times. This new tracking method provides a detailed picture of water movement and helps improve predictions about water availability in the future. This data package was originally published in January 2025. It was updated May 2025 (v2; new and modified files) and October 2025 (v3; new and modified files). File and folder names were not revised to indicate changes. See the change history section in the readme for more details. This data package contains the data and scripts used to develop the SPIT framework WT-WRF-Hydro (Water Tagging Weather Research and Forecasting Hydrologic) model and is associated with the following GitHub repository: https://github.com/zbutler33/SPIT-Framework. This data package contains five parent folders: (1) “Manipulated_outputs”, (2) “Metadata”, (3) “Observed”, (4) “Outputs”, and (5) “Scripts”. Each of these parent folders contains additional subfolders and files. Please see the FLMD (“v*_Butler_2024_WT_WRF_Hydro_flmd.csv”) for a list of all the files contained in this data package and descriptions for each. See the data dictionary (“v*_Butler_2024_WT_WRF_Hydro_dd.csv”) for definitions and units of all of the tabular (files ending in “.csv” and ".tsv") column headers.

54 ENVIRONMENTAL SCIENCES↗

Perspectives on Data Reproducibility and Replicability in Paleoclimate and Climate Science

This paper summarizes the current state of reproducibility and replicability in the fields of climate and paleoclimate science, including brief histories of their development and applications in climate science, new and recent approaches towards improvement of reproducibility and replicability, and challenges. Recommendations for addressing those challenges include: development of searchable, auto-updated, interlinked, multi-archive public paleoclimate repositories for raw and processed digital datasets; cross-center standardized code base cases, improved data storage techniques, and a focus on replicability for climate simulation storage and access; and support of the development and community awareness of findable, accessible, interoperable and reusable (FAIR) principles by funding agencies and publishers. This paper is largely based on the May 2018 presentations of a panel of researchers to the Committee on Reproducibility and Replicability in Science, part of the National Academies of Science, Engineering, and Medicine. The commentary and recommendations made here are in alignment with those of its Consensus Study Report on Reproducibility and Replicability in Science (2019).

data repositories↗

Simplifying Analysis of Hierarchical HDF5 and NetCDF4 Files with Xarray-Datatree

NASA’s Earth Observing System Data and Information System (EOSDIS) contains thousands of Earth science datasets from satellites, models, and field campaigns. EOSDIS data are stored in formats that are well supported by the Earth Science community. These formats include the Hierarchical Data Format (HDF), with derivative flavors such as HDF-5 and the Network Common Data Format (NetCDF-4). The HDF specification allows for a directory-like hierarchy within a single file, known as "groups". Observational data and associated metadata within a single file can be distributed amongst multiple internal groups, which can also be nested to multiple levels. Working with datasets that have a group hierarchical structure can be difficult because of the nested structure of groups. Widely used packages, such as xarray, have data models that do not accommodate the hierarchical structure within HDF files, requiring users to traverse the file and open different HDF groups as separate, unrelated objects. Xarray-datatree is a Python package developed to solve the difficulty of traversing HDFs with a hierarchical group structure by creating a tree-like hierarchical data structure in xarray. The tree-like structure allows each group to be accessed once a DataTree object is instantiated. The migration of xarray-datatree into the xarray core library will reduce barriers to accessing Earth science data by eliminating the need to understand and traverse the specific hierarchy of a grouped HDF file.

Eni Awowale↗

Open Science for Plants in Space: Data Sharing, Standards, and Informatics for Reuse and Knowledge Discovery

Upcoming deep space missions rely on plants and crops for crew and ecosystem health. Access to space plant data enables scientists to gain a deeper understanding of biological responses to ionizing radiation, altered gravity, low atmospheric pressure, elevated CO2, and altered photoperiods. Open Science is the practice of making research available to all, while respecting diverse cultures, fostering collaborations with equity. 2023 is the ‘Year of Open Science’, and NASA has a 5-year Transform to Open Science (TOPS) mission designed to rapidly transform the agency toward an inclusive culture of open science. NASA’s Open Science Data Repository (OSDR) developed by NASA’s Biological and Physical Sciences Division provides access to data from space-relevant biological experiments. OSDR combines two databases, GeneLab and Ames Life Sciences Data Archive (ALSDA) to maximize access to standardized ‘omics (e.g., transcriptomics, proteomics) and phenotypic data (e.g., microscopy, biomass), respectively. OSDR started in 2014 with the creation of the first space-relevant FAIR (Findable, Accessible, Interoperable, Reusable) biological ‘omics repository (GeneLab), providing detailed metadata on investigation, sample, and assay levels. Today, GeneLab hosts 62 plant datasets which have led to 5 published peer-reviewed meta-analysis publications. Most of these publications were collaboration efforts under the OSDR Analysis Working Groups (AWGs). AWGs provide great opportunities for investigators to collaborate and set new standards for space-relevant data and metadata. The AWGs are welcoming any ASPB members interested in providing plant expertise for space biology. The addition of ALSDA to OSDR is also expanding analysis capability beyond ‘omics. Now is the time to get involved as a Subject Matter Expert as we establish the framework for modern plant data archiving through the AWGs. Investigators are invited to submit their space-relevant plant datasets to OSDR and visit the site to learn about the tools OSDR has to offer (osdr.nasa.gov/bio).

FAIR↗

Canopy spectral reflectance, Kougarok and Teller sites, Seward Peninsula, Alaska, 2016

Measurements of full-spectrum (350-2500 nm) canopy spectral reflectance of Arctic plant species at the Teller and Kougarok NGEE-Arctic sites, Seward Peninsula, Alaska. Spectra were collected in July 2016 using an SVC HR-2014i spectroradiometer together with a Spectralon white plate to calibrate each measurement under variable illumination conditions. The locations of the 43 measurement targets are provided as latitude and longitude recorded by the spectroradiometer internal GPS. This data package comprises .csv data and metadata files, and the SVC instrument output (.sig in .zip). The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Data Sharing in Radiobiology; Towards FAIR

The value of scientific data depends on their findability, accessibility, integrability and reusability according to the FAIR principles. Together with the sustainability of data preservation and access, these principles underpin the long term benefits of scientific research. Within the domain of radiobiology we have a huge array of data types, themes and complexities which make standardisation of metadata, data structure and data integration very challenging. Moreover, it is clear that, for example, in the area of disaster preparedness, the ready discovery and availability of multiple types of data, for example on biological effects of exposure, climatology, ecology, human behavioural and attitudinal studies, is important for an integrated scientific approach. Because these data are spread over many databases, journal supplementary information resources and even the computers of the investigators, their discovery and reuse can be challenging. Despite exhortations from funding agencies and scientific institutions over the past two decades there is still a serious deficit in the willingness and in some cases the ability of investigators to share data, and although much may not be formally „Public domain“, information about the existence of the data, their metadata, and how to obtain them should always be available. We report the progress of work on three databases, the STORE and the NASA GeneLab and LSDA repositories to leverage the Radiation Biology Ontology (RBO), a structured terminology for metadata that can be used by all radiation biology-relevant databases to unite federated and automated data searches across multiple databases, for example using web services, and through semantic web technologies supporting data discovery. The initial primary use-cases for RBO were archiving data in the STORE database (https://www.storedb.org/), the repository used for the RadoNorm and Pianoforte Projects among others, and in the NASA Open Science Data Repository (https://osdr.nasa.gov/bio). The scope of radiobiology research ranges from basic physics to radiation oncology to sociolegal studies; no existing ontology had the necessary breadth or depth to fulfill this need. In addition, a formal ontology has the advantage of being usable for machine learning and, importantly, for tasks like data integration, knowledge extraction from the scientific literature and for query extension and data classification. Standardisation of metadata is one of the primary objectives of the FAIR principles for open data; RBO is an important landmark for FAIR-compliant radiation biology data sharing. The RBO is developed using the open-source tools of GitHub and the OBO Foundry-led Ontology Development Kit, and published through GitHub and the NIH/NCBI BioPortal website. This initial phase of concept modeling has yielded an ontology that has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies with relevance to radiation biology (for example, concepts from the ISO standard Basic Formal Ontology, the Environment Ontology and the Gene Ontology). We welcome input into the development of RBO and encourage its adoption.

ontologies↗

Data Sharing in Radiation Biology: Towards FAIR

The value of scientific data depends on their findability, accessibility, integrability and reusability according to the FAIR principles. Together with the sustainability of data preservation and access, these principles underpin the long term benefits of scientific research. Within the domain of radiobiology we have a huge array of data types, themes and complexities which make standardisation of metadata, data structure and data integration very challenging. Moreover, it is clear that, for example, in the area of disaster preparedness, the ready discovery and availability of multiple types of data, for example on biological effects of exposure, climatology, ecology, human behavioural and attitudinal studies, is important for an integrated scientific approach. Because these data are spread over many databases, journal supplementary information resources and even the computers of the investigators, their discovery and reuse can be challenging. Despite exhortations from funding agencies and scientific institutions over the past two decades there is still a serious deficit in the willingness and in some cases the ability of investigators to share data, and although much may not be formally "Public domain“, information about the existence of the data, their metadata, and how to obtain them should always be available. We report the progress of work on three databases, the STORE and the NASA GeneLab and LSDA repositories to leverage the Radiation Biology Ontology (RBO), a structured terminology for metadata that can be used by all radiation biology-relevant databases to unite federated and automated data searches across multiple databases, for example using web services, and through semantic web technologies supporting data discovery. The initial primary use-cases for RBO were archiving data in the STORE database (https://www.storedb.org/), the repository used for the RadoNorm and Pianoforte Projects among others, and in the NASA Open Science Data Repository (https://osdr.nasa.gov/bio). The scope of radiobiology research ranges from basic physics to radiation oncology to sociolegal studies; no existing ontology had the necessary breadth or depth to fulfill this need. In addition, a formal ontology has the advantage of being usable for machine learning and, importantly, for tasks like data integration, knowledge extraction from the scientific literature and for query extension and data classification. Standardisation of metadata is one of the primary objectives of the FAIR principles for open data; RBO is an important landmark for FAIR-compliant radiation biology data sharing. The RBO is developed using the open-source tools of GitHub and the OBO Foundry-led Ontology Development Kit, and published through GitHub and the NIH/NCBI BioPortal website. This initial phase of concept modeling has yielded an ontology that has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies with relevance to radiation biology (for example, concepts from the ISO standard Basic Formal Ontology, the Environment Ontology and the Gene Ontology). We welcome input into the development of RBO and encourage its adoption.

ontologies↗

Meteorological and Soil Data from Ecohydrology Sensor Towers at Pump House and Snodgrass Mountain in East River Watershed, Colorado, 2019-2025

This data package includes hourly meteorological and soil sensor data at eight ecohydrology monitoring sites in East River Watershed, Colorado as part of the Watershed Function Scientific Focus Area (WFSFA) research led by Lawrence Berkeley National Lab (LBNL). Four field sites were located on the hillslope of East River (ER) near Pump House (PH) at Mount Crested Butte (ER-PHS1 to 4), and the other four are in the Snodgrass Mountain (SG) area (SG-EHS5 to 8). In terms of vegetation cover, three sites are in montane grasslands (ER-PHS1, ER-PHS2, and SG-EHS5), three are below evergreen conifer canopy (ER-PHS3, SG-EHS6, and SG-EHS7), and two are below deciduous aspen canopy (ER-PHS4 and SG-EHS8). The monitoring period began in October 2019 at the East River sites, in October 2020 at SG-EHS5 and SG-EHS6, and in October 2021 at SG-EHS7 and SG-EHS8. In September 2024, all four East River sites were fully retired. The four Snodgrass Mountain sites remain active. Each site is equipped with a comprehensive suite of meteorological sensors on a tripod and soil sensors that measure weather, energy fluxes, and soil variables. This data package includes measurements from ten different types of sensors and up to thirteen individual sensors per site, including (1) a weather station (measurement height ranges from 2.8~3.8 meters (m) above ground), (2) a quantum sensor for photosynthetic active radiation (PAR) (2.4~3.3m), (3) a net radiometer (1.7~2.1m), (4) an infrared radiometer (1.6~2.2m), (5) a sonic distance sensor (1.5~1.9m), (6) a soil carbon dioxide (CO2) flux chamber (0m), (7) a soil heat flux plate (-0.05m below ground), (8) a soil oxygen sensor (-0.3m), (9) a soil water potential sensor (-0.3m), and (10) soil water content sensors at 3~4 depths (-1.15 ~ -0.1m). A total of twenty-three variables is reported in this data package, including (1) atmospheric variables: air temperature (TA), atmospheric pressure (PA), vapor pressure (VP), and vapor pressure deficit (VPD), (2) precipitation variables: rain precipitation (P) and snow depth (D_SNOW), (3) energy fluxes variables: four-component net radiation (NETRAD) (shortwave/longwave incoming/outgoing radiation, SW_IN, SW_OUT, LW_IN, LW_OUT), photosynthetic photon flux density (PPFD), and soil heat flux (G), (4) soil variables: soil water content (SWC), soil water potential (SWP), soil temperature (TS), soil bulk electrical conductivity (COND_SOIL), and soil gaseous oxygen concentration (O2_SOIL), (5) wind variables: two-dimensional wind speed (WS), gust speed (WS_MAX), and wind direction (WD), and (6) surface variables: surface infrared temperature (T_CANOPY) and soil CO2 flux (CO2_SOIL). Please see the Methods section for data processing and QA/QC steps taken to generate the hourly datasets. The following files are included in this data package (notes on version: v{x}-{y}, where x is the metadata version, and y is the data version, when applicable): (1) “metadata_site_v{x}-{y}.csv” - a site metadata file that summarizes location information of all sites, including site ID, description, coordinates, timeframe, elevation, and vegetation cover, (2) “metadata_instrument_v{x}-{y}.csv” - an instrument metadata file that summarizes sensor information of all sites, including sensor manufacturer and model, measurement height, and sampling and averaging interval of all variables, (3) "data_{SITE_ID}_v{x}-{y}.csv" - eight data files that contain hourly data of each site indicated by {SITE_ID} in the filename, (4) “/figure/data_{SITE_ID}_v{x}-{y}.png" - eight figures that help visualize data of each site indicated by {SITE_ID} in the filename, (5) “/photo/*” - photos of each site indicated by {SITE_ID} in the filename, and (6) four file level metadata (flmd.csv) and data dictionary (*_dd.csv) files that summarize file, header, column, and variable information of all files. Notes: (1) Measurement height: Each variable name is followed by conventional positional qualifiers “H_V_R”, where H indicates the relative horizontal positions of that specific variable, V the vertical positions, and R the replicates. In this data package, only the vertical qualifier V varies, and V increases from the highest vertical position (V=1) to the lowest. Variables with the same qualifier are not necessarily measured by the same sensor, and the same variable with the same qualifier across different sites are not necessarily measured at the same height. Please refer to “metadata_instrument.csv” for the sensor information and measurement heights, and whether a variable is measured below the canopy. (2) Variable availability: Snow depth is not available at ER-PHS3 and SG-EHS7. SWC, soil temperature, and soil bulk EC at the deepest depth (<-1m) are not available at SG-EHS6 and SG-EHS7. The missing value code for numeric variables is -9999, except for SWP. For SWP, the missing value code is +9999, because SWP values are negative. (3) Sampling frequency: Please refer to “metadata_instrument.csv” for the increase of sampling frequency of some variables from 30-min to 1-min at ER-PHS1 to 4 in July 2020. (4) Sensors: While the methods of each sensor are not detailed, all sensors are commercially available, and their methods can be found in their manuals. Please refer to “metadata_instrument.csv” for the sensor manufacturer and model information. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

RC-SFA Data Management Templates and Guidance for Standardized, Reusable AI-Ready Data Packages

This data package provides templates and supporting documentation developed by the River Corridor Science Focus Area (RC-SFA; https://www.pnnl.gov/projects/river-corridor) to communicate its approach to managing and publishing AI-ready data. The package is intended to help data users and data producers understand the structures, metadata practices, and quality-control approaches that support consistent, reusable, and machine-actionable data products across RC-SFA studies. Rather than focusing on a single experimental dataset, this package documents the data management framework used to make RC-SFA data easier to find, ingest, navigate, and interpret. The materials in this package reflect RC-SFA practices for standardized data package organization, including the use of a human- and machine-readable README, file-level metadata, data dictionaries, descriptive file naming, method identifiers, and automated and review-based quality assurance procedures. Together, these components illustrate how RC-SFA extends FAIR data principles toward AI-readiness by prioritizing deep metadata, consistency across data packages, and support for informed downstream reuse by both humans and computational tools. This dataset is comprised of (1) readme; (2) presentation slides with an overview of RC-SFA approach and guidance; (3) document of RC-SFA best practices; (4) data dictionary (dd); (5) file level metadata (flmd); and a subfolder containing templates for dd and flmd. All files are .csv and .pdf. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.

AI-readiness↗

Simple, Script-Based Science Processing Archive

The Simple, Scalable, Script-based Science Processing (S4P) Archive (S4PA) is a disk-based archival system for remote sensing data. It is based on the data-driven framework of S4P and is used for data transfer, data preprocessing, metadata generation, data archive, and data distribution. New data are automatically detected by the system. S4P provides services such as data access control, data subscription, metadata publication, data replication, and data recovery. It comprises scripts that control the data flow. The system detects the availability of data on an FTP (file transfer protocol) server, initiates data transfer, preprocesses data if necessary, and archives it on readily available disk drives with FTP and HTTP (Hypertext Transfer Protocol) access, allowing instantaneous data access. There are options for plug-ins for data preprocessing before storage. Publication of metadata to external applications such as the Earth Observing System Clearinghouse (ECHO) is also supported. S4PA includes a graphical user interface for monitoring the system operation and a tool for deploying the system. To ensure reliability, S4P continuously checks stored data for integrity, Further reliability is provided by tape backups of disks made once a disk partition is full and closed. The system is designed for low maintenance, requiring minimal operator oversight.

Lynnes, Christopher↗