Search NASA⌕ Search

SEARCH · Search NASA

Results for “data curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

The Zooplankton International Geospatial (ZIG) dataset: A global repository of spatiotemporal freshwater zooplankton community composition data to support ecological research

Zooplankton play critical roles in aquatic ecosystem function and food webs. Nevertheless, global syntheses of their abundance and community dynamics are challenging due to methodological differences across monitoring programs, taxonomic inconsistencies, and a lack of standardized metadata. To reconcile these challenges, we assembled, curated, validated, and harmonized the Zooplankton International Geospatial (ZIG) dataset, which includes co-located and contemporaneous zooplankton, water chemistry, and limnological data from 307 lakes and reservoirs. ZIG includes waterbodies from each major lake thermal region and range in size from 0.8-2,805,8600 hectares. Temporal coverage for individual waterbodies ranges between 1-60 years of data (median = 4 years) with sampling from once annually to weekly. ZIG is publicly available and can be used to understand freshwater biodiversity change and its drivers at unprecedented scales, and we consider it to be a cornerstone for future investigations of freshwater biology, chemistry, and ecology.

Figary, Stephanie [Cornell University, Ithaca, NY]↗

PhysBERT: A text embedding model for physics scientific literature

The specialized language and complex concepts in physics pose significant challenges for information extraction through Natural Language Processing (NLP). Central to effective NLP applications is the text embedding model, which converts text into dense vector representations for efficient information retrieval and semantic analysis. In this work, we introduce PhysBERT, the first physics-specific text embedding model. Pre-trained on a curated corpus of 1.2 × 106 arXiv physics papers and fine-tuned with supervised data, PhysBERT outperforms leading general-purpose models on physics-specific tasks, including the effectiveness in fine-tuning for specific physics subdomains.

Hellert, Thorsten (ORCID:0000000227970926)↗

2024 Buildings Technology Baseline: Dataset Documentation

The Buildings Technology Baseline is a curated and regularly updated dataset of current and projected performance, retail, and installed price data for all major building energy technologies needed to enable cost/benefit analyses. Building technology analyses require an up-to-date understanding of installation costs and cost-effectiveness of key building energy efficiency technologies. The dataset was assembled by Guidehouse during fiscal year 2024. Data was gathered from the 2024 National Residential Efficiency Measures Database (NREMDB), the 2023 Energy Information Administration Updated Buildings Sector Appliance and Equipment Costs and Efficiencies ("EIA Building Data Report"), DOE Lighting Market Model, the 2023 RSMeans database, and the 2020 Grid-Interactive Efficient Building Technology Cost, Performance, and Lifetime Characteristics ("GEB Data Report"), Lawrence Berkeley National Laboratory, various literature, as well as new data from online retailers, stakeholder interviews, and contractor databases in 2023 and 2024. The dataset has been reviewed by subject matter experts at NREL and DOE. The 2024 dataset release is intended to be a starting point for interested users to provide feedback. This database is not intended to provide specific cost estimates for a specific project. The cost estimates do not include any rebates or tax incentives that may be available for the measures. Rather, it is meant to help determine which measures may be more cost-effective. The National Renewable Energy Laboratory (NREL) makes every effort to ensure accuracy of the data; however, NREL does not assume any legal liability or responsibility for the accuracy or completeness of the information.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

AI-Ready Data Pilot Project Report

The proliferation of artificial intelligence in scientific research has created an urgent need to define "AI-ready data" for researchers and, more importantly, provide resources to help them produce AI-ready data. At Pacific Northwest National Laboratory, we conducted a pilot study with three data scientists evaluating three CSV datasets from different scientific domains, followed by semi-structured interviews capturing assessment practices. Our findings reveal that AI-readiness evaluation is intuition-based, with practitioners asking "How fast can I go from raw data to my machine learning pipeline?" Data scientists consistently prioritized workflow efficiency, human interpretability, and quality stewardship signals. From these insights, we developed a practical evaluation framework comprising data requirements, metadata standards, and validation tests that provides actionable criteria for producing and curating AI-ready datasets, addressing the gap between theoretical understanding and practical implementation.

97 MATHEMATICS AND COMPUTING↗

Stratospheric dust collections: Valuable resources for space and atmospheric scientists

The stratospheric collection at the Johnson Space Center Curatorial Facility offers a unique opportunity to study well-documented, individual particles (or groups of particles) from a wide variety of sources. The nature of the collection and curation process, as well as the timeliness of some sampling periods, ensures that all data obtained from stratospheric particles is a valuable resource for scientists from a wide range of disciplines. A few samples of the uses of these stratospheric dust collections are outlined. An understanding of global parameters at a particular point in time in the stratosphere can be obtained from a study of complete collection surfaces. For example, an accurate assessment of particle concentration over a wide range of sizes was experimentally determined for the stratospheric cloud formed one month after the eruption of El Chichon. Additional studies on the El Chichon cloud over a six-month period showed that volcanic ash settles out of the stratosphere at a rate determined primarily by particle shape and density. Another study during a volcanically quiescent period has shown that total particle number density during the summer of 1981 was approx. 2.7 x 10(-1) cm(-3), for particles 1 micron diameter. However, 95% of these particles were 5 micrometers diameter. With the above classification scheme, an estimate of micrometeorite number density at 20km altitude can also be made. Continuation of these types of studies, for shorter collection periods at regular intervals, can provide important experimental data on the contributions of orbital debris, rocket firings and transient events on the total stratospheric particle budget.

Mackinnon, I. D. R.↗

Using X-Ray Computed Tomography to Image Apollo Drive Tube 73002

The Apollo missions collected 382 kg of rock, regolith, and core samples from six locations on the nearside of the Moon. Today, just over 84% by mass of the Apollo collection remains in pristine condition within the curation facility at Johnson Space Center. Most Apollo samples have been well characterized, however there are several types of samples that have remained wholly or largely unstudied since their return, and/or that have been curated under special conditions. These sample types are: (1) unopened samples sealed under vacuum on the Moon; (2) unopened (but unsealed) drive tubes; (3) Apollo 17 samples frozen shortly after their return; and (4) Apollo 15 samples opened and stored in a helium atmosphere since their return. Last summer, NASA solicited proposals for the Apollo Next Generation Sample Analysis Program (ANGSA), and 9 teams were selected to study: (1) unsealed, unopened drive tube 73002; (2) sealed, unopened drive tube 73001 (paired with 73002); and (3) a subset of the frozen and He-purged samples [1]. The first sample opened as part of the ANGSA program was drive tube 73002. This is a 30 cm long, 4 cm diameter drive tube collected on a landslide deposit near Lara Crater at the Apollo 17 landing site. It was part of a 60 cm long double drive tube collected, and the bottom half of the tube (73001) was sealed under vacuum on the Moon [2]. Prior to opening sample 73002, the sample was imaged with a high resolution Xray Computed Tomography (XCT) scan of the entire tube. Additional XCT scans have been made of “large” clasts removed from the core as part of the dissection process [3]. Here we present a first look at the XCT data from 73002, and talk about the utility of the scans as part of the curation process, including the potential for future science returns from the high resolutions scans.

Zeigler, R. A.↗

Using X-Ray Computed Tomography to Image Apollo Drive Tube 73002

The Apollo missions collected 382 kg of rock, regolith, and core samples from six locations on the nearside of the Moon. Today, just over 84% by mass of the Apollo collection remains in pristine condition within the curation facility at Johnson Space Center. Most Apollo samples have been well characterized, however there are several types of samples that have remained wholly or largely unstudied since their return, and/or that have been curated under special conditions. These sample types are: (1) unopened samples sealed under vacuum on the Moon; (2) unopened (but unsealed) drive tubes; (3) Apollo 17 samples frozen shortly after their return; and (4) Apollo 15 samples opened and stored in a helium atmosphere since their return. NASA solicited proposals for the Apollo Next Generation Sample Analysis Program (ANGSA), and 9 teams were selected to study: (1) unsealed, unopened drive tube 73002; (2) sealed, unopened drive tube 73001 (paired with 73002); and (3) a subset of the frozen and He-purged samples [1]. The first sample opened as part of the ANGSA program was drive tube 73002. This was originally a ~30 cm long, 4 cm diameter drive tube collected on a landslide deposit near Lara Crater at the Apollo 17 landing site. It was part of a ~60 cm long double drive tube collected, and the bottom half of the tube (73001) was sealed under vacuum on the Moon [2]. Prior to opening sample 73002, the sample was imaged with a high resolution X-ray Computed Tomography (XCT) scan of the entire tube. Additional XCT scans have been made of “large” clasts removed from the core as part of the dissection process [3]. Here we present the whole tube and close-up XCT data from 73002, and talk about the utility of the scans as part of the curation process, including the potential for future science returns from the high resolutions scans.

R A Zeigler↗

The Genesis Non-flight Implant Sample Archive

Introduction: For the past 20 years the Genesis science community has been analyzing solar wind captured in sample collectors now curated at Johnson Space Center (JSC). Non-flown flight like collector materials (non-flight) considered reference material was always allocated by the JSC Genesis team for testing experimental procedures. However, to expedite Genesis science, distribution of non-flown reference materials were also provided by, D. S. Burnett (Genesis PI, Caltech) and A. J. G. Jurewicz (ASU) who were funded to aid all Genesis PIs by providing engineering test materials archived at Caltech pre-flight as well as ion implants into these materials [1]. Implanted materials (non-flight reference, flight-like semiconductor materials, and NIST SRM glass) received from the Caltech collection during a pilot project are being used to develop an accession procedure for the purpose of making these materials available to researchers in the future. These implant samples are important because they can be utilized as measurement standards [2]. This is a work in progress and this work reports progress in documentation and imaging. The pilot program received 63 implanted reference materials with known implanted ions. Genesis Non-Flight Database: The existing non-flight material database was modified to allow for the description of implant sessions and to allow correlation of the implant information to individual specimens. An implant session has three main components: a known ion, a known dose (ion/cm2), and a known energy (keV). Procedure: Laboratory Procedure. Upon receiving non-flight analytical standards (independently calibrated implants), or implant samples (nominally calibrated implants), they will be imaged in their current containers to capture any information written before being brought into the Genesis Sample Laboratory. Documentation is provided by the donor and used for correlating the data between the implant sessions and implant samples. This information is not verified by Genesis Curation. The microscope used to image samples in the laboratory has an automated stage with a stainless-steel plate to hold samples and a camera that utilizes the software Surveyor to take photomosaic images of the implant samples. The camera attached to the microscope is a QICAM High Performance IEEE 1394 FireWire Digital CCD Camera. This setup, shown in Fig. 1, allows for photomosaics of small-scale samples with a 10x objective or 25x objective. Examples of images taken with this microscope setup are shown in Fig.2 and Fig. 3. After imaging, samples get stored in a clean fluoroware container or polypropylene vial depending on sample size for long term storage. Database Procedure. The Genesis Non-Flight Database currently allows for the addition of implant sessions. The required fields for adding an implant session are the implant session ion, the implant session dose, and the implant session energy. There are also fields for the date and vendor of the sessions as well as a field for if the session was calibrated or not. Once an implant session has been inserted into the database, additional documentation from the session can be attached. The database will automatically assign an Implant Session Number. After inserting an implant session, the implanted samples can be inserted. The fields required for an implant sample are the material of the sample and the size of the sample. The Generic name for the implant sample with be automatically generated by the database once it has been inserted. Implant sample numbers begin with 3X to distinguish from flown Genesis sample numbers. There are also fields for processors to add the location (room, cabinet, tray) and container type (fluoroware, polypropylene vial). Once an implant session and an associated sample(s) have been inserted into the database the sample can be then tied to all associated implant sessions. An implant sample can be tied to multiple implant sessions. Future work: For curation, future work consists of creating an online catalogue for implant samples as well as updating the database with more samples as we receive them. For the scientific community these implant samples have many potential uses. For Genesis research, they are directly applicable for testing and, in some cases, as analytical standards. For non-Genesis research, implants (especially calibrated implants) could be used for calibrating other, unique planetary materials [1].

C D Calva↗

Produced Water DNA Database (PW-DNA): Utilizing KBase to generate an environmental specific curated molecular database

The deep subsurface is estimated to host the majority of Earth’s microbial biomass yet remains one of the most challenging environments to access and study. One common approach to investigate these microbial communities is through the analysis of produced water from subsurface reservoirs, where researchers can assess water and gas chemistry along with molecular (DNA/RNA) sequence data. Advances in high-throughput sequencing have greatly expanded our understanding of these environments and their biotechnological potential. However, further progress requires large-scale, integrative meta-analyses across diverse datasets. To address this need, we developed the Produced Water-DNA (PW-DNA) Database, a curated, publicly available resource that consolidates microbial DNA/RNA sequences, geochemical data, and relevant metadata from in situ hydrocarbon environments such as coal beds, oil reservoirs, and natural gas systems. The PW-DNA database delivers three core benefits to the research community: (1) it improves data sharing by linking environmental microbial datasets with corresponding geochemical parameters, enabling more robust filtering and analysis; (2) it connects with complementary research databases to promote broader dissemination and interoperability; and (3) it supports technological innovation by serving as a resource for identifying microbial trends and exploring genetic potential. While individual studies have highlighted basin-specific microbial communities and functional redundancy in biogeochemical cycling, a comprehensive, system-wide perspective is needed to better understand connectivity and novelty across subsurface ecosystems. By designing the PW-DNA in the KBase platform, we provide a reproducible, visual framework for integrating large-scale genomic and geochemical data, enabling researchers to perform more informed analyses and experimental design. Ultimately, this resource enhances the ability to identify, characterize, and interpret microbial functions across diverse subsurface environments, thereby accelerating discovery in subsurface microbiology and biotechnology.

59 BASIC BIOLOGICAL SCIENCES↗

Missing microbial eukaryotes and misleading meta-omic conclusions

Meta-omics is commonly used for large-scale analyses of microbial eukaryotes, including species or taxonomic group distribution mapping, gene catalog construction, and inference on the functional roles and activities of microbial eukaryotes in situ. Here, we explore the potential pitfalls of common approaches to taxonomic annotation of protistan meta-omic datasets. We re-analyze three environmental datasets at three levels of taxonomic hierarchy in order to illustrate the crucial importance of database completeness and curation in enabling accurate environmental interpretation. We show that taxonomic membership of sequence clusters estimates community composition more accurately than returning exact sequence labels, and overlap between clusters can address database shortcomings. Clustering approaches can be applied to diverse environments while continuing to exploit the wealth of annotation data collated in databases, and selecting and evaluating these databases is a critical part of correctly annotating protistan taxonomy in environmental datasets. We argue that ongoing curation of genetic resources is crucial in accurately annotating protists in in situ meta-omic datasets. Moreover, we propose that precise taxonomic annotation of meta-omic data is a clustering problem rather than a feasible alignment problem.

59 BASIC BIOLOGICAL SCIENCES↗

International Agreement on Planetary Protection

The maintenance of a NASA policy, is consistent with international agreements. The planetary protection policy management in OSS, with Field Center support. The advice from internal and external advisory groups (NRC, NAC/Planetary Protection Task Force). The technology research and standards development in bioload characterization. The technology research and development in bioload reduction/sterilization. This presentation focuses on: forward contamination - research on the potential for Earth life to exist on other bodies, improved strategies for planetary navigation and collision avoidance, and improved procedures for sterile spacecraft assembly, cleaning and/or sterilization; and backward contamination - development of sample transfer and container sealing technologies for Earth return, improvement in sample return landing target assessment and navigation strategy, planning for sample hazard determination requirements and procedures, safety certification, (liaison to NEO Program Office for compositional data on small bodies), facility planning for sample recovery system, quarantine, and long-term curation of 4 returned samples.

Source record↗

FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation

Retrieval-augmented generation (RAG) has emerged as a promising paradigm for improving factual accuracy in large language models (LLMs). We introduce a benchmark designed to evaluate RAG pipelines as a whole, evaluating a pipelines ability to ingest several modalities of information. We present (1) a curated dataset of 93 questions designed to evaluate a pipeline's ability to ingest textual data, tables, images, multimodal data, and cross-document multimodal data; (2) a phrase-level recall metric for correctness; (3) a nearest-neighbor embedding classifier in an attempt to classify pipeline hallucinations; (4) a comparative evaluation of 2 pipelines built with open-source retrieval mechanisms and 4 closed-source foundational models; and (5) a third-party human evaluation of the alignment of our correctness and hallucination metrics. We find that closed-source pipelines significantly outperform open-source pipelines in both the correctness and halucination metrics, with a wider performance gap in questions relying on multimodal and cross-document information. We also find after a human evaluation of our correctness and hallucination metric compared with our questions and pipeline responses, average agreement was 4.62 for correctness 4.53 for hallucination detection on a 1-5 Likert scale with 5 being strongly agree with our determination.

Hildebrand, Samuel [ORNL] (ORCID:0009000465963104)↗

Educational Consortium for Energy-related Data Science & Computation in Building Engineering Programs

The project spearheaded by Pennsylvania State University aims to address the growing need for integrating energy-focused computation and data science into building engineering education. As the demand for energy-efficient building designs and operations increases, the educational sector must adapt to equip future engineers with the necessary skills. This initiative responds to this need by developing a consortium that unites multiple institutions to enhance curriculum development, dataset curation, and resource sharing, thereby ensuring students are well-prepared for the evolving energy sector. The primary goal of the project is to establish a consortium that will develop and disseminate educational materials and training programs focused on energy-related data science and computation. Key accomplishments include the creation of a beta website for resource sharing, the development of training programs and standalone modules, and the curation of datasets accessible to the public. This effort will culminate in a curriculum that incorporates advanced modeling technologies and data science skills into building engineering programs.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Reference Site Conditions for Floating Wind Arrays in the United States

Floating offshore wind farm design is highly site-specific, requiring detailed information about the specific conditions of a project area for realistic design studies. Unfortunately, publicly available site condition data for potential floating offshore wind project sites in the United States is scarce. To support U.S. offshore wind research, we developed reference site condition datasets, including metocean and seabed information, for four potential floating wind project areas in the U.S.: Humboldt Bay, Morro Bay, the Gulf of Maine, and the Gulf of Mexico. These datasets were compiled using publicly available data. Our metocean analysis, covering wind, waves, and surface currents, utilized measurement data from 2000 to 2020. Sources included the National Renewable Energy Laboratory’s National Offshore Wind Dataset for wind data, National Data Buoy Center buoys for wave data, and the High Frequency Radar Network for surface currents. These data were integrated into hourly time series used to compute extreme return periods up to 500 years, monthly statistics, and joint probability clusters for fatigue analysis. Soil conditions were evaluated using the usSEABED database and bathymetry grids were interpolated from the NCEI Digital Elevation Model Global Mosaic. In addition to providing curated reference site condition datasets for four U.S. areas, our assessment highlights the need for more publicly available metocean and soil condition data.

17 WIND ENERGY↗

The NASA Ames Research Center Institutional Scientific Collection: History, Best Practices and Scientific Opportunities

The NASA Ames Life Sciences Institutional Scientific Collection (ISC), which is composed of the Ames Life Sciences Data Archive (ALSDA) and the Biospecimen Storage Facility (BSF), is managed by the Space Biosciences Division and has been operational since 1993. The ALSDA is responsible for archiving information and animal biospecimens collected from life science spaceflight experiments and matching ground control experiments. Both fixed and frozen spaceflight and ground tissues are stored in the BSF within the ISC. The ALSDA also manages a Biospecimen Sharing Program, performs curation and long-term storage operations, and makes biospecimens available to the scientific community for research purposes via the Life Science Data Archive public website (https:lsda.jsc.nasa.gov). As part of our best practices, a viability testing plan has been developed for the ISC, which will assess the quality of archived samples. We expect that results from the viability testing will catalyze sample use, enable broader science community interest, and improve operational efficiency of the ISC. The current viability test plan focuses on generating disposition recommendations and is based on using ribonucleic acid (RNA) integrity number (RIN) scores as a criteria for measurement of biospecimen viablity for downstream functional analysis. The plan includes (1) sorting and identification of candidate samples, (2) conducting a statiscally-based power analysis to generate representaive cohorts from the population of stored biospecimens, (3) completion of RIN analysis on select samples, and (4) development of disposition recommendations based on the RIN scores. Results of this work will also support NASA open science initiatives and guides development of the NASA Scientific Collections Directive (a policy on best practices for curation of biological collections). Our RIN-based methodology for characterizing the quality of tissues stored in the ISC since the 1980s also creates unique scientific opportunities for temporal assessment across historical missions. Support from the NASA Space Biology Program and the NASA Human Research Program is gratefully acknowledged.

ALSDA↗

NASA GeneLab Platform Utilized for Space Radiation Dosimetry Biological Response Compared to Radiation Ground Studies

Ionizing radiation from Galactic Cosmic Rays (GCR) is one of the major risk factors factor that will impact health of astronauts on extended missions outside the protective effects of the Earth's magnetic field. Currently there are gaps in our knowledge of the health risks associated with chronic low dose, low dose rate ionizing radiation, specifically ions associated with high (H) atomic number (Z) and energy (E). The NASA GeneLab project (genelab.nasa.gov) aims to provide a detailed library of Omics datasets associated with biological samples exposed to HZE. The GeneLab Data System (GLDS) includes datasets from both spaceflight and ground-based studies, a majority of which involved exposure to ionizing radiation. Recently GeneLab has also assessed radiation dosimetry data with omics datasets associated with samples flown to the International Space Station (ISS). The combination of the detailed information on radiation exposure for ground-based studies and curated dosimetry information for spaceflight experiments allows GeneLab to be the first comprehensive Omics database for space related research from which an investigator can generate hypotheses to direct future experiments utilizing both ground and space biological radiation data. We demonstrate the usefulness of these datasets by analyzing multiple GeneLab datasets associated with both radiation ground-based studies and spaceflight studies. The radiation ground based studies we analyzed includes both in vivo and in vitro work with a range ions from protons to iron particles with doses from 0.1Gy to 2Gy. These datasets were compared to both in vivo and in vitro datasets from samples flown to the ISS and on shorter shuttle missions with total doses of 0.1 mGy to 30 mGys. From this analysis we were able to associate distinct biological signatures associating specific ions to specific biological response to radiation exposure in space. For example, we discovered radiation biological response related to cardiovascular effects from proton ground studies are the dominating response for samples related to cardiovascular effects on the ISS. With this work we will provide a summary of how different ions will impact different biological response in space and how this can be used in future studies to assess optimal ground experiments to simulate space radiation.

Beheshti, Afshin↗

Biospecimen Culling: Temporal RNA Integrity Analysis Across Spaceflight Missions Dating from 1985 to 2011

The Ames Life Science Data Archive (ALSDA) at NASA Ames Research Center is managed by the Space Biosciences Division and has been operational since 1993. The ALSDA is responsible for archiving information and biospecimens collected from life science spaceflight experiments and matching ground control experiments. They are stored in the Ames biobank, which is located in the Biospecimen Storage Facility (BSF). The ALSDA also manages a Biospecimen Sharing Program, performs curation and long-term storage operations, and makes biospecimens available to the scientific community for research purposes via the Life Science Data Archive public website (https:lsda.jsc.nasa.gov). The BSF maintains both fixed and frozen spaceflight and ground tissues, collected from recent and past spaceflight missions. Due to the ever increasing demand for space to preserve current and future flight biospecimens, the ALSDA has initiated the development of a culling plan for biospecimens currently stored in the BSF. Culling enables the ALSDA to assess the quality of archived samples, and supports the development of standardized culling procedures that improve the operational efficiency of the BSF. The culling plan focuses on generating disposition recommendations for samples in the BSF, and currently is based on measuring ribonucleic acid (RNA) integrity number (RIN). The culling process includes (1) sorting and identification of candidate samples for RIN analysis, (2) completion of RIN analysis on select samples, and (3) development of disposition recommendations for specimens based on the RIN values. Furthermore, our approach allows for unique scientific opportunities, including development of a RIN-based methodology for culling, and temporal assessment of the quality of the tissues that have been stored in BSF since the 1980s. Results of this work will also support NASA open science initiatives.

biospecimen↗

Opening doors to physical sample tracking and attribution in Earth and environmental sciences

Physical samples and their associated data and metadata underpin scientific discoveries across disciplines and can enable new science when appropriately archived. However, there are significant gaps in current practices and infrastructure that prevent accurate provenance tracking, reproducibility, and attribution. For most samples, descriptive metadata are often sparse, inaccessible, or absent. Samples and associated data and metadata may also be scattered across numerous physical collections, data repositories, laboratories, data files, and papers with no clear linkage or provenance tracking as new information is generated over time. The Earth Science Information Partners (ESIP) Physical Samples Curation Cluster has therefore developed guidance for scientific authors on ‘Publishing Open Research Using Physical Samples.’ This involved synthesizing existing practices, gathering community feedback, and assessing real-world examples. We identified improvements needed to enable authors to efficiently cite and link Earth science samples and related data, and track their use. Our goal is to help improve discoverability, interoperability, and reuse of physical samples, and associated data and metadata. Though primarily focused on the needs of Earth and environmental sciences, these guidelines are broadly applicable.

58 GEOSCIENCES↗