Search NASASearch

SEARCH · Search NASA

Results for “metadata standards”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Metadata Standards for the NSE: Extended Field Standards

This standard presents a set of optional metadata fields for managed digital objects within the Nuclear Security Enterprise (NSE) and provides a deeper look at data representation in metadata by looking at the representation of 1) Records Management required metadata, and 2) common representations of technical/scientific data. Metadata standardization is a critical enabler for effectively sharing data, documents, and other digital objects between NSE sites, and for tracing the digital thread at the object level. Standardization is necessary for both schemas and vocabularies, meaning that both field standards and value standards must be specified. This document serves as a complementary field standard, recommending an optional set of fields that should be uniformly built for all managed digital objects within the NSE. This document specifically focuses on extending the shared discovery layer defined in the first white paper by introducing additional descriptive and data representation fields that improve cross-site search and interpretation.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION

Metadata Standards for the NSE: Core Fields

This standard presents a core set of metadata fields required for each managed digital object within the Nuclear Security Enterprise (NSE). Metadata standardization is a critical enabler for two primary objectives: 1) effectively sharing data, documents, and other digital objects between NSE sites; and 2) supporting digital engineering through the digital thread at the object level. Standardization is necessary for both schemas and vocabularies, meaning that both field standards and value standards must be specified. This document serves as a foundational field standard, recommending a core set of fields that should be uniformly required for all managed digital objects within the NSE.

99 GENERAL AND MISCELLANEOUS

NEPATEC v2.0: Standardized Metadata and Text Corpus of National Environmental Policy Act Documents

The National Environmental Policy Act of 1969, as amended (NEPA), is a major environmental law in the United States, requiring Federal agencies to consider and document potential environmental impacts before deciding on a proposed action. Modernization of NEPA and permitting processes faces significant challenges due to the lack of standardized formats and interoperable systems for organizing and sharing NEPA-related information across agencies. Much of the information gathered during NEPA reviews is written into documents such as categorical exclusions, environmental assessments, and environmental impact statements, then filed in predominately independent agency file stores that may or may not be publicly accessible. The application of metadata and data standards, such as those recommended by the Council on Environmental Quality (CEQ), to NEPA documents offers a shared vocabulary and structure for key entities like projects, processes, and documents that can streamline information exchange and enhance collaboration across systems. In this work, we publicly release NEPATEC2.0, an expanded corpus of NEPA documents with associated metadata. NEPATEC2.0 encompasses approximately 120,000 documents from 60,000 projects prepared by more than 60 different agencies. Modeled to align with CEQ metadata standards, NEPATEC2.0 promotes consistency in environmental reviews and supports the ongoing effort to modernize permitting technologies by facilitating more transparent, efficient, and data-driven decision-making. Importantly, NEPATEC2.0 demonstrates the possibilities and limitations of large language model-based prompting to extract information from NEPA documents at scale.

54 ENVIRONMENTAL SCIENCES

PermitTEC v0.1: Standardized Metadata Corpus of NEPA Litigation Documents

The National Environmental Policy Act of 1969, as amended (NEPA), mandates that federal agencies assess and document potential environmental impacts before deciding on proposed actions. While significant progress has been made in cataloging and standardizing NEPA documents themselves, the legal challenges that frequently arise from these decisions remain poorly cataloged and largely inaccessible for systematic analysis. Litigation challenging NEPA compliance can substantially delay project timelines, reshape agency decision-making, and establish precedents that influence future environmental reviews — yet no standardized, machine-readable corpus exists that links litigation records to the NEPA projects they contest.

54 ENVIRONMENTAL SCIENCES

Data Cards for Standardized Metadata Across DOE-Aligned Data Initiatives: Toward Transparent, Interoperable, and Governed Dataset Documentation

As data-intensive research, advanced computing, and artificial intelligence become increasingly central to scientific and operational workflows, the need for consistent, transparent, and machine-actionable documentation has grown correspondingly. Multiple DOE-aligned communities—including Office of Science, Genesis Mission, American Science Cloud (AmSC), National Nuclear Security Administration (NNSA) stewardship and governance, and related cross-laboratory collaborations—have independently developed metadata practices to support discovery, access, reuse, repository deposit, and compliance.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION

Diagnosing the representation of surface and layered soil moisture in Earth system models

Surface soil moisture (mrsos) and vertically integrated soil moisture (mrsol) over the top 10 cm should, by definition, be physically consistent in Earth System Models (ESMs). However, an evaluation of nine CMIP6 models reveals substantial inconsistencies: in some models, mrsos and integrated mrsol agree globally; in others, they align only in specific regions; and in a few, they diverge across all grid cells. These discrepancies arise from a combination of factors, including metadata errors, inconsistent variable definitions, or diagnostic sequencing within the model. We demonstrate how such issues can lead to significant biases, even when both variables are present and seemingly well-defined. As model complexity increases and multi-model comparisons become more common, assumptions about variable equivalence may lead to flawed conclusions. This study highlights the need for routine consistency checks, improved metadata standards, and community-wide practices that ensure reliability of derived variables across ESM outputs, particularly in preparation for CMIP7.

Earth system models

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Genetics & Heredity

Challenges and Vision for Standardization of Biopolymer Data Sets for Machine Learning

Machine learning (ML) is transforming materials research, yet potential for biopolymer discovery remains constrained by fragmented data and nonstandardized reporting. Biopolymers differ significantly from synthetic polymers, requiring specialized approaches to represent their biosynthetic origins, hierarchical structures, and application-specific metrics. In this Perspective, we identify three core challenges limiting biopolymer representation: information encoding, data quality, and data sharing. We describe the most pressing issues and propose commensurate approaches to address each key challenge. Recommendations include the design and adoption of biopolymer-specific fingerprinting and representation frameworks, development of hybrid human-large language model (LLM) data extraction strategies, and expanding Findable, Accessible, Interoperable, Reusable (FAIR)-compliant repositories. We propose a robust foundation to define interoperable, high-quality data sets that capture the full context of biopolymer materials. Standardized metadata, shared ontologies, and community-driven infrastructure would enable scalable, reproducible workflows and accelerate the ML-driven development of biopolymers.

36 MATERIALS SCIENCE

The Zooplankton International Geospatial dataset: A global repository of spatiotemporal freshwater zooplankton community composition data from lakes and reservoirs to support ecological research

Zooplankton transfer substantial energy in aquatic food webs and are used as indicators of environmental change. Syntheses of zooplankton community dynamics globally require datasets that span a wide range of environmental gradients; however, these datasets are limited due to methodological differences across programs, taxonomic inconsistencies, and a lack of standardized metadata. To reconcile these challenges, we created the Zooplankton International Geospatial (ZIG) dataset, which includes original zooplankton, water physical and chemical variables, and lake morphometric data from 311 inland lakes and reservoirs. ZIG includes waterbodies ranging in size from 0.005 to 82,100 km2 and spanning broad latitudinal (−47.26 to 64.90) and longitudinal ranges (−165.04 to 176.53). Temporal coverage for individual waterbodies ranges between 1 and 60 yr with sampling frequency ranging from annually to weekly. With its extensive coverage and content, we consider ZIG to be a cornerstone for future investigations of global scale lake biodiversity change.

Figary, Stephanie [Cornell University, Ithaca, NY]

Hosting downscaled decision-relevant community data products in ESGF2-US

As regionally-relevant high-resolution Earth system data is increasingly relied upon across scientific, policy, and practitioner communities, there is an urgent need for coordinated and federated infrastructure to store, manage, standardize, and distribute decision-relevant community data products. Substantial effort is required to ensure that these products, which are often critical for regional impact assessments and decision-making, are findable, accessible, interoperable, and reusable. The Earth System Grid Federation US project (ESGF2-US) is addressing this challenge by expanding its open-source, distributed platform to support the hosting and dissemination of downscaled Earth system datasets. This expansion includes aligning new downscaled datasets with developing community standards for metadata and file structure, consistent with existing ESGF archives. This includes ensuring CF-compliance, applying CMORization where appropriate, and developing tools to streamline user access. In this paper, we highlight the technical and coordination work required to bring downscaled data into ESGF2-US and aim to inform the broader Earth system data user community about the growing availability and utility of these curated resources.

ESGF

Observational Data for Next-Generation Climate Model Evaluation: Requirements, Considerations, and Best Practices

Climate model simulations are an important source of information about our planet’s climate system and also enable informed decision-making under different future scenarios. As a new archive of results from the next generation of climate models is anticipated to become available with the Coupled Model Intercomparison Project phase 7 (CMIP7), the need to develop efficient and robust methods to evaluate models is paramount. Observations are an integral part of model evaluation, providing a means to quantify and understand the degree to which climate models can faithfully reproduce Earth system processes. Such analysis is critical for constraining climate projections, identifying areas of focus for model development, and assisting analysts in deciphering the utility of models for specific applications. Observations of Earth system come from a diversity of sources, span different space–time domains, and are produced by different communities, and each dataset features different data structures and formats, metadata standards, and its own unique uncertainties. Uncertainties in an observational dataset may stem from gaps in temporal and spatial coverage, instrumentation errors, or assumptions in retrieval and processing methods. How then does one ensure that observational data are ready for use and utilized in the most appropriate way for robust, rapid, and routine climate model evaluation? The CMIP7 Model Benchmarking Task Team with input from the broader climate modeling, model evaluation, and observational data communities present a vision and considerations for best practices toward the optimal and appropriate use of observational data to support next-generation climate model evaluation.

Climate models

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY

AI-Ready Data Pilot Project Report

The proliferation of artificial intelligence in scientific research has created an urgent need to define "AI-ready data" for researchers and, more importantly, provide resources to help them produce AI-ready data. At Pacific Northwest National Laboratory, we conducted a pilot study with three data scientists evaluating three CSV datasets from different scientific domains, followed by semi-structured interviews capturing assessment practices. Our findings reveal that AI-readiness evaluation is intuition-based, with practitioners asking "How fast can I go from raw data to my machine learning pipeline?" Data scientists consistently prioritized workflow efficiency, human interpretability, and quality stewardship signals. From these insights, we developed a practical evaluation framework comprising data requirements, metadata standards, and validation tests that provides actionable criteria for producing and curating AI-ready datasets, addressing the gap between theoretical understanding and practical implementation.

97 MATHEMATICS AND COMPUTING

FAIR Data Meets FAIR Software

Modern scientific research is increasingly defined by the interplay between data, software, and the workflows that connect them. Yet while the FAIR (Findable, Accessible, Interoperable, Reusable) principles have become foundational for scientific data stewardship, the same level of structure and expectation has only recently begun to extend to research software. This talk covers why and how FAIR principles are being applied to data and software to support data reuse. It outlines the gaps in current sharing norms, the growing federal emphasis on persistent identifiers and public access, and the opportunities created when datasets, computational workflows, code, and models are linked through rich, standardized metadata. Practical implementation pathways for the EIC and JLab communities are described, including datacards for structured dataset documentation and provenance-aware workflows. By aligning data lifecycle management with FAIR-aligned software practices, the scientific community can advance toward autonomous knowledge graphs, generative workflows, and high-quality, AI-ready scientific datasets.

McSpadden, Diana [Thomas Jefferson National Accele

A total of 19 months of daily weather logging on the US east coast: the WFIP3 event log

The Third Wind Forecast Improvement Project (WFIP3) is a multi-institutional field campaign designed to advance the understanding and prediction of the offshore atmospheric boundary layer along the US east coast. Extending from February 2024 through August 2025, WFIP3 combines long-term coastal and offshore measurements with targeted modeling and forecasting efforts. This data paper presents the WFIP3 event log, a curated record of 578 d of meteorological phenomena and field observations that complements the campaign's extensive high-frequency datasets. The event log provides both manually documented daily weather discussions and automatically derived indicators of atmospheric processes – including low-level jets, wind ramps, extreme wind veer, and weak wind conditions – based on observations from scanning lidars deployed at three coastal and offshore sites. The dataset offers structured metadata, standardized time and site identifiers, and consistent terminology to facilitate its integration with WFIP3's observational and modeling data products. The log supports diverse applications, from model evaluation and forecast verification to the selection of case studies on offshore boundary-layer dynamics. The WFIP3 event log is publicly available through the US Department of Energy's Wind Data Hub, providing the research community with a transparent and enduring contextual reference for the interpretation and use of WFIP3 measurements.

17 WIND ENERGY

The Zooplankton International Geospatial (ZIG) dataset: A global repository of spatiotemporal freshwater zooplankton community composition data to support ecological research

Zooplankton play critical roles in aquatic ecosystem function and food webs. Nevertheless, global syntheses of their abundance and community dynamics are challenging due to methodological differences across monitoring programs, taxonomic inconsistencies, and a lack of standardized metadata. To reconcile these challenges, we assembled, curated, validated, and harmonized the Zooplankton International Geospatial (ZIG) dataset, which includes co-located and contemporaneous zooplankton, water chemistry, and limnological data from 307 lakes and reservoirs. ZIG includes waterbodies from each major lake thermal region and range in size from 0.8-2,805,8600 hectares. Temporal coverage for individual waterbodies ranges between 1-60 years of data (median = 4 years) with sampling from once annually to weekly. ZIG is publicly available and can be used to understand freshwater biodiversity change and its drivers at unprecedented scales, and we consider it to be a cornerstone for future investigations of freshwater biology, chemistry, and ecology.

Figary, Stephanie [Cornell University, Ithaca, NY]

Retrospective on decadal progress of the NOAA/NPS ocean noise reference station network

The National Oceanic and Atmospheric Administration (NOAA), in partnership with the U.S. National Park Service (NPS), established the Ocean Noise Reference Station Network (NRS) in 2014 as a foundational component of NOAA’s Ocean Noise Strategy. This long-term effort aims to characterize baseline ocean ambient sound conditions across diverse marine environments and to inform management of noise impacts on protected species and habitats within U.S. waters. The NRS is now composed of 13 autonomous passive acoustic monitoring stations strategically positioned across the U.S. Exclusive Economic Zone (EEZ), extending from Arctic regions to tropical waters in depths ranging from 33 to 4,790 m. These locations include several National Marine Sanctuaries and National Parks, such as the recently designated Chumash Heritage National Marine Sanctuary off the coast of California. Each station is equipped to continuously sample low-frequency underwater sound at five kHz, enabling the detection of anthropogenic, geophysical, and biological acoustic signals. To date the network has sampled over 72 years of calibrated acoustic data. The spatial breadth and consistent methodology of the NRS allow for comparative acoustic assessments across diverse marine ecosystems. In addition to applied research functions, the NRS has served as a platform for education and training, offering opportunities for students to develop skills for marine science and data analysis. Looking forward, the NRS project team is focused on network expansion, improved data delivery, and broader integration with collaborative scientific initiatives. NRS recordings are being archived in partnership with NOAA’s National Centers for Environmental Information to enhance accessibility and long-term utility. Efforts are underway to develop standardized metadata and summary products to accompany raw audio files, making the data more usable for a wide range of stakeholders in the ocean science community. The NRS is evolving into a fully integrated national framework for ocean sound monitoring that supports scientific inquiry, management decision-making, national security interests, and public engagement with ocean acoustic environments.

Long-term monitoring