Search NASA⌕ Search

SEARCH · Search NASA

Results for “metadata validation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

WGISS-45 International Directory Network (IDN) Report

The objective of this presentation is to provide IDN (International Directory Network) updates on features and activities to the Committee on Earth Observation Satellites (CEOS) Working Group on Information Systems and Services (WGISS) and provider community. The following topics will be will be discussed during the presentation: Transition of Providers DIF-9 (Directory Interchange Format-9) to DIF-10 Metadata Records in the Common Metadata Repository (CMR); GCMD (Global Change Master Directory) Keyword Update; DIF-10 and UMM-C (Unified Metadata Model-Collections) Schema Changes; Metadata Validation of Provider Metadata; docBUILDER for Submitting IDN Metadata to the CMR (i.e. Registration); and Mapping WGClimate Essential Climate Variable (ECV) Inventory to IDN Records.

WGISS↗

Evolution in Metadata Quality: Common Metadata Repository's Role in NASA Curation Efforts

Metadata Quality is one of the chief drivers of discovery and use of NASA EOSDIS (Earth Observing System Data and Information System) data. Issues with metadata such as lack of completeness, inconsistency, and use of legacy terms directly hinder data use. As the central metadata repository for NASA Earth Science data, the Common Metadata Repository (CMR) has a responsibility to its users to ensure the quality of CMR search results. This poster covers how we use humanizers, a technique for dealing with the symptoms of metadata issues, as well as our plans for future metadata validation enhancements. The CMR currently indexes 35K collections and 300M granules.

metadata quality↗

Syntactic and Semantic Validation without a Metadata Management System

The ability to maintain quality information is essential to securing the confidence in any system for which the information serves as a data source. NASA's Global Change Master Directory (GCMD), an online Earth science data locator, holds over 9000 data set descriptions and is in a constant state of flux as metadata are created and updated on a daily basis. In such a system, the importance of maintaining the consistency and integrity of these-metadata is crucial. The GCMD has developed a metadata management system utilizing XML, controlled vocabulary, and Java technologies to ensure the metadata not only adhere to valid syntax, but also exhibit proper semantics.

Pollack, Janine↗

Livewire: A Model Platform for Data Quality Assessment and AI Readiness Across DOE Missions

High-quality, well-governed data is essential for accelerating discovery and achieving operational excellence across DOE and national laboratory missions. The Livewire Data Platform is a DOE-supported platform that offers automated assessments of data quality, standardization, provenance, and Artificial Intelligence (AI) readiness. It allows researchers and data practitioners to systematically and easily evaluate datasets against established governance criteria and prepare them for advanced analytics. Livewire addresses critical challenges in DOE's data ecosystem with integrated capabilities for metadata validation, provenance tracking, and schema alignment. This platform's automated workflows assist users in identifying data quality gaps, enhancing interoperability between datasets collected from various stakeholders, and ensuring compliance with DOE data standards, all while reducing manual curation efforts. Additionally, we will discuss its AI readiness framework, which is being developed to prepare datasets for training models, developing advanced analytic tools, and machine learning applications. Using some of the more than one hundred tabular datasets on Livewire, processed with this open-source methodology, we will demonstrate how Livewire can serve as a model for scalable, standards-driven data management. This approach provides a pathway to leverage existing and future datasets within the DOE, boosting innovation and efficiency across national laboratories.

33 - ADVANCED PROPULSION SYSTEMS↗

Dronebase Photovoltaic (PV) Fleet Imagery Quantitative Evaluation (CRADA CRD-22-22941 Final Report)

Combine the Dronebase aerial imagery with corresponding sites in the NLR Photovoltaic (PV) Fleets database. By combining these two data sources in an aggregated, anonymized fashion, we can perform the following analyses: quantifying power loss due to outages caused by stuck trackers, string outages, and shading/snow, validate site metadata, including tilt and azimuth, and correlate.

14 SOLAR ENERGY↗

Validation of ADAR System 5500 Digital Imagery: Delivery Task Order #1, Task Request #857 - Brookings, SD

This work was performed under NASA's Verification and Validation Program as an independent check of data supplied by Positive Systems, Inc. through the Earth Science Enterprise's Scientific Data Purchase (SDP) Program. This document serves as the basis for reporting results associated with validation of multispectral imagery according to the specifications of contract NAS 13-98049. The validation was performed under the Positive Systems Imaging System Validation Work Instruction CRSP-WI-28: Spectral registration, spatial resolution, endlaps, sidelaps, and image quality were evaluated. The validation was proceded by Shipment Verification, as described in the Work Instruction CRSP-WI-22: Every image was passed through an automatic ingest verification and thumbnail review process to identify omissions, problems with media integrity, and gross errors in data quality. Validation of metadata files is not within the scope of this report, but it was performed separately.

Blonski, Slawomir↗

Quantifying Error in Photovoltaic Installation Metadata: Preprint

In this research, we quantify the level of metadata error for a fleet of 2860 photovoltaic (PV) systems, using metadata values provided by fleet owners. Using satellite imagery and time series analysis techniques available in open-source Python packages Panel-Segmentation and PVAnalytics, respectively, we evaluate the accuracy of PV system metadata such as location, azimuth, tilt, and mounting configuration (fixed tilt vs. tracking). We find that approximately 75% of provided latitude-longitude coordinates are within 190 meters of the actual solar installation. We were unable to link 7.8% of latitude-longitude coordinates to any solar installation via satellite imagery analysis. We evaluate the level of error in owner-provided mounting configuration (fixed tilt vs. single-axis tracking), finding only 8 systems with an incorrect mounting configuration. When evaluating azimuth and tilt parameters, we find that approximately 64% of the data is correct, with data for 860 systems (approximately 30%) not provided by system owners. To illustrate the importance of having correct solar metadata, we evaluate how incorrect metadata affects solar performance estimates by modeling system AC energy output at ground-truth vs. incorrect latitude-longitude coordinates, mounting configurations, and azimuth-tilt configurations. Energy output estimates can vary significantly if incorrect metadata parameters are used, with incorrect mounting configuration leading to the largest discrepancy with over 20% variation in expected energy output.

azimuth↗

Strategies for community-sourced biocuration in bioinformatics: a case study on MIBiG 4.0

Biocuration is essential to transform molecular sequence data into standardized, machine-readable resources. Such curated datasets enable comparative analysis, predictive modeling, and data integration across bioinformatics platforms. While professional biocuration is resource-intensive and usually limited to institutional settings, community-driven approaches can mobilize large-scale annotation of specialized datasets and are more resilient to disruptions in scientific funding. Here, we present a model for community-powered curation applied to the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) repository. Through a framework of workflows for metadata capture, annotation validation, and contributor coordination, the MIBiG 4.0 initiative recruited 267 scientists across 178 institutions from 33 countries, volunteering an estimated 4000 h of work. These efforts expanded the MIBiG repository by 22% and enhanced its usability in downstream molecular data analyses in comparative genomic analyses, natural product discovery, and machine learning applications. We provide strategies and actionable lessons for adopting this model, supporting the sustainability of curated bioinformatics resources central to nucleic acid research and related fields.

biocuration↗

AI-Ready Data Pilot Project Report

The proliferation of artificial intelligence in scientific research has created an urgent need to define "AI-ready data" for researchers and, more importantly, provide resources to help them produce AI-ready data. At Pacific Northwest National Laboratory, we conducted a pilot study with three data scientists evaluating three CSV datasets from different scientific domains, followed by semi-structured interviews capturing assessment practices. Our findings reveal that AI-readiness evaluation is intuition-based, with practitioners asking "How fast can I go from raw data to my machine learning pipeline?" Data scientists consistently prioritized workflow efficiency, human interpretability, and quality stewardship signals. From these insights, we developed a practical evaluation framework comprising data requirements, metadata standards, and validation tests that provides actionable criteria for producing and curating AI-ready datasets, addressing the gap between theoretical understanding and practical implementation.

97 MATHEMATICS AND COMPUTING↗

A Community Convention for Ecological Forecasting: Output Files and Metadata Version 1.0

This paper summarizes the open community conventions developed by the Ecological Forecasting Initiative (EFI) for the common formatting and archiving of ecological forecasts and the metadata associated with these forecasts. Such open standards are intended to promote interoperability and facilitate forecast communication, distribution, validation, and synthesis. For output files, we first describe the convention conceptually in terms of global attributes, forecast dimensions, forecasted variables, and ancillary indicator variables. We then illustrate the application of this convention to the two file formats that are currently preferred by the EFI, netCDF (network common data form), and comma-separated values (CSV), but note that the convention is extensible to future formats. For metadata, EFI's convention identifies a subset of conventional metadata variables that are required (e.g., temporal resolution and output variables) but focuses on developing a framework for storing information about forecast uncertainty propagation, data assimilation, and model complexity, which aims to facilitate cross-forecast synthesis. The initial application of this convention expands upon the Ecological Metadata Language (EML), a commonly used metadata standard in ecology. To facilitate community adoption, we also provide a Github repository containing a metadata validator tool and several vignettes in R and Python on how to both write and read in the EFI standard. Lastly, we provide guidance on forecast archiving, making an important distinction between short-term dissemination and long-term forecast archiving, while also touching on the archiving of code and workflows. Overall, the EFI convention is a living document that can continue to evolve over time through an open community process.

Michael C. Dietze↗

Automated annotation of scientific texts for ML-based keyphrase extraction and validation

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lack the essential metadata required for researchers to find, curate, and search them effectively. The lack of metadata poses a significant challenge in the utilization of these data sets. Machine learning (ML)–based metadata extraction techniques have emerged as a potentially viable approach to automatically annotating scientific data sets with the metadata necessary for enabling effective search. Text labeling, usually performed manually, plays a crucial role in validating machine-extracted metadata. However, manual labeling is time-consuming and not always feasible; thus, there is a need to develop automated text labeling techniques in order to accelerate the process of scientific innovation. This need is particularly urgent in fields such as environmental genomics and microbiome science, which have historically received less attention in terms of metadata curation and creation of gold-standard text mining data sets. In this paper, we present two novel automated text labeling approaches for the validation of ML-generated metadata for unlabeled texts, with specific applications in environmental genomics. Our techniques show the potential of two new ways to leverage existing information that is only available for select documents within a corpus to validate ML models, which can then be used to describe the remaining documents in the corpus. The first technique exploits relationships between different types of data sources related to the same research study, such as publications and proposals. The second technique takes advantage of domain-specific controlled vocabularies or ontologies. In this paper, we detail applying these approaches in the context of environmental genomics research for ML-generated metadata validation. Our results show that the proposed label assignment approaches can generate both generic and highly specific text labels for the unlabeled texts, with up to 44% of the labels matching with those suggested by a ML keyword extraction algorithm.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Location Identifiers, Metadata, and Map for Field Measurements at the East-Taylor Watershed Community Observatory, Colorado, USA (Version 3.3)

This dataset contains identifiers, metadata, and a map of the locations where field measurements have been conducted at the East-Taylor Watershed Community Observatory located in the Upper Colorado River Basin, United States. This is version 3.3 of the dataset and replaces the prior version 3.2 (see below for details on changes between the versions). Dataset description: The East River-Taylor Watershed is the primary field site of the Watershed Function Scientific Focus Area (WFSFA) and the Rocky Mountain Biological Laboratory. Researchers from several institutions generate highly diverse hydrological, biogeochemical, climate, vegetation, geological, remote sensing, and model data at the East-Taylor Watershed in collaboration with the WFSFA. Thus, the purpose of this dataset is to maintain an inventory of the field locations and instrumentation to provide information on the field activities in the East-Taylor Watershed and coordinate data collected across different locations, researchers, and institutions. The dataset contains (1) a README file with information on the various files, (2) three csv files describing the metadata collected for each surface point location, plot and region registered with the WFSFA, (3) csv files with metadata and contact information for each surface point location registered with the WFSFA, (4) a csv file with with metadata and contact information for plots, (5) a csv file with metadata for geographic regions and sub-regions within the watershed, (6) a compiled xlsx file with all the data and metadata which can be opened in Microsoft Excel, (7) a kml map of the locations plotted in the watershed which can be opened in Google Earth, (8) a jpg image of the kml map which can be viewed in any photo viewer, and (9) a zipped file with the registration templates used by the SFA team to collect location metadata. The zipped template file contains two csv files with the blank templates (point and plot), two csv files with instructions for filling out the location templates, and one compiled xlsx file with the instructions and blank templates together. Additionally, the templates in the xlsx include drop down validation for any controlled metadata fields. Persistent location identifiers (Location_ID) are determined by the WFSFA data management team and are used to track data and samples across locations. Dataset uses: This location metadata is used to update the Watershed SFA’s publicly accessible Field Information Portal (an interactive field sampling metadata exploration tool; https://wfsfa-data.lbl.gov/watershed/), the kml map file included in this dataset, and other data management tools internal to the Watershed SFA team. Version Information: The latest version of this dataset publication is version 3.3. This version contains 167 new point locations, 1 new plot, and 2 new geographic regions. Overall, there are a total of 1439 point locations, 75 plots, and 54 geographic regions. Additionally, the kml map of locations and image now includes two boundaries (Upper Ohio Creek (UO) and Carbon Creek (CA)) outside of the East River watershed (USGS HUC-10) and accompanying stream network that represents areas of focus. Refer to methods for further details on the version history. This dataset will be updated on a periodic basis with new measurement location information. Researchers interested in having their East-Taylor Watershed measurement locations added to this list should reach out to the WFSFA data management team at wfsfa-data@googlegroups.com. Acknowledgments: Please cite this dataset if using any of the location metadata in other publications or derived products. If using the location metadata for the 2018 NEON hyperspectral campaign, additionally cite Chadwick et al. (2020). doi:10.15485/1618130. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Part of this work was performed at SLAC Accelerator Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-76SF00515.

2018 NEON and 2025 CHESS Campaigns↗

Master Metadata Repository and Metadata-Management System

A master metadata repository (MMR) software system manages the storage and searching of metadata pertaining to data from national and international satellite sources of the Global Ocean Data Assimilation Experiment (GODAE) High Resolution Sea Surface Temperature Pilot Project [GHRSSTPP]. These sources produce a total of hundreds of data files daily, each file classified as one of more than ten data products representing global sea-surface temperatures. The MMR is a relational database wherein the metadata are divided into granulelevel records [denoted file records (FRs)] for individual satellite files and collection-level records [denoted data set descriptions (DSDs)] that describe metadata common to all the files from a specific data product. FRs and DSDs adhere to the NASA Directory Interchange Format (DIF). The FRs and DSDs are contained in separate subdatabases linked by a common field. The MMR is configured in MySQL database software with custom Practical Extraction and Reporting Language (PERL) programs to validate and ingest the metadata records. The database contents are converted into the Federal Geographic Data Committee (FGDC) standard format by use of the Extensible Markup Language (XML). A Web interface enables users to search for availability of data from all sources.

Armstrong, Edward↗

Metadata for a systematic description of signal data

This chapter aims to provide a comprehensive overview of metadata types that may be useful during system design, optimization, and automation. Metadata are grouped into three main categories: (a) metadata describing signal generation, (b) metadata describing signal quality, and (c) contextual information in the form of annotations. Each of these categories is introduced and explained in three separate sections. Importantly, this chapter mainly answers what is considered metadata. To a lesser degree, recommendations are made regarding the selection of metadata for long-term storage. Chapter 4 will explain where and how to store metadata. Chapters 5 and 6 explain how to collect certain metadata through dedicated sensor validation tests (Chapter 5) or algorithmic analysis (Chapter 6).

Alferes, Janelcy↗

Structuring and storing signals with their metadata: practical considerations

This chapter aims to provide a comprehensive overview on structuring signal data and their metadata, highlighting key considerations for optimal management and storage. Specifically, it focuses on: (a) the relevance of data organization; (b) what data to store and what to keep; and (c) data management methods. Thus, this chapter answers where and how to store metadata efficiently. Chapter 3 explains what is considered metadata. Chapters 5 and 6 present how to collect certain metadata through dedicated sensor validation tests (Chapter 5) or algorithmic analysis (Chapter 6).

Nicolaï, Niels↗

pyQuARC: Open Source Library for Earth Observation Metadata Quality Assessment

Metadata quality is essential to effective data discovery and has become increasingly vital as more Earth Science data sets become available. The Common Metadata Repository (CMR) hosts metadata describing NASA’s Earth Observation data products, which are archived across 12 Distributed Active Archive Centers (DAACs). The Analysis and Review of CMR (ARC) Team, located at Marshall Space Flight Center, conducts metadata quality assessments to ensure that these data products are discoverable, accessible, and usable. To achieve these goals, the ARC team has developed a metadata quality assessment framework to evaluate metadata completeness, correctness, and consistency. ARC uses a combination of manual and automated methods to assess these three components and identify areas of improvement; the team then collaborates with the DAACs to resolve any findings. To streamline this process, ARC is currently developing a host of scripts, known as pyQuARC, to automate metadata quality assessments as much as possible. pyQuARC is an open source library for Earth Observation Metadata Quality Assessment, and the tool utilizes ARC’s metadata quality assessment framework to make basic validation checks, pinpoint inconsistencies between dataset-level (i.e. collection) and file-level (i.e. granule) metadata, and identify opportunities for more descriptive and robust information. Since pyQuARC is also customizable, other users can make modifications as needed, and future metadata standards can also be implemented. Once pyQuARC is fully developed, it will support multiple schema types to serve the broader EOSDIS metadata community. This presentation will provide an overview of pyQuARC and its process of development while showcasing the tool’s valuable features and uses.

Jenny Wood↗

Atomistic Simulation of Glasses and Amorphous Materials: Challenges and Opportunities for the Next Decade

Atomistic simulations have become indispensable tools for understanding glass structure, dynamics, and properties, yet persistent challenges limit their predictive power. This perspective examines three interconnected issues, namely glass formation procedures, interatomic potential development, and machine learning applications, which emerged from the 5th International Workshop on Challenges of Atomistic Simulations of Glasses and Amorphous Materials. We identify convergent community priorities for (i) standardized validation protocols, (ii) curated benchmark datasets with complete metadata, and (iii) open repositories for glasses. A systematic was forward is provided by a hierarchical validation framework for assessing the structural fidelity, property prediction, and behavioral realism of simulation techniques. Looking ahead, transformative advances are promised by the fusion of classical techniques with machine learning based approaches, for instance, by integrating swap Monte Carlo with machine-learning (ML) potentials, leveraging foundation models through transfer learning, and finetuning ML potentials with experimental data. Progress depends on the community committing to validated models, reproducible protocols, and sustained data sharing.

Krishnan, N. M. Anoop↗