Search NASASearch

SEARCH · Search NASA

Results for “metadata evaluation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

GHRSST-14 DAS-TAG Report

The DAS-TAG provides the informatics and data management expertise in emerging information technologies for the GHRSST community. It provides expertise in data and metadata formats and standards, fosters improvements for GHRSST data curation, experiments with new data processing paradigms, and evaluates services and tools for data usage. It provides a forum for producer and distributor data management issues and coordination.

data processing

Development of a Discrepancy Checker for the Digital Twin in a Supervisory Control System for a Thermal Energy Delivery System

Defined as a virtual representation of a physical object, process, or service, and used to support real-world decision-making, a digital twin (DT) can be utilized to combine classical and novel frameworks in sensors, state predictions, and multi-input/multi-output systems, and to enable optimal autonomous operations. However, a DT’s usefulness largely depends on its ability to adequately mirror the state of its physical counterpart, and this adequacy should be reflected by the level of uncertainty in the underlying simulation models when estimating and predicting quantities of interest (QOIs). Moreover, simulation models in a DT may involve multiple fidelities of representations—ranging from physics-based models to data-driven ones—but classical uncertainty quantification (UQ) methods struggle to handle numerous uncertainty sources, nor are they designed for real-time applications. This work presents a UQ-based discrepancy checking and diagnosis tool for a DT-based supervisory control system applied to a thermal energy delivery system (TEDS) at Idaho National Laboratory. The discrepancy checker was developed using metadata from an automated DT development process, and these metadata included different combinations of physical model forms and model parameters, training data and hyperparameters for surrogate models, and design parameters for supervisory control systems. Next, correlations between the uncertainty results and the metadata were established and then applied to the DT operations. The discrepancy checker evaluates the discrepancies between model predictions from virtual and sensor measurements and backtraces them to the corresponding major sources of uncertainty. The discrepancy checker showed reasonable performance in detecting discrepancies and diagnosing sources of uncertainty in testing scenarios.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Community Involvement in Enhancing the Global Change Master Directory (GCMD) Controlled Vocabularies (Keywords)

NASA's Global Change Master Directory (GCMD) develops and expands a hierarchical set of controlled vocabularies (keywords) covering the Earth sciences and associated information (data centers, projects, platforms, instruments, etc.). The purpose of the keywords is to describe Earth science data and services in a consistent and comprehensive manner, allowing for the precise searching of metadata and subsequent retrieval of data and services. The keywords are accessible in a standardized SKOSRDFOWL representation and are used as an authoritative taxonomy, as a source for developing ontologies, and to search and access Earth Science data within online metadata catalogues. The keyword development approach involves: (1) receiving community suggestions, (2) triaging community suggestions, (3) evaluating the keywords against a set of criteria coordinated by the NASA ESDIS Standards Office, and (4) publication/notification of the keyword changes. This approach emphasizes community input, which helps ensure a high quality, normalized, and relevant keyword structure that will evolve with users changing needs. The Keyword Community Forum, which promotes a responsive, open, and transparent processes, is an area where users can discuss keyword topics and make suggestions for new keywords. The formalized approach could potentially be used as a model for keyword development.

governance

Commercial Smallsat Data Acquisition Program: Airbus U.S. Synthetic Aperture Radar Quality Assessment Summary

Quality assessment of the Airbus X-band Synthetic Aperture Radar (SAR) satellite products was conducted by the Commercial Smallsat Data Acquisition (CSDA) program’s radar subject matter experts, following the Joint NASA/ESA (European Space Agency) assessment draft guidelines. All three Airbus SAR spacecraft (TerraSAR-X, TanDEM-X, and PAZ) are based on the TerraSAR-X platform, and each have an active phased array antenna that is 4.8 x 0.7 m in the along-track and cross-track dimensions, respectively. TerraSAR-X and TanDEM-X are in a helical orbit, creating a bistatic imaging geometry, in addition to being capable of independent monostatic observations. The PAZ mission follows TerraSAR-X and TanDEM-X in the same 11-day orbit with a 5.5-day lag. TerraSAR-X and TanDEM-X are designed, developed, and operated through a Public-Private Partnership, while PAZ is a dual-use mission (civil and defense agencies), funded and owned by the Spanish Ministry of Defense and managed by Hisdesat (Hisdesat Servicios Estratégicos, S.A.), a Spanish private communications company. The assessment presented in this document is divided into two main parts: documentation review and the assessment of test datasets. The documentation review in sections 2.1 through 2.4 includes the assessment of the Airbus documentation provided to the CSDA evaluation team. The grading of these documents is given in columns 1-4 of the maturity matrix shown in section 1.1. Section 2.5 summarizes the evaluation performed by NASA using the data purchased through the CSDA program. The grading for this is given in the last column of the maturity matrix. Section 3 provides more detailed explanations on the methods and the results of the data analysis performed by NASA. Only the documents provided by Airbus for the evaluation were considered for the review. Additional documentation with more detailed description of the calibration and validation procedures may be available online but were not considered for this evaluation. The product information provided in the available documentation (RD-1, RD-2) and the product metadata together provided adequate information to work with the data. The product details in the metadata included the required information to work with the data in the common XML file format. Metrological traceability documentation was not provided to CSDA. All relevant characterization of the SAR system and data were provided, and the metadata include all relevant ancillary information. Documentation provided to CSDA included limited pre-flight and post-launch calibration information.

Batuhan Osmanoglu

Soil biogeochemical properties and metrics of tree-mycorrhizal dominance for a 25-Ha forest in South Central Indiana, USA.

This data package contains a dataset used in the papers “Seeing the forest for all the trees: Mycorrhizal-associated nutrient economies are modulated by stem density and the synchrony between overstory and understory communities” and “Mycorrhizal associations of tree species influence soil nitrogen dynamics via effects on soil acid–base chemistry”. Four csv files are included along with a dataset. The dataset features chemical soil properties for a single sampling campaign within the 25 Ha Lilly-Dickey Woods Smithsonian Forest Global Earth Observatory (ForestGEO) plot in South Central Indiana, USA (ldw_dat_raw.csv). Also included are separate files focused on pH (pH_data.csv), carbon and nitrogen (CN_data.csv), and nitrification rates (Nitrification_data.csv). These variables are commonly associated with the tree-mycorrhizal dominance of forest stands. In these data subsets, each soil variable was matched to a 10 meter radius neighborhood wherein metrics of tree-mycorrhizal dominance (basal area, stem count, importance value, etc.) were calculated. Models between these soil variables and dominance metrics were used to investigate how different assessments of mycorrhizal associated nutrient economies (MANE) capture these relationships. This research was performed as a part of the Smithsonian ForestGEO project. This data package can be used to explore spatial variability in soil chemistry within a mature hardwood forest, or it can be combined with the included tree data, other fine-scale spatial information, or other tree inventory data for the site to evaluate how soil chemistry varies with tree community composition or edaphic or topographic properties.

Craig, Matthew [ORNL] (ORCID:0000000288907920)

Human Host Cellular Response to HCoV-229E Infection Proteomics (ACS-JM-DP2)

The purpose of this experiment was to evaluate the human host cellular response to wild-type Human coronavirus strain 229E (HCoV-229E) infection. Sample data was obtained for mock and infected immortalized human lung epithelial cells (A549) (MOI 5) nuclear extracts, immortalized human lung fibroblasts cells (MRC5) (MOI5) nuclear extracts, and primary human airway epithelial (HAE) (MOI 3) cells from lung tissue and processed for proteome analysis. Processed datasets are openly accessible from the download button and contain secondary processed proteomic results files and supporting metadata materials. Experimental proteomics samples were prepared using Limited Proteolysis (LiP) methods for Label-free quantification (LFQ) and global proteomic evaluation. Sample data was acquired using a Q-Exactive HF-X mass spectrometer and was processed and compiled using MaxQuant software (v.1.6.17.0). Processed proteomic data downloads include a sample naming key, processed MaxQuant results/parameters, and protein annotated relative abundance files. See corresponding primary data accessions below and Viral Experiment LiP Analysis source code supporting data transparency and reuse. Experimental transcriptomics samples were collected in parallel and processed for RNA sequencing (RNA-Seq) as summarized under ACS-DP1 (https://data.pnnl.gov/group/nodes/dataset/34069).

59 BASIC BIOLOGICAL SCIENCES

DEDUPKV: A Space-Efficient and High-Performance Key-Value Store via Fine-Grained Deduplication

Log-Structured Merge Tree (LSM-tree) based key-value stores excel in write-intensive environments but suffer from data duplication, consuming up to 49% of storage space in LSM-tree-based key-value store deployments. Traditional solutions like compression and coarse-grained file system-level deduplication introduce overhead or have limited effectiveness. In this study, we propose DedupKV, a fine-grained deduplication framework tailored for LSM-tree, maximizing data reduction efficiency while minimizing write stalls and read overheads. DedupKV features three key innovations: (1) FLUSH-integrated inline deduplication, which removes duplicates during memory-to-storage writes; (2) WAL file-based offline deduplication, repurposing write-ahead logs to avoid double writes; and (3) elastic execution, dynamically balancing inline and offline deduplication based on memory pressure and workload intensity. Additionally, dynamic granularity management reduces deduplication metadata overhead. We implemented these four ideas in RocksDB for the first time and conducted experiments in a Linux environment. Our evaluation shows that WAL file-based offline deduplication and DedupKV outperform BlobDB by 33% and 23%, respectively, in write-heavy workloads, while reducing write amplification by 1.2 ×, 2 ×, and 1.6 × for real KV datasets.

Jamil, Safdar [Sogang University]

Compilation of Experimental Yield Data for Spontaneous Fission of 252 Cf

We present a comprehensive compilation and curation of experimental fission yield (FY) data for the spontaneous fission of 252 Cf, extracted from the EXFOR database. The compilation follows a structured methodology developed for prior compilations of neutron-induced fission yields, and incorporates both independent (IFY) and cumulative (CFY) yields. A total of 62 datasets were reviewed, with entries spanning from 1955 to 2021. A significant portion of the literature reports pre-neutron emission yields, which were excluded from the present compilation due to limitations in format compatibility. Each accepted dataset was processed into a standardized JSON format, including metadata, uncertainties, and bibliographic references. Where available, decay radiation information was used to update the FY data using the latest ENSDF evaluations; 237 data points were corrected accordingly. These corrections are fully traceable and preserve original values. The result is a curated dataset suitable for use in nuclear data evaluations. This work is part of an ongoing effort to modernize the handling of FY data and provide evaluators with high-quality, machine-readable experimental inputs

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

FAIRness and Usability for Open-access Omics Data Systems

Omics data sharing is crucial to the biological research community, and the last decade or two has seen a huge rise in collaborative analysis systems, databases, and knowledge bases for omics and other systems biology data. We assessed the "FAIRness" of NASA's GeneLab Data Systems (GLDS) along with four similar kinds of systems in the research omics data domain, using 14 FAIRness metrics. The range of overall FAIRness scores was 6-12 (out of 14), average 10.1, and standard deviation 2.4. The range of Pass ratings for the metrics was 29-79%, Partial Pass 0-21%, and Fail 7-50%. The systems we evaluated performed the best in the areas of data findability and accessibility, and worst in the area of data interoperability. Reusability of metadata, in particular, was frequently not well supported. We relate our experiences implementing semantic integration of omics data from some of the assessed systems for federated querying and retrieval functions, given their shortcomings in data interoperability. Finally, we propose two new principles that Big Data system developers, in particular, should consider for maximizing data accessibility.

Berrios, Daniel C.

FAIRness and Usability for Open-access Omics Data Systems

Omics data sharing is crucial to the biological research community, and the last decade or two has seen a huge rise in collaborative analysis systems, databases, and knowledge bases for omics and other systems biology data. We assessed the “FAIRness” of NASA’s GeneLab Data Systems (GLDS) along with four similar kinds of systems in the research omics data domain, using 14 FAIRness metrics. The range of overall FAIRness scores was 6-12 (out of 14), average 10.1, and standard deviation 2.4. The range of Pass ratings for the metrics was 29-79%, Partial Pass 0-21%, and Fail 7-50%. The systems we evaluated performed the best in the areas of data findability and accessibility, and worst in the area of data interoperability. Reusability of metadata, in particular, was frequently not well supported. We relate our experiences implementing semantic integration of omics data from some of the assessed systems for federated querying and retrieval functions, given their shortcomings in data interoperability. Finally, we propose two new principles that Big Data system developers, in particular, should consider for maximizing data accessibility.

Berrios, Daniel C.

AIACHNE's contribution for Nuclear Energy Agency Working Party on International Nuclear Data Evaluation Co-operation Subgroup 50

The AIACHNE (AI/ML Informed cAlifornium CHi Nuclear data Experiment) project aims at designing an experiment for the 252 Cf Prompt Fission Neutron Spectrum (PFNS) that explores systematic biases in an experimental database retrieved from the EXFOR databases. To that end, machine learning (ML) methods were applied to pint-point measurement features likely related to bia. From that information, we selected a feature that should be explored by the AIACHNE experiment. Measurement features are metadata encapsulating all pertinent information about the physical measurement and analysis techniques. Examples are, for instance, what neutron and fission detectors were used for the physical metadata, and what background reduction techniques were employed for analysis techniques. Such metadata were retrieved both from EXFOR entries as well as the literature of data sets described in detail in Ref. [2]. The prerequisite for applying machine learning techniques is casting the metadata into a format that can be parsed by the algorithm. This step might seem trivial but requires to find a unique language where metadata that carry the same physics meaning across several experiments must have the same identifier. One example is, for instance, the neutron detector. As seen in Figure 1, the machine learning code identified the use of 6 Li detectors as being related to bias in some datasets of the AIACHNE 252 Cf PFNS experimental database. In fact, here are several experiments that used neutron detectors containing 6Li in the database, for instance for the example below. EXFOR format has a unique keywords describing detectors such as “SCIN” or “GLASD”. One may think that these keywords are already sufficient descriptors for ML to uniquely find an issue. However, “SCIN” (used for [3, 4]) and “GLASD” (used for [5]) fail to inform the algorithm what is the active material in the detector. And, the key common issue leading to bias in 252 Cf related to neutron detectors is not whether it is a glass detector or a scintillator. No, the issue is that 6 Li was within both detector types and that even small mistakes in the detector response functions around approximately 200 keV are amplified by the 6 Li(n,α) resonance there leading to bias in data as highlighted in Fig. 1 and Ref. [1]. Hence, the features describing the neutron detector must call out the active material in the detector, rather than the existing EXFOR detector keyword, that the ML algorithm can find physically meaningful features related to bias. The AIACHNE team used a precursor of the WPEC (Working Party on International Nuclear Data Evaluation Co-operation) SG(Subgroup)-50 format to store the metadata for the ML analysis.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Livewire: A Model Platform for Data Quality Assessment and AI Readiness Across DOE Missions

High-quality, well-governed data is essential for accelerating discovery and achieving operational excellence across DOE and national laboratory missions. The Livewire Data Platform is a DOE-supported platform that offers automated assessments of data quality, standardization, provenance, and Artificial Intelligence (AI) readiness. It allows researchers and data practitioners to systematically and easily evaluate datasets against established governance criteria and prepare them for advanced analytics. Livewire addresses critical challenges in DOE's data ecosystem with integrated capabilities for metadata validation, provenance tracking, and schema alignment. This platform's automated workflows assist users in identifying data quality gaps, enhancing interoperability between datasets collected from various stakeholders, and ensuring compliance with DOE data standards, all while reducing manual curation efforts. Additionally, we will discuss its AI readiness framework, which is being developed to prepare datasets for training models, developing advanced analytic tools, and machine learning applications. Using some of the more than one hundred tabular datasets on Livewire, processed with this open-source methodology, we will demonstrate how Livewire can serve as a model for scalable, standards-driven data management. This approach provides a pathway to leverage existing and future datasets within the DOE, boosting innovation and efficiency across national laboratories.

33 - ADVANCED PROPULSION SYSTEMS

Evolving a NASA Digital Object Identifiers System with Community Engagement

To demonstrate how the ESDIS (Earth Science Data and Information System) DOI (Digital Object Identifier) system and its processes have evolved over these years based on the recommendations provided by the user community (whether the community members create and manage DOI information or use DOIs in the data citations). The user community is comprised of people with common interests and needs for data identifiers who are actively involved in the creation and usage process. Engagement describes the interactive context wherein the community provides information, evaluates the proposed processes, and provides guidance in the area of identifiers.

Identifiers

TOLNet’s FAIR Journey: Yesterday, Today, and Tomorrow

The Tropospheric Ozone Lidar Network (TOLNet) has generated over a decade of ozone vertical profile data products over North America and contributed to several air quality focused field studies. The science value of the TOLNet data has been demonstrated in numerous peer-reviewed publications on air quality and ozone relevant research. As the broad scientific community has moved towards adopting FAIR Principles to make data more findable, accessible, interoperable, and (re)usable, the TOLNet team has been consistently making data more FAIR. This effort has many challenges, partially reflecting on the FAIR principles being domain agnostic while the implementation needs to be domain specific. The FAIR principles declare the dependence on the community standards, domain-relevant metadata, and rich metadata. This presentation uses the TOLNet data and data system as an example to explore the best practices to implement FAIR principle. Particularly, we will examine the metadata and the “richness” to support findability and usability as well as machine-to-machine actionability via API. Last year, as part of our FAIR journey, we launched the TOLNet website (https://tolnet.larc.nasa.gov/) and the API (https://tolnet.larc.nasa.gov/api/). Part of this process included extracting and cataloging metadata across the entire TOLNet mission timeframe. This enabled users to search through the mission by various metadata criteria, improving the findability and accessibility. And computers could connect directly to the TOLNet API to extract both metadata and data, providing a level of interoperability never present before for TOLNet data. On top of that, all new TOLNet data is now automatically validated using the API to ensure it complies with GEOMS standards, aiding in reusability. It takes both technology and scientists working together to make progress. The next step is to evaluate the current TOLNet offerings against NASA’s Practical Guide for Open, Free & FAIR NASA Earth Science Data Products (https://doi.org/10.5067/DOC/ESCO/ESDSWG-0002V1).

TOLNet

PPI DataHub Project Data Package: S. elongatus PCC 7942 Limited Proteolysis and Thermal Proteome Profiling Structural Proteomics (JM-PB-DP3)

The purpose of this experiment was to investigate structural alterations in proteins involved in central carbon metabolism and photosynthetic electron transfer pathways in Synechococcus elongatus PCC 7942. Sample data was obtained from S. elongatus cell lysates using three complementary mass spectrometry (MS) techniques using limited proteolysis (LiP-MS), thermal proteome profiling (TPP-MS), and redox enrichment (Redox-MS) in evaluating alterations solvent accessibility and structural stability caused by light perturbation at the molecular level. Experimentally processed sample data for LiP and TPP proteomic datasets were derived from the same cell culture stock, prepared simultaneously in parallel, and acquired by mass spectrometry. Processed datasets are openly accessible from the download button and contain secondary processed proteomic results files, computed outputs, and supporting metadata materials. Experimental samples processed for LiP-MS label-free quantification (LFQ) or TPP-MS tandem mass tag (TMT) 10-plex were acquired using a Q-Exactive HF-X mass spectrometer and processed/compiled using either MSGF+ (v2024.03.26) or ​​​​PlexedPiper for proteome evaluation. Additional software supporting downstream proteomic analysis include FragPipe (v.4.0), MSFragger (v.22.1), and an adapted Microbial Isolate LiP Analysis Workflow (located at Zenodo). Processed proteomic data downloads include a sample naming key, normalized quantification results files, and processed protein annotated abundance files.

59 BASIC BIOLOGICAL SCIENCES

User Evaluation of the NASA Technical Report Server Recommendation Service

We present the user evaluation of two recommendation server methodologies implemented for the NASA Technical Report Server (NTRS). One methodology for generating recommendations uses log analysis to identify co-retrieval events on full-text documents. For comparison, we used the Vector Space Model (VSM) as the second methodology. We calculated cosine similarities and used the top 10 most similar documents (based on metadata) as 'recommendations'. We then ran an experiment with NASA Langley Research Center (LaRC) staff members to gather their feedback on which method produced the most 'quality' recommendations. We found that in most cases VSM outperformed log analysis of co-retrievals. However, analyzing the data revealed the evaluations may have been structurally biased in favor of the VSM generated recommendations. We explore some possible methods for combining log analysis and VSM generated recommendations and suggest areas of future work.

Nelson, Michael L.

User Evaluation of the NASA Technical Report Server Recommendation Service

We present the user evaluation of two recommendation server methodologies implemented for the NASA Technical Report Server (NTRS). One methodology for generating recommendations uses log analysis to identify co-retrieval events on full-text documents. For comparison, we used the Vector Space Model (VSM) as the second methodology. We calculated cosine similarities and used the top 10 most similar documents (based on metadata) as recommendations . We then ran an experiment with NASA Langley Research Center (LaRC) staff members to gather their feedback on which method produced the most quality recommendations. We found that in most cases VSM outperformed log analysis of co-retrievals. However, analyzing the data revealed the evaluations may have been structurally biased in favor of the VSM generated recommendations. We explore some possible methods for combining log analysis and VSM generated recommendations and suggest areas of future work.

Nelson, Michael L.

Evaluation of Best Practices in Mitigating Startup Costs on Leadership-Class Supercomputers

Supercomputers at Department of Energy (DOE) National Laboratories face a widening range of workloads, from traditional modeling and simulation to Artificial Intelligence model training or complex multi-stage workflows, and beyond. At DOE Leadership Computing Facilities like the Oak Ridge Leadership Computing Facility (OLCF), these workloads demand concurrent access to large portions of the supercomputer’s resources. Launching a job across massive supercomputers is challenging from the start; the file system struggles with a large backlog of metadata requests as tens of thousands of processes read thousands of the same files, and the compute job cannot start until this is completed. There are multiple existing approaches to calm this metadata storm, ranging from vendor-developed tools like sbcast to National Laboratory-developed tools like Spindle and Copper. In this paper, we benchmark and discuss three common approaches to improving compute job launch latencies on Frontier: Slurm’s sbcast tool, Spindle, and Copper. We evaluate these tools by measuring the launch latencies of four workloads: OSU Microbenchmark’s osu_init, Pynamic, Python import mpi4py, and Python import torch. We provide discussion of the results, highlighting data that meet expectations and that do not meet expectations.

Hagerty, Nick [ORNL] (ORCID:0000000330014414)