Search NASA⌕ Search

SEARCH · Search NASA

Results for “system metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Constraints on Future Analysis Metadata Systems in High Energy Physics

In high energy physics (HEP), analysis metadata comes in many forms—from theoretical cross-sections, to calibration corrections, to details about file processing. Correctly applying metadata is a crucial and often time-consuming step in an analysis, but designing analysis metadata systems has historically received little direct attention. Among other considerations, an ideal metadata tool should be easy to use by new analysers, should scale to large data volumes and diverse processing paradigms, and should enable future analysis reinterpretation. This document, which is the product of community discussions organised by the HEP Software Foundation, categorises types of metadata by scope and format and gives examples of current metadata solutions. Important design considerations for metadata systems, including sociological factors, analysis preservation efforts, and technical factors, are discussed. A list of best practices and technical requirements for future analysis metadata systems is presented. These best practices could guide the development of a future cross-experimental effort for analysis metadata tools.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Best Practices for Energy Management Information Systems Metadata Schemas

Fact sheet highlights best practices for metadata schemas and standard naming conventions, which improve the ability of an energy management information system (EMIS) to consistently analyze, visualize, and derive value from operational data. This best practice document is part of a series of fact sheets created to help accelerate the market adoption and use of EMIS in the federal sector. It provides an overview of best practices for using metadata in an EMIS stack, which can significantly reduce the initial time required to deploy an EMIS.

data tags↗

Data Catalog Project - A Browsable, Searchable, Metadata System

Modern experiments are typically conducted by large, extended, where researchers rely on other team members to produce much of the data they use. The experiments record very large numbers of measurements which can be difficult for users to find, access and understand. We are developing a system for users to annotate their data products with structured metadata, providing data consumers with a discoverable, browsable data index. Machine understandable metadata captures the underlying semantics of the recorded data, which can then be consumed by both programs, and interactively by users. Collaborators can use these metadata to select and understand recorded measurements.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Best Practices for Energy Management Information Systems Metadata Schemas

The Federal Energy Management Program (FEMP) promotes best practices for impactful utilization of Energy Management Information Systems (EMIS) at federal facilities. This best practice document is part of a series of fact sheets created to help accelerate the market adoption and use of EMIS in the federal sector. The use of standard naming conventions and a metadata schema, which may be referred to as 'data tags,' improves the ability of the EMIS to consistently analyze, visualize, and derive value from operational data. To help address these issues, this fact sheet highlights best practices for use of metadata in an EMIS stack, which can significantly reduce the initial time required to deploy an EMIS.

EMIS↗

MetaCat - metadata catalog for data management systems

Metadata management is one of three major areas and parts of functionality of scientific data management along with replica management and workflow management. Metadata is the information describing the data stored in a data item, a file or an object. It includes the data item provenance, recording conditions, format and other attributes. MetaCat is a metadata management database designed and developed for High Energy Physics experiments. As a component of a data management system, it’s main objectives are to provide efficient metadata storage and management and fast data items selection functionality. MetaCat is supposed to work on the scale of 100 million files (or objects) and beyond. The article will discuss the functionality of MetaCat and technological solutions used to implement the product.

Mandrichenko, Igor↗

Scalable filesystem enumeration and metadata operations

Systems, apparatus, and methods are disclosed for performing scalable operations in a file system, including POSIX-like file systems. Metadata entries in a namespace or directory tree are sharded across multiple file metadata servers. An enumeration operation, such as listing a directory, is parallelized across the multiple file metadata servers, while retaining standard functionality transparently to clients. Other enumeration operations include no-output operations such as changing file attributes or deleting a file, and cumulative operations such as counting total disk space usage. The parallelization is compatible with tree-level parallelization and storage-level parallelization. Disclosed technologies can be applied to other fields requiring scalable enumeration, such as database and network applications.

Grider, Gary A.↗

Scalable augmented enumeration and metadata operations for large filesystems

Systems, apparatus, and methods are disclosed for performing scalable operations in a file system. Metadata entries in a namespace or directory tree are sharded across multiple file metadata servers. An augmented enumeration operation, such as listing a directory, is parallelized across the multiple file metadata servers, transparently to clients. Exemplary augmentation features can include filtering and sorting. Augmentation features can be executed concurrently with enumeration, prior to enumeration, after enumeration, or as a combination of these, and can utilize pre-built index structures or holding structures for intermediate results. Augmented enumeration operations can also include no-output operations such as changing file attributes or deleting a file, and cumulative operations such as counting total disk space usage. The parallelization is compatible with tree-level parallelization and storage-level parallelization. Disclosed technologies can be applied to other fields requiring scalable enumeration, such as database and network applications.

Grider, Gary A.↗

"PoliMOR: A Policy Engine \"Made-to-Order\" for Automated and Scalable Data Management in Lustre"

Modern supercomputing systems are increasingly reliant on hierarchical, multi-tiered file and storage system architectures due to cost-performance-capacity trade-offs. Within such multi-tiered systems, data management services are required to maintain healthy utilization, performance, and capacity levels. We present PoliMOR, a pragmatic and reliable policy-driven data management framework. PoliMOR is composed of modular, single-purpose agents that gather file system metadata and enforce policies on storage systems. PoliMOR facilitates automated and scalable data management with customizable agents tailored to HPC facility-specific storage systems and policies. Our evaluations demonstrate the scalability and performance of PoliMOR both by its individual agents and as a collective entity. We believe PoliMOR is widely applicable across HPC facilities with large-scale data management challenges and will garner interest from the HPC community, given its flexible and open-source nature.

George, Anjus↗

Photovoltaic Data Acquisition (PVDAQ) Public Datasets

The NREL PVDAQ is a large-scale time-series database containing system metadata and performance data from a variety of experimental PV sites and commercial public PV sites. The datasets are used to perform on-going performance and degradation analysis. Some of the sets can exhibit common elements that effect PV performance (e.g. soiling). The dataset consists of a series of files devoted to each of the systems and an associated set of metadata information that explains details about the system hardware and the site geo-location. Some system datasets also include environmental sensors that cover irradiance, temperatures, wind speeds, and precipitation at the site.

Array↗

Automated Metadata Extraction: Challenges and Opportunities

Proper application of the FAIR data principles is what separates a vibrant data ecosystem, in which research data are frequently shared and reused, from a lifeless data graveyard. Automated metadata extraction systems have been proposed as a means of bolstering the findability, interoperability, and reusabil- ity of data repositories with little or no human intervention. These extraction systems mine metadata by crawling a repository and applying lightweight extractors that, for various types of file (e.g., image, CSV file), extract or synthesize relevant attributes. In practice, however, the automated creation of generally useful metadata is fraught with challenges. Data consumers may have different perspectives as to what metadata representations are useful, the standards for recording metadata tend to change over time, and the software model for processing updates can introduce unnecessary human and computational effort. Thus, generalizing extraction for a broad audience of data consumers is a difficult and relatively unsolved problem.In this work, we explore these challenges faced by extraction systems in the context of constructing our own extraction system for science data. We first define the metadata extraction problem and provide context to the issues faced in generalizing metadata. Additionally, we identify potential research directions to help alleviate many of these challenges for all automated extraction systems. Ultimately, this work represents a first step in designing ubiquitous metadata extraction systems that can maximize the value of research data while minimizing the human efforts required in doing so.

Skluzacek, Tyler↗

The fast camera (Fastcam) imaging diagnostic systems on the DIII-D tokamak

Two camera systems are installed on the DIII-D tokamak at the toroidal positions of 90° (90° system) and 225° (225° system), respectively. The cameras have two types of relay optics, namely, a coherent optical fiber bundle and a periscope system. The periscope system provides absolute intensity calibration stability while sacrificing resolution (10 lp/mm), while the fiber system provides high resolution (16 lp/mm) while sacrificing calibration stability. The periscope is available only for the 90° system. The optics of the 225° system were designed for view stability, repeatability, and easy maintenance. The cameras are located inside optimized neutron, x ray and magnetic shielding in order to reduce electronics damage, reboots, and magnetic and neutron interference, increasing the overall system reliability. An automated filter wheel, providing remote filter change, allows for remote wavelength selection. A software suite automates camera acquisition and data storage, allowing for remote operation and reduced operator involvement. System metadata is used to streamline the data analysis workflow, particularly for intensity calibration. Here, the spatial calibration uses multiple observable wall features, resulting in a reconstruction accuracy ≤2 cm.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

PVInsight (Final Technical Report)

Data generated from real-world photovoltaic (PV) systems represent significant opportunities for the industry–from digital operations and maintenance to real-time planning and forecasting. However, these data also come with substantial, unique challenges. A particular challenge is the analysis of unlabeled PV performance data, which we define as time-series measurements of real power production (or sometimes current or voltage) that do not have corresponding meteorological measurements (irradiance, temperature, etc.) or system configuration information (sometimes called system metadata). These challenges are amplified in the distributed rooftop sector, in which data quality, completeness, and metadata can be very poor.

14 SOLAR ENERGY↗

Improved CdTe PLR Estimates: Self-Shading and Spectral Mismatch

The RdTools year-on-year method of estimating performance loss rate (PLR) employs a simple normalization to remove the confounding effect of irradiance and temperature variation. However, the normalization's assumption that PV production scales linearly with in-plane broadband irradiance is a worse approximation for CdTe and other technologies with larger spectral sensitivities than it is for the more common c-Si technology. Additionally, CdTe systems using single-axis trackers (-20% of installed US utility-scale capacity) are subject to self-shading in the morning and afternoon, introducing another nonlinearity between PV output and broadband irradiance. Ignoring these effects may subject the estimated PLR to increased uncertainty, and perhaps bias, depending on the character of their short- and long-term variability. In this work we show that including self-shading and spectral mismatch models in the normalization for tracking CdTe systems can result not only in tighter PLR confidence intervals but different median PLRs as well. The shading and spectral models are kept simple to maintain consistency with the RdTools ethos of not requiring detailed system metadata or unusual measurements.

CdTe↗

Quantifying Error in Photovoltaic Installation Metadata: Preprint

In this research, we quantify the level of metadata error for a fleet of 2860 photovoltaic (PV) systems, using metadata values provided by fleet owners. Using satellite imagery and time series analysis techniques available in open-source Python packages Panel-Segmentation and PVAnalytics, respectively, we evaluate the accuracy of PV system metadata such as location, azimuth, tilt, and mounting configuration (fixed tilt vs. tracking). We find that approximately 75% of provided latitude-longitude coordinates are within 190 meters of the actual solar installation. We were unable to link 7.8% of latitude-longitude coordinates to any solar installation via satellite imagery analysis. We evaluate the level of error in owner-provided mounting configuration (fixed tilt vs. single-axis tracking), finding only 8 systems with an incorrect mounting configuration. When evaluating azimuth and tilt parameters, we find that approximately 64% of the data is correct, with data for 860 systems (approximately 30%) not provided by system owners. To illustrate the importance of having correct solar metadata, we evaluate how incorrect metadata affects solar performance estimates by modeling system AC energy output at ground-truth vs. incorrect latitude-longitude coordinates, mounting configurations, and azimuth-tilt configurations. Energy output estimates can vary significantly if incorrect metadata parameters are used, with incorrect mounting configuration leading to the largest discrepancy with over 20% variation in expected energy output.

azimuth↗

Unified architecture for data-driven metadata tagging of building automation systems

This article presents a Unified Architecture (UA) for automated point tagging of Building Automation System (BAS) data, based on a combination of data-driven approaches. Advanced energy analytics applications—including fault detection and diagnostics and supervisory control—have emerged as a significant opportunity for improving the performance of our built environment. Effective application of these analytics depends on harnessing structured data from the various building control and monitoring systems, but typical BAS implementations do not employ any standardized metadata schema. While standards such as Project Haystack and Brick Schema have been developed to address this issue, the process of structuring the data, i.e., tagging the points to apply a standard metadata schema, has, to date, been a manual process. This process is typically costly, labor-intensive, and error-prone. In this work we address this gap by proposing a UA that automates the process of point tagging by leveraging the data accessible through connection to the BAS, including time-series data and the raw point names. The UA intertwines supervised classification and unsupervised clustering techniques from machine learning and leverages both their deterministic and probabilistic outputs to inform the point tagging process. Furthermore, we extend the UA to embed additional input and output data-processing modules that are designed to address the challenges associated with the real-time deployment of this automation solution. We test the UA on two datasets for real-life buildings: (i) commercial retail buildings and (ii) office buildings from the National Renewable Energy Laboratory (NREL) campus. We report the proposed methodology correctly applied 85–90% and 70–75% of the tags in each of these test scenarios, respectively for two significantly different building types used for testing UA's fully-functional prototype. The proposed UA, therefore, offers promising approach for automatically tagging BAS data as it reaches close to 90% accuracy. Further building upon this framework to algorithmically identify the equipment type and their relationships is an apt future research direction to pursue.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Shaping the Future of Self-Driving Autonomous Laboratories Workshop

The "Shaping the Future of Self-Driving Autonomous Laboratories" workshop, held in Denver on November 7-8, 2024, brought together leading experts from materials science and computing to address the growing need to revolutionize scientific research through AI-driven autonomous laboratories. The workshop identified critical challenges, including the integration of heterogeneous data, development of AI systems that understand fundamental physical principles, and comprehensive safety protocols. Key recommendations emerged around developing universal laboratory equipment interfaces, implementing automated metadata collection systems, and creating hybrid AI approaches that combine data-driven learning with scientific principles. The workshop emphasized maintaining human oversight while leveraging automation, transforming scientific education to prepare the next generation of researchers, and establishing a national consortium leveraging DOE facilities as anchors for broader collaboration with academia and industry. Participants stressed the urgency of addressing the growing disconnect between human decision-making timescales and modern instrumentation capabilities, highlighting the need for strategic automation while preserving essential human insight and oversight in the research process.

36 MATERIALS SCIENCE↗

Centralized Interactive Phenomics Resource: an integrated online phenomics knowledgebase for health data users

Development of clinical phenotypes from electronic health records (EHRs) can be resource intensive. Several phenotype libraries have been created to facilitate reuse of definitions. However, these platforms vary in target audience and utility. Here, we describe the development of the Centralized Interactive Phenomics Resource (CIPHER) knowledgebase, a comprehensive public-facing phenotype library, which aims to facilitate clinical and health services research. The platform was designed to collect and catalog EHR-based computable phenotype algorithms from any healthcare system, scale metadata management, facilitate phenotype discovery, and allow for integration of tools and user workflows. Phenomics experts were engaged in the development and testing of the site. The knowledgebase stores phenotype metadata using the CIPHER standard, and definitions are accessible through complex searching. Phenotypes are contributed to the knowledgebase via webform, allowing metadata validation. Data visualization tools linking to the knowledgebase enhance user interaction with content and accelerate phenotype development. The CIPHER knowledgebase was developed in the largest healthcare system in the United States and piloted with external partners. The design of the CIPHER website supports a variety of front-end tools and features to facilitate phenotype development and reuse. Health data users are encouraged to contribute their algorithms to the knowledgebase for wider dissemination to the research community, and to use the platform as a springboard for phenotyping. CIPHER is a public resource for all health data users available at https://phenomics.va.ornl.gov/ which facilitates phenotype reuse, development, and dissemination of phenotyping knowledge.

60 APPLIED LIFE SCIENCES↗