Search NASA⌕ Search

SEARCH · Search NASA

Results for “Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

bibcheck

SAND2026-16981O Bibcheck is designed to extract bibliographies from research papers and perform metadata searches to identify errors. It assists authors in checking their bibliographies for metadata errors during the writing process and helps reviewers identify errors in bibliographies of papers under review. The software uses large language models (LLMs) to extract bibliography entries from PDF documents, classifies the type of bibliography entry, and verifies referenced works. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Pearson, Carl [Sandia National Lab. (SNL-CA), Live↗

GRinding Automated Classification Engine

This work is an ML-driven framework for automated surface analysis of microscopy images. We create a training dataset by imaging stainless steel samples to benchmark four developed deep neural network architectures. These models, based on a YOLOv8n-cls backend, integrate image features and process metadata using various fusion methods to distinguish between acceptable and unacceptable surface finishes. This code is associated with publication "Classifying Alloy Surface Preparation Quality with Metadata-Infused Machine Learning for Rapid Alloy Discovery" for project APEX LDRD-ER (25-ERD-039)

Gongora, AldairE [Lawrence Livermore National Labo↗

A Data Processing Pipeline To Extract A Knowledge Graph From Heterogeneous Data For Socio-technical Analysis Of Critical Infrastructure Influence

The code is written in Python and consists of the following pipeline that is implemented in Apache Airflow. This pipeline intends to understand the companies that are directly or indirectly involved with a type of critical infrastructure system at some point in that system's lifecycle. The pipeline takes a configuration file that specifies a list of initial companies to consider, a geographic region of interest, and a set of SEC form types as well as other data sources (e.g. CrunchBase) from which to extract entities and relations. There are four main components to this pipeline as currently implemented: Entity Extraction, Network Construction, Analysis, and Visualization. First, Entity Extraction, is implemented as the `topear-extract_organizations` Apache Airflow workflow. Given an initial query that specifies a geographic region of interest and a time interval, the software will extract CI facilities of interest and organizations that have a direct influence relationship to those facilities (e.g. ownership). During the course of the LDRD, we focused on Electric Vehicle charging stations and this information is available via the Department of Energy (DOE) database on fueling stations maintained by NREL. Within the context of the DOE CESER project, we have focused on Battery Energy Storage Systems (BESS). Second, the Network Extraction component will iteratively construct a social network graph given the set of organizations and people extracted in the previous step. Organizations (and eventually People if desired) are then fed as a query to the `topgear-construct_social_network` Apache Airflow workflow which given a set of initial companies and data sets (e.g. SEC EDGAR form types, OpenCorporates, Crunchbase). This Airflow workflow will iteratively query such data sources to discover relationships with new organizations and people. For example, this module can iteratively query SEC EDGAR for metadata that documents the number of each type of form for the given set of companies and their location. This forms metadata represents a catalog of data sources from SEC EDGAR for the extracted social network knowledge graph. The pipeline then downloads these forms from the website and saves them in a build directory for further processing. These documents are then parsed for entities and relations. Again, we note that in additional to SEC data sources, this step can also pull in information on organizations via API services such as CrunchBase and OpenCorporates or bulk data sources. At the end of this step, the resultant social network, the Critical Infrastructure network, and the edges that encode relationships between organizations and CI facilities, form the Adversarial Socio-Technical Network (ASTN) that informs the analysis. Third, the Analysis component processes these generated ASTN. Previously, that has included the ability to compare prevalence of different vendors for a given infrastructure component type across different regions as well as identify common public and private investors across those vendors. This was demonstrated for EV Charging Stations across several different metropolitan areas within an IEEE PES GridEdge publication. More recently, we have looked at ways to identify infrastructure owners and operators of BESS with the most nameplate capacity across different states as well as other indictors of risk resulting from changes in ownership over time. Finally, the Visualization component consists of an HTML/CSS/JS framework by which users can interact geospatial, operational, and organizational relationships across a given portfolio of Critical Infrastructure facilities. The objective is to provide a library of UI/UX modules that can be repurposed for stakeholder-specific dashboards. All of the modules are related via a common event model that enables UI actions in one view to percolate across the other views.

Weaver, Gabriel [Idaho National Laboratory (INL), ↗

The need for standardization and improved open (meta)data practices in metaproteomics

Metaproteomics enables functional insight into microbial communities by identifying and quantifying proteins in complex samples. Yet, heterogeneous analytical workflows and the lack of standardization across experimental and bioinformatics stages hinder reproducibility and comparability, limiting integration with other omics data. We here present a community-developed reporting checklist tailored to the specific needs of metaproteomics. We also outline current efforts to enable structured and interoperable metadata capture, drawing on standards from proteomics and microbiome research wherever possible. By promoting transparent reporting and advancing metadata practices, our recommendations aim to align metaproteomics more closely with FAIR principles and support reproducible and interoperable research practices.

Armengaud, Jean [Universite Paris-Saclay, France]↗

EAGLE-I Power Outage Data 2024

The provided EAGLE-I historic dataset includes power outage information at the county level for 2024 at 15-minute intervals collected by the EAGLE-I program at ORNL. The data has been collected from utility's public outage maps using an ETL process. The dataset details FIPS code, county name, state name, total number of customers without power, total customers per county, and a date/timestamp. For detailed metadata, refer to the linked metadata DOI.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Frontier Job-Centric Telemetry Dataset

Comprehensive analysis of high-performance computing (HPC) systems requires linking workload execution to system behavior. This kind of analysis is vital for diagnosing performance issues, managing capacity, detecting anomalous workloads, and understanding how applications interact with system hardware. This job-centric telemetry dataset unifies scheduler job records with node-level measurements, enabling direct association between workloads and their corresponding power, thermal, and performance characteristics. It contains sanitized, scheduler related metadata for 152,400 individual jobs that ran on the Frontier supercomputer and ended on selected days throughout 2024 and 2025, a subpopulation of ~6.8% of the total number of allocated jobs with non-zero run time on the system over that same period. Each is linked with files that contain telemetry time series records of the power utilization and temperature behavior of its allocated nodes and their processors during the run time of the job. Where available, a portion of the job files also contain network performance time series. Jobs are sampled from select days that reflect normal levels of user activity and possess job size distributions with large numbers of leadership class jobs (>20% of Frontier nodes). Jobs in this dataset attempt to best represent successful user workflows.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

EAGLE-I Power Outage Data 2025

The provided EAGLE-I historic dataset includes power outage information at the county level for 2025 at 15-minute intervals collected by the EAGLE-I program at ORNL. The data has been collected from utility's public outage maps using an ETL process. The dataset details FIPS code, county name, state name, total number of customers without power, and a date/timestamp. For detailed metadata, refer to the linked metadata DOI.

EAGLE-I↗

Spatialyze: A Geospatial Video Analytics System with Spatial-Aware Optimizations

Videos that are shot using commodity hardware such as phones and surveillance cameras record various metadata such as time and location. We encounter suchgeospatial videoson a daily basis and such videos have been growing in volume significantly. Yet, we do not have data management systems that allow users to interact with such data effectively. In this paper, we describe Spatialyze, a new framework for end-to-end querying of geospatial videos. Spatialyze comes with a domain-specific language where users can construct geospatial video analytic workflows using a 3-step, declarative,build-filter-observeparadigm. Internally, Spatialyze leverages the declarative nature of such workflows, the temporal-spatial metadata stored with videos, and physical behavior of real-world objects to optimize the execution of workflows. Our results using real-world videos and workflows show that Spatialyze can reduce execution time by up to 5.3×, while maintaining up to 97.1% accuracy compared to unoptimized execution.

Computer Science↗

DAISY Benchmark Performance Data

This repository contains the underlying data from benchmark experiments for Drifting Acoustic Instrumentation SYstems (DAISYs) in waves and currents described in "Performance of a Drifting Acoustic Instrumentation SYstem (DAISY) for Characterizing Radiated Noise from Marine Energy Converters" (https://link.springer.com/article/10.1007/s40722-024-00358-6). DAISYs consist of a surface expression connected to a hydrophone recording package by a tether. Both elements are instrumented to provide metadata (e.g., position, orientation, and depth). Information about how to build DAISYs is available at https://www.pmec.us/research-projects/daisy. The repository's primary content is three compressed archives (.zip format), each containing multiple MATLAB binary data files (.mat format). A table relating individual data files to figures in the paper, as well as the structure of each file, is included in the repository as a Word document (Data Description MHK-DR.docx). Most of the files contain time series information for a single DAISY deployment (file naming convention: [site]_DAISY_[Drift #].mat) consisting of processed hydrophone data and associated metadata. For a limited number of DAISY deployments, the hydrophone package was replaced with an acoustic Doppler velocimeter (file naming convention: [site]_DAISY_[Drift #]_ADV.mat). Data were collected over several years at three locations: (1) Sequim Bay at Pacific Northwest National Laboratory's Marine & Coastal Research Laboratory (MCRL) in Sequim, WA, the energetic tidal channel in Admiralty Inlet, WA (Admiralty Inlet), and the U.S. Navy's Wave Energy Test Site (WETS) in Kaneohe, HI. Brief descriptions of data files at each location follow. - MCRL - (1) Drift #4 and #16 contrast the performance of a DAISY and a reference hydrophone (icListen HF Reson), respectively, in the quiescent interior of Sequim Bay (September 2020). (2) Drift #152 and #153 are velocity measurements for a drifting acoustic Doppler velocimeter in in the tidally-energetic entrance channel inside a flow shield and exposed to the flow, respectively (January 2018). (3) Two non-standard files are also included: DAISY_data.mat corresponds to a subset of a DAISY drift over an Adaptable Monitoring Package (AMP) and AMP_data.mat corresponds to approximately co-temporal data for a stationary hydrophone on the AMP (February 2019). - Admiralty Inlet - (1) Drift #1-12 correspond to tests with flow shielded DAISYs, unshielded DAISYs, a reference hydrophone, and drifting acoustic Doppler velocimeter with 5, 10, and 15 m tether lengths between surface expression and hydrophone recording package (July 2022). (2) Drift #13-20 correspond to tests of flow shielded DAISYs with three different tether materials (rubber cord, nylon line, and faired nylon line) in lengths of 5, 10, and 15 m (July 2022). - WETS - (1) Drift #30-32 correspond to tests with a heave plate incorporated into the tether (standard configuration for wave sites), rubber cord only, and rubber cord, but with a flow shielded hydrophone (November 2022). (2) Drift #49-58 and Drift #65-68 correspond to measurements around mooring infrastructure at the 60 m berth where time-delay-of-arrival localization was demonstrated for different DAISY arrangements and hydrophone depths (November 2022).

16 TIDAL AND WAVE POWER↗

Reproductive and leaf litterfall fluxes in forest ecosystem sites globally (1950-2022)

Forest allocation of net primary productivity (NPP) to reproduction is poorly quantified globally, despite its critical role in forest regeneration and a well-supported trade-off with allocation to growth. Although field measurements of total NPP are rare, our work finds that a proxy for reproductive carbon allocation constructed from leaf (L) and reproductive (R) litterfall fluxes, R/(R+L), is strongly correlated with R/NPP, facilitating analysis across a wide range of sites where biometric estimates of NPP are not available (R² = 0.85; Hanbury-Brown et al., 2022, Ward et al., in prep). To investigate relationships between ecosystem-scale reproductive allocation (RA) and climate, soil fertility, and stand age gradients, we conducted a literature search and synthesized 824 observations of annual average leaf and reproductive litterfall fluxes across forest sites globally. The zip file includes 1) a folder Data/ containing the litterfall data ("GlobalForestRA_data.csv") and metadata ("GlobalForestRA_metadata.doc") files. The data file includes geographic coordinates, long-term mean annual temperature and precipitation (1970-2000, extracted from WorldClim2.1), leaf and reproductive litterfall fluxes, sampling interval and protocols, forest characteristics (dominant leaf morphology, information pertaining to forest age and successional stage, and disturbance history) and soil properties (% sand, %silt, %clay, total phosphorus (P), nitrogen (N), cation exchange capacity (CEC) and pH) extracted from SoilGrids250 and from on-site measurements, where available. The metadata file contains information about each variable reported in the data file, including data sources, processing methods, and all references. The Data folder contains two additional files used to create Figure 1; these are described in greater detail in the README.2) R scripts GloalForestRA_analysis.r and GlobalForestRA_SI.r and a folder /Functions used to produce results, figures, and tables in the manuscript Ward et al. (in press)3) a README file describing how the data and R scripts can be used to reproduce statistical results, figures, and tables found in the manuscript. Ward et al. (in press)This repository can also be found at: https://github.com/r-ward/Global_Analysis_ForestRA.Ward, R.E., Zhang-Zheng, H. Aernethy, K., Adu-Bredu, S., Arroyo, L., Bailey, A. et al. (in press). Forest age rivals climate to explain reproductive allocation patterns in forest ecosystems globally. Ecology Letters. Hanbury-Brown, A.R., Ward, R.E. & Kueppers, L.M. (2022). Forest regeneration within Earth system models: current process representations and ways forward. New Phytol., 235, 20–40.Ward et al. (2025), Forest age rivals climate to explain reproductive allocation patterns in forest ecosystems globally, in prep.

54 ENVIRONMENTAL SCIENCES↗

Changuinola peat soil characteristics and gas emission raw data October 2019

This dataset comprises radiocarbon and geochemical measurements from peat and porewater samples collected across various depths at a site in Bocas del Toro, Panama. The study focuses on carbon cycling dynamics in tropical peatlands by examining carbon isotopic signatures (¹⁴C and ¹³C) and elemental compositions of bulk peat, dissolved organic carbon (DOC), carbon dioxide (CO₂), and methane (CH₄). Key parameters include radiocarbon ages and isotopic ratios (δ¹³C) of bulk peat, concentrations of carbon (%C) and nitrogen (%N), and radiocarbon content of porewater gases and dissolved organic carbon (DOC). The data provide insights into the vertical and spatial distribution of carbon sources and possible preservation and decomposition processes within tropical peat profiles, offering critical information for understanding carbon storage and greenhouse gas emissions in these ecosystems.This dataset is comprised of one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) carbon isotopic signatures (¹⁴C and ¹³C); (5) concentrations of carbon (%C) and nitrogen (%N); (6) radiocarbon content of porewater carbon dioxide (CO₂), and methane (CH₄) ; (7) porewater DOC; (8) bulk peat sampling protocol; (9) porewater sampling protocol; (10) porewater gas collection methods; and (11) gas extraction methods. All files are in .csv format and can be opened with any software that supports this file types.

54 ENVIRONMENTAL SCIENCES↗

High resolution characterization of soil dissolved organic matter with FTICR-MS (Fourier-transform ion cyclotron resonance mass spectrometry) from soil samples in control and warming plots in Blodgett Forest, CA (2014 and 2018)

The pathways of carbon transport and loss through and from soils—soil organic matter (SOM) depolymerization to dissolved organic carbon and mineralization to carbon dioxide (CO2)—are fundamentally driven by microbial activity, which is strongly regulated by environmental conditions. As part of Lawrence Berkeley National Laboratory (LBNL) Terrestrial Ecosystem Science (TES) Belowground Biogeochemistry Science Focus Area (SFA), we have established a novel whole-soil long-term warming experiment at the University of California (UC) Blodgett Forest Research Station (Sierra Nevada) in 2014, where we study the role of biogeochemical, microbial and geochemical process interactions in SOM decomposition and stabilization.This package contains Fourier transform ion cyclotron resonance mass spectrometry (21 Tesla FTICR-MS) data measured in negative and positive ionization mode from water and methanol soil extracts. Soil samples were collected in 2014/06/03 and 2018/06/04 from 3 replicated paired plots that had been subjected to experimental warming since June 2014 to simulate a predicted climate change scenario for northern California. The following files are included: (1) fticr_neg_h2oMeoh_data_raw.csv: raw data from combined water (H2O) and methanol (MeOH) extracts in negative ion mode, (2) fticr_neg_h2oMeoh_data_processed.csv: processed data from combined water (H2O) and methanol (MeOH) extracts in negative ion mode, (3) fticr_neg_metadata.csv: metadata for samples/measurements in negative ion mode, (4) fticr_pos_h2oMeoh_data_raw.csv: raw data from combined water (H2O) and methanol (MeOH) extracts in positive ion mode, (5) fticr_pos_h2oMeoh_data_processed.csv: processed data from combined water (H2O) and methanol (MeOH) extracts in positive ion mode, (6) fticr_pos_metadata.csv: metadata for samples/measurements in positive ion mode.

54 ENVIRONMENTAL SCIENCES↗

Data & Code from Phoenix CPPP Phase 2 Analysis

This data and code package supports the analysis presented in “Beyond Surface Cooling: Comprehensive Field Assessment of Reflective Pavement Thermal Performance in Phoenix, Arizona” and provides fully reproducible workflows for evaluating the thermal performance of cool pavement treatments in a hot urban environment. The dataset integrates multi-modal field measurements collected across residential and nonresidential settings, including mobile air temperature traverses, stationary air temperature monitoring, residential mean radiant temperature (MRT) measurements, subsurface temperature profiles, and controlled testbed observations. The data package contains raw and processed datasets in comma-separated value (CSV) format, accompanying metadata files describing site characteristics and measurement protocols, and R scripts (.R files) used for data cleaning, time synchronization, spatial and temporal matching, quality control filtering, statistical comparison, and figure generation. All analyses were conducted using R (version ≥ 4.2.0) with commonly available packages (e.g., tidyverse, lubridate, data.table, ggplot2). No proprietary software is required to reproduce results. Field campaigns were designed to quantify the effects of high-reflectance pavement coatings on surface temperature, near-surface air temperature, subsurface heat propagation, and radiative heat exposure. Temporal alignment procedures include standardized timestamp conversion and nearest-neighbor matching of high-frequency sensor measurements to stop-based metadata within defined tolerance windows to ensure comparability across instruments. The workflows generate summary statistics, treatment–control contrasts, depth-dependent thermal gradients, and time-series visualizations used in the associated publication. By integrating mobile, stationary, radiative, and subsurface measurements within a unified and transparent processing framework, this package enables comprehensive evaluation of cool pavement performance across multiple thermal exposure pathways and supports reuse in future urban heat mitigation and climate resilience studies.

AIR TEMPERATURE↗

Stream Chemistry, Synoptic Surveys, East Fork Poplar Creek Watershed, TN, USA; April 2023 to February 2025

Impacts of developed land cover on stream chemistry can be difficult to discern from natural variability, particularly in carbonate watersheds where weathering of urban infrastructure and lithology generate similar signatures. We evaluated how spatial patterns of stream chemistry varied across perennial and non-perennial tributaries spanning an urban-to-forested gradient in a mid-order, carbonate-dominated watershed. This data package contains a processed and compiled summary of stream chemistry and properties obtained from 12 synoptic surveys of 54 stream sites across the East Fork Poplar Creek watershed located near Oak Ridge, TN, United States. The sites include non-perennial tributaries, perennial tributaries, and the main stem and span forested to urban (highly developed) land cover gradients. The data package includes the processed and flagged chemical data (WaDE_SynopticSummary_FinalChemistry), metadata describing data flagging and analysis (WaDE_SynopticSummary_Metadata), information about each site and its contributing subcatchment (WaDE_SynopticSummary_SiteInformation), and a comparison of instrument and field detection limits used to determine method detection limits for the study (WaDE_SynopticSummary_DetectionLimitComparison). Stream chemistry includes stream parameters measured in situ using multiparameter probes (dissolved oxygen, pH, specific conductance, temperature) and solutes including nutrients (nitrate, ammonium, soluble reactive phosphorus), dissolved organic carbon, dissolved inorganic carbon, major cations (calcium, magnesium, potassium, sodium), major anions (chloride, sulfate), and a broad suite of minor and trace elements.

EARTH SCIENCE > TERRESTRIAL HYDROSPHERE > SURFACE ↗

HP-FLEX: Field demonstration of the semantics-driven configuration of a Model Predictive Control system to make heat pumps flexible

Model Predictive Control (MPC) has demonstrated significant potential for optimizing building operations and enabling demand flexibility. However, the widespread adoption of MPC is hindered by complex manual configuration and commissioning processes that must be conducted by control experts working alongside building operators. These challenges drive up costs and reduce scalability, particularly when technical human resources and building automation systems are limited, such as in small and medium commercial buildings (SMCBs). This paper demonstrates how semantic standards, specifically ASHRAE 223P, can accelerate the adoption of MPC applications for load flexibility in SMCBs. The authors present a replicable control framework, titled “HP-FLEX” that leverages a building’s semantic model to bootstrap the required data configuration for an MPC controller developed for optimizing heat pump systems as flexible grid resources. The semantic model helps streamline the deployment workflow, particularly for site setup, data/control commissioning, and model setup. This integration enhances portability, transferability, and scalability of the HP-FLEX MPC, which has been developed to support MPC-based supervisory HVAC controllers in SMCBs. Additionally, the paper details the required building and thermostat metadata information to enable the HP-FLEX MPC based on field demonstrations. The new workflow was tested in a small commercial building located in California, U.S., and demonstrated a load shifting performance of 9% based on a dynamic pricing signal that varies by the hour. This work provides a practical pathway for transitioning sophisticated building applications from custom to standardized semantic representations, supporting the broader adoption of advanced control strategies like MPC. The framework also establishes a foundation for evolving metadata requirements as applications mature while maintaining compatibility with industry standards.

Paul, Lazlo↗

Introduction

This report provides a comprehensive overview of metadata to describe sensor signals in wastewater treatment plants and methods to obtain such metadata. In this introduction, we explain the original motivation behind the MetaCO task group. This includes a description of historical challenges (data volume, data velocity) for which mature technology is now available, and newer challenges, which relate to data structure (data variety) and data quality (veracity). We conclude the chapter with an expression of gratitude to all involved.

Aguado, Daniel↗

Conclusions and outlook

In this chapter, we provide notes on the state of metadata collection and organization practices and reflect on the achievements of the MetaCO task group. We also include an utility perspective on how to use this report in daily data management practice and a short guide is provided to show how metadata practices can be initiated. The chapter, and the report, is concluded within an outlook, sketching future progress and opportunities on the horizon.

Aguado, Daniel↗

Automating the Analysis of Large Language Models Responses through Zero-Shot Question Answering

Recent advancements in Large Language Models (LLMs) have shown significant potential in various applications, yet their evaluation, particularly in zero-shot question answering scenarios, remains a challenging task. In this study, our objective was to explore precision metrics for Large Language Models (LLM) and design and implement a software pipeline to automatically evaluate LLMs' outputs under zero-shot question answering. Zero-shot question answering involves a model providing answers to questions about topics it hasn't seen during training. It leverages the principles of zero-shot learning by relying on semantic understanding and generalization from related knowledge. The data used was metadata from medical databases on congenital heart disease. We explored eleven LLM metrics and selected three for our evaluation: BLEU, BERTScore, and MoverScore. BLEU calculates a score based on the overlap of n-grams (contiguous sequences of n items, typically words) between the machine-generated translation and the reference translations. Higher BLEU scores indicate better correspondence between the machine-generated and human-generated translations. BERTScore is a metric used to evaluate the quality of machine-generated text by measuring the similarity of token embeddings produced by BERT (Bidirectional Encoder Representations from Transformers) between the generated text and reference text. MoverScore is a metric that quantifies the dissimilarity between the distributions of word embeddings from machine-generated text and reference text, emphasizing semantic similarity over exact token overlap. We also introduced HBKI, a composite metric summarizing these approaches. We tested five models —GPT-3, Llama-2, Gemini 1.5 Pro, Solar 10.7B, and Mixtral-8x7b. Our software pipeline, designed and implemented using Object-Oriented Programming principles, allows users to customize the selection and extraction of features for topics of interest in their own research. Our results show that MoverScore delivered the most precise evaluation of the LLM's outputs, while Mixtral-8x7b achieved the best overall performance in extracting metadata from the databases.

97 MATHEMATICS AND COMPUTING↗