Search NASA⌕ Search

SEARCH · Search NASA

Results for “Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

GraphTango: A Hybrid Representation Format for Efficient Streaming Graph Updates and Analysis

Abstract Streaming graph processing performs batched updates and analytics on a time-evolving graph. The underlying representation format of the graph largely determines the throughputs of these updates and analytics phases. Existing representation formats usually employ variations of hash tables or adjacency lists. However, a recent study showed that the adjacency-list-based approaches perform poorly on heavy-tailed graphs, and the hash table-based approaches suffer on short-tailed graphs. We propose GraphTango, a hybrid representation format that provides excellent update and analytics throughput regardless of the graph’s degree distribution. GraphTango dynamically switches among three different formats based on a vertex’s degree: (i) Low-degree vertices store the edges directly with the neighborhood metadata, confining accesses to a single cache line, (2) Medium-degree vertices use adjacency lists, and (3) High-degree vertices use hash tables as well as adjacency lists. In this case, the adjacency list provides fast traversal during the analytics phase, while the hash table provides constant-time lookups during the update phase. We further optimized the performance by designing an open-addressing-based hash table that fully utilizes every fetched cache line. In addition, we developed a thread-local lock-free memory pool that allows fast growing/shrinking of the adjacency lists and hash tables in a multi-threaded environment. We evaluated GraphTango with the help of the SAGA-Bench framework and compared it with four other representation formats: Stinger, Degree-aware Robin Hood Hashing, and two adjacency list-based formats with different workload balancing scheme. On average, GraphTango provides 4.5x higher insertion throughput, 3.2x higher deletion throughput, and 1.1x higher analytics throughput over the next best format. Furthermore, we integrated GraphTango with the state-of-the-art graph processing frameworks DZiG and RisGraph. Compared to the vanilla DZiG and vanilla RisGraph , [ GraphTango + DZiG ] and [ GraphTango + RisGraph ] reduces the average batch processing time by 2.3x and 1.5x, respectively.

Ahmed, Alif↗

A Case Study of Multimodal, Multi-institutional Data Management for the Combinatorial Materials Science Community

Although the convergence of high-performance computing, automation, and machine learning has significantly altered the materials design timeline, transformative advances in functional materials and acceleration of their design will require addressing the deficiencies that currently exist in materials informatics, particularly a lack of standardized experimental data management. The challenges associated with experimental data management are especially true for combinatorial materials science, where advancements in automation of experimental workflows have produced datasets that are often too large and too complex for human reasoning. The data management challenge is further compounded by the multimodal and multi-institutional nature of these datasets, as they tend to be distributed across multiple institutions and can vary substantially in format, size, and content. Furthermore, modern materials engineering requires the tuning of not only composition but also of phase and microstructure to elucidate processing–structure–property–performance relationships. To adequately map a materials design space from such datasets, an ideal materials data infrastructure would contain data and metadata describing (i) synthesis and processing conditions, (ii) characterization results, and (iii) property and performance measurements. In this work, we present a case study for the low-barrier development of such a dashboard that enables standardized organization, analysis, and visualization of a large data lake consisting of combinatorial datasets of synthesis and processing conditions, X-ray diffraction patterns, and materials property measurements generated at several different institutions. While this dashboard was developed specifically for data-driven thermoelectric materials discovery, we envision the adaptation of this prototype to other materials applications, and, more ambitiously, future integration into an all-encompassing materials data management infrastructure.

36 MATERIALS SCIENCE↗

A semantics-driven framework to enable demand flexibility control applications in real buildings

Decarbonising and digitalising the energy sector requires scalable and interoperable Demand Flexibility (DF) applications. Semantic models are promising technologies for achieving these goals, but existing studies focused on DF applications exhibit limitations. These include dependence on bespoke ontologies, lack of computational methods to generate semantic models, ineffective temporal data management and absence of platforms that use these models to easily develop, configure and deploy controls in real buildings. This paper introduces a semantics-driven framework to enable DF control applications in real buildings. The framework supports the generation of semantic models that adhere to Brick and SAREF while using metadata from Building Information Models (BIM) and Building Automation Systems (BAS). The work also introduces a web platform that leverages these models and an actor and microservices architecture to streamline the development, configuration and deployment of DF controls. The paper demonstrates the framework through a case study, illustrating its ability to integrate diverse data sources, execute DF actuation in a real building, and promote modularity for easy reuse, extension, and customisation of applications. The paper also discusses the alignment between Brick and SAREF, the value of leveraging BIM data sources, and the framework's benefits over existing approaches, demonstrating a 75% reduction in effort for developing, configuring, and deploying building controls.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Explainable discrepancy checker and diagnosis for digital Twin-based supervisory control system

By virtually representing a physical object and process, a digital twin (DT) enables optimal autonomous operations by combining classical and novel frameworks in sensors, state predictions, and multi-input/multi-output systems. A DT’s values depend on how well models estimate quantities of interest and on how uncertainty is handled. Moreover, DTs often combine physics-based and data-driven models with mixed fidelities, where classical uncertainty quantification (UQ) struggles with many sources of uncertainty and real-time constraints. Here, this work presents a UQ-based discrepancy checking and diagnosis tool for a DT-based supervisory control system. The tool is developed using metadata from an automated DT development process to learn correlations between sources of uncertainties and outcomes. During operation, it compares predictions with measurements, attributes discrepancies to dominant sources, and recommends parameter and configuration updates. We verify the workflow on a synthetic temperature-control problem and deploy it on a virtual Thermal Energy Delivery System, reducing mismatch and improving control robustness.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

From models to reality: a systematic review on simulated and measured residential heat pump energy savings

High-performance HVAC solutions are central to residential energy management. A substantial share of these are electric, reversible-cycle systems, with heat pumps representing the largest portion of current and near-term adoption. This review synthesizes peer-reviewed and grey literature on residential space heating and cooling heat pumps. The academic literature is dominated by modeling (73.8%), with limited field measurement (13.1%). Grey literature from United States serve as a supplemental resource providing measured savings. Conversions from electric-resistance heating consistently show the largest site energy reductions, while oil/propane baselines yield moderate savings, and gas baseline scenario often deliver small and region-dependent savings. This study cross-checks the grey literature measured data with simulation data filtered from the ResStock dataset. The comparison indicates a discrepancy between simulations and measured data: simulated site EUIs are typically lower than measured EUIs, but percentage energy savings fall in similar ranges, implying simulations capture directional effects while underestimating energy use. Factors associated with variability and model–measurement differences include system characterization and control representation (e.g., backup heat engagement, thermostat/setpoint strategies, commissioning/installation quality), occupant behavior, weather normalization, metering scope, and envelope characterization. This paper also outlines the proposed methodology for comparing simulation and measured data for heat pumps. It emphasizes the metrics used for comparison and units harmonization, building characteristics matching, and compact metadata are needed for simulations to match measured data. The proposed methodology is expected to improve the credibility of simulated savings as measured evidence grows.

Yu, Lili↗

Envelope-driven comfort risk in residential demand response

Residential demand response (DR) is a valuable resource for grid reliability, but remains challenging because the highly heterogeneous residential building stock leads to widely varying and hard-to-predict load and comfort responses during DR events. Although prior research has estimated the technical potential of DR-capable technologies for achieving energy demand savings, little is known about how they affect thermal comfort. In particular, it remains unclear how indoor thermal conditions due to DR depend on the thermal envelope characteristics of the housing stock. To address this gap, this study provides a systematic, location-specific assessment of indoor thermal performance during DR-events across the US housing stock using both typical DR weather data and detailed building metadata. We evaluate how envelope characteristics influence indoor temperatures during realistic simulated summer and winter DR events across 37 US locations, applying both temperature threshold and rate of temperature change criteria to estimate region-level probabilities of discomfort. Additionally, we show the impact of distinct weather patterns that intensify or abate thermal stress on comfort outcomes. Results show a near-universal overheating risk in summer DR events, where comfort outcomes are strongly influenced by rapid risk of comfort violations. In contrast, overall winter DR discomfort risk is lower, risk escalation is more gradual and shows greater sensitivity to event duration. These findings offer a data-driven quantification of comfort risk across diverse climates and building envelopes, demonstrating the need for region-specific DR scheduling and discomfort mitigation strategies tailored to local weather patterns and the performance of existing residential buildings.

Demand response↗

Thermal property characterization of phase change materials in building applications: A systematic review of fundamentals, recent progress, and future directions

Phase change materials (PCMs) can reduce building peak loads and enable demand-responsive thermal energy storage (TES), but their deployment depends on reliable measurement and interpretation of thermal properties across laboratory, intermediate, and application scales. Here, this review systematically examines characterization methods, testing protocols, and recent advances for neat PCMs and PCM composites, emphasizing thermal conductivity, enthalpy-related properties (phase change temperature, latent heat, specific heat), and cycling stability. For thermal conductivity, we compare steady-state and transient techniques and note limitations when phase transition and contact resistance affect measurements. For enthalpy–temperature characterization, we discuss differential scanning calorimetry together with intermediate- and bulk-scale methods, including T-history, heat flow meter testing, and three-layer calorimetry (3LC), to generate application-relevant enthalpy–temperature profiles. Cycling stability is organized into four experimental families: thermoelectric–air, fully thermoelectric, water-bath, and in situ chamber approaches, with attention to separating reversible supercooling from true degradation such as phase segregation. We highlight emerging noncontact diagnostics, including infrared thermography and embedded sensing, for spatially resolved validation and multiscale interpretation. Finally, we review the growing use of AI and machine learning for property prediction, inverse characterization from experimental signals, and real-time state estimation in building-integrated TES. Key needs include harmonized protocols, interlaboratory benchmarking, uncertainty reporting, and metadata-rich datasets to accelerate reproducible PCM qualification for grid-flexible buildings.

AI↗

A terminology for scientific workflow systems

The term “scientific workflow” has evolved over the last two decades to encompass a broad range of compositions of interdependent compute tasks and data movements. It has also become an umbrella term for processing in modern scientific applications. Today, many scientific applications can be considered as workflows made of multiple dependent steps, and hundreds of workflow systems have been developed to manage and run these scientific workflows. However, no turnkey solution has emerged from the field to address the diversity of scientific processes and the infrastructure on which they are supposed to be implemented. Instead, new research problems requiring the execution of scientific workflows with some novel feature often lead to the development of an entirely new workflow system. A direct consequence of this situation is that many existing workflow management systems (WMSs) share some salient features, offer similar functionalities, and can manage the same categories of workflows but at the same time also have some distinct capabilities that can be important for specific applications. This situation makes researchers who develop workflows face the complex question of selecting a WMS. This selection can be driven by technical considerations, to find the system that is the most appropriate for their application and for the computing and storage resources available to them, or other factors such as reputation, adoption, strong community support, or long-term sustainability. To address this problem, a group of WMS developers and practitioners joined their efforts to produce a community-based terminology of WMSs. This paper summarizes their findings and introduces this new terminology to characterize WMSs. Furthermore, this terminology is composed of fives axes: workflow structure and characteristics, composition, orchestration, data management, and metadata capture. Each axis comprises several concepts that capture the prominent features of WMSs. Based on this terminology, this paper also presents a classification of 23 existing WMSs according to the proposed axes and terms.

Community-based terminology↗

Challenges and Vision for Standardization of Biopolymer Data Sets for Machine Learning

Machine learning (ML) is transforming materials research, yet potential for biopolymer discovery remains constrained by fragmented data and nonstandardized reporting. Biopolymers differ significantly from synthetic polymers, requiring specialized approaches to represent their biosynthetic origins, hierarchical structures, and application-specific metrics. In this Perspective, we identify three core challenges limiting biopolymer representation: information encoding, data quality, and data sharing. We describe the most pressing issues and propose commensurate approaches to address each key challenge. Recommendations include the design and adoption of biopolymer-specific fingerprinting and representation frameworks, development of hybrid human-large language model (LLM) data extraction strategies, and expanding Findable, Accessible, Interoperable, Reusable (FAIR)-compliant repositories. We propose a robust foundation to define interoperable, high-quality data sets that capture the full context of biopolymer materials. Standardized metadata, shared ontologies, and community-driven infrastructure would enable scalable, reproducible workflows and accelerate the ML-driven development of biopolymers.

36 MATERIALS SCIENCE↗

fluxfinder: An R Package for Reproducible Calculation and Initial Processing of Greenhouse Gas Fluxes From Static Chamber Measurements

Fluxes of greenhouse gases are a critical component of the earth's natural climate, but anthropogenic emissions have created an imbalance and resulted in global climate change. Quantifying the emission of these gases is vital to our understanding of their sources and sinks, both natural and anthropogenic. The static chamber method, in which a system of interest is enclosed, and gas concentrations are measured over time, is widely used to estimate fluxes of greenhouse gases. With the development of instruments such as infrared gas analyzers (IRGAs) supporting high-frequency concentration data, there is a growing need for open-source workflows to calculate fluxes. Here we present fluxfinder, an R package designed to support reproducible calculations and processing of greenhouse gas fluxes measured with the static chamber method. The package includes raw data file parsing from widely used IRGAs, metadata matching, unit conversion, flux estimations, and initial quality assurance/quality control (QA/QC). Diagnostic graphical plots provide a transparent way to differentiate between measurement issues and nonlinear behavior. The package is also designed to be easily integrated with the gasfluxes package for further fitting of nonlinear concentration-time models, allowing alternative or additional flux QA/QC. The fluxfinder package offers a flexible workflow that is easily adaptable to promote open and reproducible greenhouse gas flux estimations.

Wilson, Stephanie J.↗

Prediction of Distributed River Sediment Respiration Rates Using Community-Generated Data and Machine Learning

River sediment microbial respiration is a key indicator of ecosystem functioning and the biogeochemical fluxes across this critical zone link surface and subsurface waters. As such, there is tremendous interest in measuring and mapping these respiration rates. Respiration observations are expensive and labor intensive; there is limited data available to the community. An open science, collaborative initiative is collecting samples for respiration rate analysis and multi-scale metadata; this evolving data set is being used for making machine learning (ML) predictions at unsampled sites to help inform continued community engagement. However, it is a challenge to find an optimum configuration for ML models to work with this feature-rich (i.e., 100+ possible input variables) data set. Here, we present results from a two-tiered approach to managing the analysis of this complex data set: (a) a stacked ensemble of models that automatically optimizes hyperparameters and manages the training of many models and (b) feature permutation importance to detect the most important features in the models. The major elements of this workflow are modular, portable, open, and cloud-based thus making this implementation a potential template for other applications. The models developed here predict that sediment organic matter chemistry is one of the most important features for predicting sediment respiration rate. Other larger-scale, important features fall into the categories of climatic, ecological, geological, and fluvial settings. Leveraging these larger-scale features to generate data-driven estimates of river sediment respiration rates reveals spatially consistent but heterogeneous patterns across the river network of the Columbia River Basin.

54 ENVIRONMENTAL SCIENCES↗

A Performant, Scalable Processing Pipeline for High‐Quality and FAIR Environmental Sensor Data

High-resolution environmental monitoring is necessary to record, understand, and predict biogeochemical and ecological changes particularly in coastal systems but brings significant challenges in processing and making rapidly available the resulting data. The COMPASS-FME project established a network of coastal observational sites across the Chesapeake Bay and western Lake Erie regions extensively instrumented with soil, vegetation, and weather sensors logging data every 15 min. Our data processing framework, written in R and completely open source, prioritizes rapid model-experiment iteration and makes biogeochemical data rapidly available for quality assurance/quality control, analysis, and model ingestion. This pipeline is distinguished by a standardized and modular approach to data curation, extensive metadata and documentation, and its high performance. These attributes combine to make biogeochemical data rapidly accessible across COMPASS-FME and the broader community. Flexible, powerful, and reproducible approaches to handling high-volume environmental data are crucial for accelerating biogeosciences research.

Pennington, Stephanie C. [Pacific Northwest Nation↗

Predicting the heat release variability of Li-ion cells under thermal runaway with few or no calorimetry data

Accurate measurement of the variability of thermal runaway behavior of lithium-ion cells is critical for designing safe battery systems. However, experimentally determining such variability is challenging, expensive, and time-consuming. Here, we utilize a transfer learning approach to accurately estimate the variability of heat output during thermal runaway using only ejected mass measurements and cell metadata, leveraging 139 calorimetry measurements on commercial lithium-ion cells available from the open-access Battery Failure Databank. We show that the distribution of heat output, including outliers, can be predicted accurately and with high confidence for new cell types using just 0 to 5 calorimetry measurements by leveraging behaviors learned from the Battery Failure Databank. Fractional heat ejection from the positive vent, cell body, and negative vent are also accurately predicted. We demonstrate that by using low cost and fast measurements, we can predict the variability in thermal behaviors of cells, thus accelerating critical safety characterization efforts.

25 ENERGY STORAGE↗

Enabling pan-repository reanalysis for big data science of public metabolomics data

Public untargeted metabolomics data is a growing resource for metabolite and phenotype discovery; however, accessing and utilizing these data across repositories pose significant challenges. Therefore, here we develop pan-repository universal identifiers and harmonized cross-repository metadata. This ecosystem facilitates discovery by integrating diverse data sources from public repositories including MetaboLights, Metabolomics Workbench, and GNPS/MassIVE. Our approach simplified data handling and unlocks previously inaccessible reanalysis workflows, fostering unmatched research opportunities.

El Abiead, Yasin↗

Resolving root causes of experiment discrepancies guided by machine learning

Abstract Scientists rely on accurate experimental data to explain nature and then harness this knowledge for applications addressing human needs. However, discrepancies between experiments of the same observable can impede scientific progress if one does not understand the underlying causes. Here, we developed a process that unravels data discrepancies by first using Bayesian machine learning to relate discrepancies to few of many, potentially biasing metadata features that encode experiment procedures. This machine learning output guides human experts to study discrepancy causes by simulating suspicious aspects of historical experiments or designing modern ones to address open questions. The study findings then lead to rejecting or correcting historical data on firm scientific bases. This process is demonstrated for the energy spectrum of neutrons emitted promptly (<1 ns) after fission of 252 Cf, a trusted nuclear physics Standard. It reduces the spread in experimental 252 Cf spectra by up to a factor of 6.

Neudecker, D. (ORCID:0000000339200627)↗

The Arctic Plant Aboveground Biomass Synthesis Dataset

Plant biomass is a fundamental ecosystem attribute that is sensitive to rapid climatic changes occurring in the Arctic. Nevertheless, measuring plant biomass in the Arctic is logistically challenging and resource intensive. Lack of accessible field data hinders efforts to understand the amount, composition, distribution, and changes in plant biomass in these northern ecosystems. Here, we present The Arctic plant aboveground biomass synthesis dataset, which includes field measurements of lichen, bryophyte, herb, shrub, and/or tree aboveground biomass (g m -2 ) on 2,327 sample plots from 636 field sites in seven countries. We created the synthesis dataset by assembling and harmonizing 32 individual datasets. Aboveground biomass was primarily quantified by harvesting sample plots during mid- to late-summer, though tree and often tall shrub biomass were quantified using surveys and allometric models. Each biomass measurement is associated with metadata including sample date, location, method, data source, and other information. This unique dataset can be leveraged to monitor, map, and model plant biomass across the rapidly warming Arctic.

54 ENVIRONMENTAL SCIENCES↗

Aerial imagery dataset of lost oil wells

Orphaned wells are wells for which the operator is unknown or insolvent. The location of hundreds of thousands of these wells remain unknown in the United States alone. Cost-effective techniques are essential to locate orphaned wells to address environmental problems. In this paper, we present a dataset consisting of 120,948 aerial images of recently documented orphan wells. Each of these 512 × 512 images is paired with segmentation masks that indicate the presence or absence of such well. These images, sourced from the National Agriculture Imagery Program, cover the continental United States with spatial resolutions ranging from 30 centimeters to 1 meter. Additionally, we included negative examples by selecting locations uniformly across the United States. Accompanying metadata includes the IDs and spatial resolution of the original images, which are available for free through the United States Geological Survey, and the pixel coordinates of documented orphaned wells identified in these images. This dataset is intended to support the development of deep-learning models that can help locating undocumented orphan wells from such imagery, thereby blunting the environmental damage they do.

Climate-change mitigation↗

International database of reference gamma spectra for nuclear safeguards applications

Nuclear safeguards missions use gamma spectroscopy as a non-destructive measurement technique for examining nuclear materials. Despite advances in the development of detection equipment as well as software codes, one of the concerns is the lack of well-documented spectra needed to test and validate isotopic analysis codes for their applicability. To address this need, this work introduces IDB, an international database of the reference gamma spectra of uranium, plutonium and mixed oxide nuclear material samples. IDB provides access to well-characterized sets of gamma spectra described by rich metadata, including information on the sample, measurement configuration and detector specifications. These spectra are accessible in different formats, also compatible with analysis code standards, thus promoting their sustainability and maintenance.

None, Dipti [International Atomic Energy Agency (I↗