Search NASA⌕ Search

SEARCH · Search NASA

Results for “metadata evaluation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Leveraging Pre-Built Catalogs and Object-Level Scheduling to Eliminate I/O Bottlenecks in HPC Environments

Modern High-Performance Computing (HPC) environments face mounting challenges due to the shift from large to small file datasets, along with an increasing number of users and parallelized applications. As HPC systems rely on Parallel File Systems (PFS), such as Lustre for data processing, performance bottlenecks stemming from Object Storage Target (OST) contention have become a significant concern. Existing solutions, such as LADS with its object-level scheduling approach, fall short in large-scale HPC environments due to their inability to effectively address metadata I/O bottlenecks and the growing number of I/O processes. This study highlights the pressing need for a comprehensive solution that tackles both OST contention and metadata I/O challenges in diverse HPC workloads. To address these challenges, we propose SwiftLoad, an object-level I/O scheduling framework that leverages a metadata catalog to enhance the performance and efficiency of parallel HPC utilities. The adoption of the metadata catalog mitigates the metadata I/O bottlenecks that commonly occur in HPC utilities, a challenge that is particularly pronounced in object-level I/O scheduling. SwiftLoad addresses OST contention and the uneven distribution of I/O processes across different OSTs through mathematical modeling and incorporates a Loader Configuration Module to regulate the number of I/O processes. Evaluated with two representative utilities—data deduplication profiling and data augmentation—SwiftLoad achieved performance improvements of up to 5.63x and 11.0x, respectively, on a production supercomputer.

HPC↗

Benchmarking DAOS Filesystem on Aurora

We benchmark the DAOS filesystem on Argonne's Aurora supercomputer (127 nodes, 4,064 targets) using fio, IOR, mdtest, and IO500 to characterize I/O and metadata performance across the DFS API and DFuse+POSIX. Single-client fio shows POSIX bandwidth saturating at 1–2 MiB I/O sizes, with write-heavy workloads outperforming reads. Multi-node IOR shows DFS bandwidth scaling well up to ~32 tasks/node, with write latency growing faster than read latency. An 8-node IO500 evaluation shows DFS achieving ~5x higher bandwidth and ~190x higher IOPS than POSIX. Results indicate DAOS is well-suited to read-heavy workloads like AI training data loading, given appropriately sized transfers and concurrency.

George, Rebecca [College of William and Mary, Will↗

AIACHNE's contribution for Nuclear Energy Agency Working Party on International Nuclear Data Evaluation Co-operation Subgroup 50

The AIACHNE (AI/ML Informed cAlifornium CHi Nuclear data Experiment) project aims at designing an experiment for the 252 Cf Prompt Fission Neutron Spectrum (PFNS) that explores systematic biases in an experimental database retrieved from the EXFOR databases. To that end, machine learning (ML) methods were applied to pint-point measurement features likely related to bias. From that information, we selected a feature that should be explored by the AIACHNE experiment. Measurement features are metadata encapsulating all pertinent information about the physical measurement and analysis techniques. Examples are, for instance, what neutron and fission detectors were used for the physical metadata, and what background reduction techniques were employed for analysis techniques. Such metadata were retrieved both from EXFOR entries as well as the literature of data sets described in detail in Reference 2 (at the end of the article).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Model Card for WaveDenoiser

This study used STEAD to train and evaluate this model because STEAD is among the best benchmark datasets available for local to regional data. STEAD is a global dataset with over 1 million 60 s long seismic waveforms that originated from approximately 450,000 earthquakes and background noise captured by more than 2,500 seismic stations. Each waveform in STEAD was attached with metadata such as earthquake locations, station locations, and signal arrival times when available

58 GEOSCIENCES↗

Towards Next-Generation Urban Decision Support Systems through AI-Powered Construction of Scientific Ontology Using Large Language Models—A Case in Optimizing Intermodal Freight Transportation

The incorporation of Artificial Intelligence (AI) models into various optimization systems is on the rise. However, addressing complex urban and environmental management challenges often demands deep expertise in domain science and informatics. This expertise is essential for deriving data and simulation-driven insights that support informed decision-making. In this context, we investigate the potential of leveraging the pre-trained Large Language Models (LLMs) to create knowledge representations for supporting operations research. By adopting ChatGPT-4 API as the reasoning core, we outline an applied workflow that encompasses natural language processing, Methontology-based prompt tuning, and Generative Pre-trained Transformer (GPT), to automate the construction of scenario-based ontologies using existing research articles and technical manuals of urban datasets and simulations. From these ontologies, knowledge graphs can be derived using widely adopted formats and protocols, guiding various tasks towards data-informed decision support. The performance of our methodology is evaluated through a comparative analysis that contrasts our AI-generated ontology with the widely recognized pizza ontology, commonly used in tutorials for popular ontology software. We conclude with a real-world case study on optimizing the complex system of multi-modal freight transportation. Our approach advances urban decision support systems by enhancing data and metadata modeling, improving data integration and simulation coupling, and guiding the development of decision support strategies and essential software components.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Processed Soil Respiration at the TRACE experimental Warming project, Aug 2015 - Sep 2017, Sabana, Luquillo, Puerto Rico

This data package contains processed measurements of soil carbon dioxide (CO₂) efflux collected using LI-COR LI-8100 soil respiration chambers at the Tropical Responses to Altered Climate Experiment (TRACE) located at the Sabana Field Research Station near Luquillo, Puerto Rico. The TRACE site is a mature, closed-canopy tropical wet forest within the Luquillo Experimental Forest. These data quantify soil surface CO₂ fluxes from both ambient (control) and experimentally warmed plots to evaluate how long-term soil warming affects belowground carbon cycling in tropical ecosystems. The data files include time-series tables of CO₂ flux (µmol CO₂ m⁻² s⁻¹), soil temperature (°C), and ancillary environmental variables, stored in comma-separated values (CSV) format and viewable with any text editor, spreadsheet, or statistical software (e.g., R, Python, Excel). Associated metadata describe plot identifiers, measurement intervals, and processing steps. These data were generated to address the research question: How does sustained soil warming influence soil respiration and carbon flux dynamics in tropical wet forests?

54 ENVIRONMENTAL SCIENCES↗

Performance and Reliability Assessment of the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) Data Advisor (ADA)

The Atmospheric Radiation Measurement (ARM) User Facility provides one of the world's largest openly accessible repositories of atmospheric observations through the ARM Data Discovery platform. Although the repository contains more than three decades of measurements collected from permanent observatories, mobile facilities, aircraft campaigns, and field experiments, identifying appropriate datasets can be challenging, particularly for new users unfamiliar with ARM instrumentation and datastream organization. To improve data accessibility, the ARM Data Center developed the ARM Data Advisor (ADA), an artificial intelligence-powered assistant designed to facilitate scientific data discovery, dataset interpretation, and user guidance. This report evaluates ADA's performance as a domain-specific scientific assistant using realistic atmospheric science workflows. The evaluation examines five key capabilities: data retrieval and curation efficiency, hallucination resistance, scientific reasoning, response to ambiguous queries, and content retention and session continuity. Representative prompts were developed to simulate typical interactions between researchers and the ARM Data Discovery platform, and ADA's responses were assessed for retrieval completeness, scientific accuracy, consistency, and practical usefulness. In these representative tests, ADA reduced the complexity of discovering and accessing ARM datasets by recommending appropriate datastreams, explaining instrumentation, interpreting metadata, and assisting with data processing workflows. ADA also exhibits strong domain knowledge of atmospheric science terminology and generally resists hallucination by acknowledging unavailable datasets and requesting clarification when appropriate. Overall, the results indicate that ADA represents a promising advancement in scientific data discovery within the ARM User Facility and has considerable potential to improve researcher productivity, particularly for new users and interdisciplinary scientists seeking efficient access to ARM observations.

Salvador, Christian [ORNL] (ORCID:0000000283287777↗

Fleet Utilization

A key goal of NextGen Profiles' fleet utilization study was to conduct a comprehensive, strategic, and standardized assessment of the operational behavior and utilization patterns across EV and EVSE production-ready fleets. These data-driven insights were intended to inform current fleet management strategies and support future infrastructure planning, ensuring the effective adoption and adaptation of the growing EV fleet market. The study applied a series of metrics defined in NextGen Profiles to evaluate diverse fleet operations across various use cases, emphasizing trends in charging, routing, and other critical behaviors. The fleet utilization dataset includes these three sets of metrics from 17 EV fleets, each consisting of a wide range of vehicle types and operational categories, as well as two EVSE fleets. Data were collected from a variety of sources and reformatted into a unified structure before metric computation, ensuring consistency and comparability across all fleets. To protect confidentiality, all fleet metadata are anonymized, and the publicly released metric datasets are aggregated to an hourly cadence.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Integrated Hourly Meteorological Database of 20 Meteorological Stations (1981-2022) for Watershed Function SFA Hydrological Modeling

This dataset contains (a) a script “R_met_integrated_for_modeling.R”, and (b) associated input CSV files: 3 CSV files per location to create a 5-variable integrated meteorological dataset file (air temperature, precipitation, wind speed, relative humidity, and solar radiation) for 19 meteorological stations and 1 location within Trail Creek from the modeling team within the East River Community Observatory as part of the Watershed Function Scientific Focus Area (SFA). As meteorological forcings varied across the watershed, a high-frequency database is needed to ensure consistency in the data analysis and modeling. We evaluated several data sources, including gridded meteorological products and field data from meteorological stations. We determined that our modeling efforts required multiple data sources to meet all their needs. As output, this dataset contains (c) a single CSV data file (*_1981-2022.csv) for each location (20 CSV output files total) containing hourly time series data for 1981 to 2022 and (d) five PNG files of time series and density plots for each variable per location (100 PNG files). Detailed location metadata is contained within the Integrated_Met_Database_Locations.csv file for each point location included within this dataset, obtained from Varadharajan et al., 2023 doi:10.15485/1660962. This dataset also includes (e) a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and (f) a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type. Review the (g) ReadMe_Integrated_Met_Database.pdf file for additional details on the script, methods, and structure of the dataset.The script integrates Northwest Alliance for Computational Science and Engineering’s PRISM gridded data product, National Oceanic and Atmospheric Administration’s NCEP-NCAR Reanalysis 1 gridded data product (through the `RCNEP` R package, Kemp et al., doi:10.32614/CRAN.package.RNCEP), and analytical-based calculations. Further, this script downscales the input data into hourly frequency, which is necessary for the modeling efforts.

54 ENVIRONMENTAL SCIENCES↗

Redox potential in Typha-dominated tidal brackish marsh, PIE LTER, Plum Island Sound, MA, June–December 2022

This dataset includes soil redox potential measurements collected at multiple depths within a tidal brackish wetland in the upper estuary of the Plum Island Ecosystems Long-Term Ecological Research site (PIE LTER), Plum Island Sound, Newbury, Massachusetts. Measurements were taken to evaluate temporal variation in redox potential in relation to hydrological events at three replicate locations. Data were recorded every 5 minutes using a Campbell Scientific Volt116 connected to a CR6 datalogger with SWAP instrument redox probes (ORP-30-4-B) and reference electrodes. Measurements were made at the AmeriFlux site US-PLo at four soil depths (5, 10, 15, and 30 cm). The file redox_soiltemp_2022.csv contains temperature-corrected redox values and soil temperature following Silva-Machado et al. (2024). Metadata files redox_soiltemp_dd.csv and redox_soiltemp_flmd.csv provide detailed descriptions of variables and site locations.

54 ENVIRONMENTAL SCIENCES↗

Redox potential in Typha-dominated tidal brackish marsh, PIE LTER, Plum Island Sound, MA, 2023

This dataset includes soil redox potential measurements collected at multiple depths within a tidal brackish wetland in the upper estuary of the Plum Island Ecosystems Long-Term Ecological Research site (PIE LTER), Plum Island Sound, Newbury, Massachusetts (MA). Measurements were taken to evaluate temporal variation in redox potential in relation to hydrological events at three replicate locations. Data were recorded every 5 minutes using a Campbell Scientific Volt116 connected to a CR6 datalogger with SWAP instrument redox probes (ORP-30-4-B) and reference electrodes. Measurements were made at the AmeriFlux site US-PLo at four soil depths (5, 10, 15, and 30 cm). The file redox_soiltemp_2023.csv contains temperature-corrected redox values and soil temperature following Silva-Machado et al. (2024). Metadata files redox_soiltemp_dd.csv and redox_soiltemp_2023_flmd.csv provide detailed descriptions of variables and site locations.

54 ENVIRONMENTAL SCIENCES↗

Web-based Preprocessing and Visualization of 3D FIB Tomography Data for Nuclear Fuel Characterization

Three-dimensional (3D) focused ion beam (FIB) tomography enables reconstruction of internal nuclear fuel features that can't be fully evaluated through surface imaging alone. This capability supports characterization of fuel constituents and defects under thermal and irradiation conditions relevant to microreactor development. However, large tomography datasets can create data-handling, loading, and visualization challenges, especially when image-stack preparation and file conversion must be completed with separate tools. The Computational Ultraspatial Tomography Toolkit for High-Resolution Object Analysis Tools (CUTTRHOAT) is an open-source web application being developed to display FIB tomography datasets available through the Nuclear Research Data System (NRDS). The current alpha version requires prepared HDF5 datasets and has limited integrated data-preparation capabilities. This project improves CUTTHROAT by adding dataset-folder selection, automatic input detection, dataset scanning, missing-slice identification, blank-slice insertion, and image-stack-to-HDF5 conversion. Two applications will be compared: the baseline CUTTHROAT alpha workflow and the updated application containing the integrated data-handling and preprocessing functions. Evaluation will consider dataset detection accuracy, conversion success, loading time, rendering responsiveness, application stability, and user interaction. Preliminary results demonstrate successful loading of existing HDF5 files and converted image stacks, while testing also identified performance reductions caused by excessive blank-slice generation. The updated workflow reduces reliance on external preparation tools and supports more direct movement from image stacks to color-code 3D visualization. Future work includes refining missing-slice handling, integrating additional preprocessing functions, like a denoising feature, parsing TIFF metadata for automatic voxel scaling, and adding manual X, Y, and Z voxel-spacing inputs for PNG and JPEG.

36 - MATERIALS SCIENCE↗

Lab Homes

This dataset includes processed data from the Lab Homes (LH) Test Facility located on the PNNL campus in Richland, WA. This a set of 2 identical homes that allow for the side-by-side comparison/performance evaluation of different technologies under the same weather at any given time. The dataset spans December 6, 2021 to December 27, 2021 and represents a series of tests performed; calibration, set-point excitation, pre-heating, free-floating and warm up. The measurements correspond to whole building electrical power, HVAC energy use, water heating, appliances and lighting, as well as space temperatures, space humidity, window glass surface temperatures, through glass solar radiation, and meterological data from an onsite meteorological weather station. In addition to the measurements, a metadata .json file, a .ttl file to visualize the data as per BRICK schema, and a detailed .pdf description of the dataset are also provided.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

PPI DataHub Project Data Package: S. elongatus PCC 7942 Circadian Control Bioproduction Transcriptomics (PB-DP3)

The purpose of this experiment was to evaluate how circadian clock regulation impacts carbon partitioning between storage, growth, and product synthesis in Synechococcus elongatus PCC 7942 in providing insights to strategies for enhanced bioproduction. Sample data was acquired using a Illumina HiSeq sequencer system and processed for RNA sequencing (RNA-Seq) expression analysis. Transcriptomic differential expression analysis revealed coordinated circadian clock-driven adjustment of the cell cycle and rewiring of energy and carbon metabolism. Processed RNA-Seq datasets are openly accessible from the PNNL DataHub project dataset download page and contain secondary processed RNA-seq results files and supporting metadata materials linked to relevant source code information supporting data transparency and reuse.

59 BASIC BIOLOGICAL SCIENCES↗

Geochemistry and Strontium Isotopes for Coal Creek Watershed, Colorado, 2021-2022

The geochemistry and strontium isotope data for Coal Creek Watershed, Colorado, consists of cation, anion, and 87Sr/87Sr isotope values from samples collected at 8 stream location along Coal Creek, samples from two groundwater springs within the watershed, and a shallow subsurface piezometer. All stream and spring samples were collected between June and October, 2021, and the shallow, near stream piezometer sample was collected in July of 2022. These data were collected to evaluate how groundwater contributions to Coal Creek originating from shallow vs deep flow paths respond seasonal drying. Understanding of groundwater-surface water interactions in montane systems in critical for the future of water availability in the Western US as groundwater contributions are expected to become more important for sustaining summer stream flows. This data package contains: (1) a csv of all cation samples; (2) a csv of all anion samples; (3) a csv of all 87Sr/87Sr isotope samples; and (4) a csv of locations for each sampling site. The dataset additionally includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.

54 ENVIRONMENTAL SCIENCES↗

Common practices for quantifying methane emissions from plumes detected by remote sensing

This document provides a set of community-accepted practices for quantifying methane emissions based on plumes detected via spectroscopic remote sensing. Its primary goal is to promote consistency in the generation, validation, reporting, and quality assessment of methane emission estimates derived from remote sensing radiances. Developed by subject matter experts with deep experience across all stages of the measurement process, this guidance reflects a critical evaluation of current methodologies and highlights key practices needed to produce reliable, interoperable, and traceable products. The focus is specifically on methane emissions quantified from distinct plumes originating from localized sources, rather than diffuse emissions spread over large regions, which are beyond the scope of this work. This document is intended to serve both data producers and users. For producers, it offers a framework for aligning with field-recognized standards to ensure their outputs meet rigorous quality and transparency criteria. For users, it provides a reference to assess dataset fitness-for-purpose by highlighting essential metadata, assumptions, and methodological choices that underpin emission estimates. By fostering a shared understanding of best practices, this work aims to enhance comparability, confidence, and utility of remotely sensed methane emission products.

54 ENVIRONMENTAL SCIENCES↗

Rhodotorula toruloides Nitrogen Limitation PTM Profiling Multi-Omics (TZ-DP1)

The purpose of this experiment was to evaluate the regulatory stress response of Oleaginous yeast species Rhodotorula toruloides NBRC 0880 (JGI strain IFO0880 v4.0) under nitrogen-rich and nitrogen-limited conditions over time. Time course experimental samples (0, 24, 48, and 72 hours after inoculation) were prepared using a semi-automated multi-PTM proteomic approach, using tandem mass tag 18-plex (TMT18), and lipidome remodeling for downstream multi-omics analysis. Processed datasets are openly accessible from PNNL DataHub and contain secondary processed proteomic (redox, phospho, and global TMT) and lipidomic (positive and negative ion mode) results files and experimental design metadata.

59 BASIC BIOLOGICAL SCIENCES↗

Stream Chemistry, Synoptic Surveys, East Fork Poplar Creek Watershed, TN, USA; April 2023 to February 2025

Impacts of developed land cover on stream chemistry can be difficult to discern from natural variability, particularly in carbonate watersheds where weathering of urban infrastructure and lithology generate similar signatures. We evaluated how spatial patterns of stream chemistry varied across perennial and non-perennial tributaries spanning an urban-to-forested gradient in a mid-order, carbonate-dominated watershed. This data package contains a processed and compiled summary of stream chemistry and properties obtained from 12 synoptic surveys of 54 stream sites across the East Fork Poplar Creek watershed located near Oak Ridge, TN, United States. The sites include non-perennial tributaries, perennial tributaries, and the main stem and span forested to urban (highly developed) land cover gradients. The data package includes the processed and flagged chemical data (WaDE_SynopticSummary_FinalChemistry), metadata describing data flagging and analysis (WaDE_SynopticSummary_Metadata), information about each site and its contributing subcatchment (WaDE_SynopticSummary_SiteInformation), and a comparison of instrument and field detection limits used to determine method detection limits for the study (WaDE_SynopticSummary_DetectionLimitComparison). Stream chemistry includes stream parameters measured in situ using multiparameter probes (dissolved oxygen, pH, specific conductance, temperature) and solutes including nutrients (nitrate, ammonium, soluble reactive phosphorus), dissolved organic carbon, dissolved inorganic carbon, major cations (calcium, magnesium, potassium, sodium), major anions (chloride, sulfate), and a broad suite of minor and trace elements.

EARTH SCIENCE > TERRESTRIAL HYDROSPHERE > SURFACE ↗