Search NASASearch

SEARCH · Search NASA

Results for “metadata evaluation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

DoCeph: DPU-Offloaded Messaging in Ceph for Reduced Host CPU Utilization

Ceph is a widely used distributed object store, but its messenger layer imposes substantial CPU overhead on the host. To address this limitation, we propose DoCeph, a DPU-offloaded storage architecture for Ceph that disaggregates the system by offloading the communication-intensive messaging component to the DPU while retaining the storage backend on the host. The DPU efficiently manages communication, using lightweight RPC for metadata operations and DMA for data transfer. Moreover, DoCeph introduces a pipelining technique that overlaps data transmission with buffer preparation, mitigating hardware-imposed transfer size limitations. We implemented DoCeph on a Ceph cluster with NVIDIA BlueField-3 DPUs. Evaluation results indicate that DoCeph cuts host CPU usage by up to 92% while sustaining stable throughput and providing larger performance benefits for object writes over 1 MB.

Park, Kuri [Sogang University]

Search Enhancements using Natural Language Processing Techniques

NASA Goddard Earth Sciences Data and Information Services Center (GESDISC) is one of the 12 NASA Science Mission Directorate Data Centers. The main goal of GESDISC is to provide earth science data, information, and services to the earth science data community. Consequently, data discovery is at the center of our mission and our search engine is the primary tool for our users to interact, find, and access our data. Existing search approaches are largely focused on hard-matching of keywords in the search query with dataset metadata. Here we propose to expand the search by introducing a complementary natural language processing (NLP) search. At the heart of our proposed NLP search, we trained a joint embedding using scientific text corpus and a curated set of dataset metadata. The embedding learns the association between words in our dataset metadata and those of the scientific text corpus. This enables us to go beyond simple hard-matching of a query and data set metadata and have a notion of “similarity” between the search query and the datasets. We further integrated our NLP search into the Elastic Search (ES) framework leveraging similarity search capabilities offered through the “dense_vector” field type. Our preliminary evaluations show that our proposed NLP search has the potential to be utilized to complement the existing search engine and serve as a base for a dataset recommendation system.

Armin Mehrabian

Monthly Mean In Situ Surface Flux Observations Paired with Satellite-Derived and Reanalysis-Based Flux Data for the Great Lakes Region, 2001–2020

Surface radiative and turbulent heat fluxes over the Great Lakes strongly influence regional hydrological and meteorological processes, and their accurate representation is critical for numerical weather prediction and coupled atmosphere–lake modeling. However, direct flux observations are spatially sparse across the region, so gridded reanalysis and satellite-derived products are often used for climatological analyses and model evaluation despite differences in their flux representations. This dataset provides processed, quality-controlled, monthly mean surface flux observations from the Great Lakes Evaporation Network (GLEN), AmeriFlux, and the National Data Buoy Center, paired with spatiotemporally matched flux estimates from two reanalysis products, the fifth generation European Centre for Medium-Range Weather Forecasts (ECMWF) reanalysis dataset (ERA5) and the Modern Era Reanalysis for Research and Applications, version 2 (MERRA-2), and two satellite-derived products, the Clouds and Earth's Radiant Energy Systems Energy Balanced and Filled (CERES-EBAF) and the Cloud, Albedo and Surface Radiation dataset from AVHRR data - Edition 3 (CLARA-A3). The dataset includes sixteen observational stations with variable temporal coverage within 2001–2020. For each station, a CSV file contains monthly time series of available flux variables, including surface downwelling shortwave radiation (SW), surface downwelling longwave radiation (LW), sensible heat (SH) flux, and latent heat flux (LH), alongside matched gridded product values where available. Columns in the CSV file correspond to different variables sourced from each dataset, with column titles structured as "{dataset}_{variable}". Columns with relevant metadata are also provided in each CSV file, including station latitude and longitude, monthly timestamps, and the name of the sourced observational data. These files are structured for direct use in common analysis tools, including Microsoft Excel, Python pandas, and Python matplotlib. This dataset supports climatological analysis of the Great Lakes regional surface energy budget, evaluation of satellite-derived and reanalysis-based flux products, and development or validation of flux representations in numerical weather prediction and coupled atmosphere–lake models.

Great Lakes

Diagnosing the representation of surface and layered soil moisture in Earth system models

Surface soil moisture (mrsos) and vertically integrated soil moisture (mrsol) over the top 10 cm should, by definition, be physically consistent in Earth System Models (ESMs). However, an evaluation of nine CMIP6 models reveals substantial inconsistencies: in some models, mrsos and integrated mrsol agree globally; in others, they align only in specific regions; and in a few, they diverge across all grid cells. These discrepancies arise from a combination of factors, including metadata errors, inconsistent variable definitions, or diagnostic sequencing within the model. We demonstrate how such issues can lead to significant biases, even when both variables are present and seemingly well-defined. As model complexity increases and multi-model comparisons become more common, assumptions about variable equivalence may lead to flawed conclusions. This study highlights the need for routine consistency checks, improved metadata standards, and community-wide practices that ensure reliability of derived variables across ESM outputs, particularly in preparation for CMIP7.

Earth system models

User Centered, Application Independent Visualization of National Airspace Data

This paper describes an application independent software tool, IV4D, built to visualize animated and still 3D National Airspace System (NAS) data specifically for aeronautics engineers who research aggregate, as well as single, flight efficiencies and behavior. IV4D was origin ally developed in a joint effort between the National Aeronautics and Space Administration (NASA) and the Air Force Research Laboratory (A FRL) to support the visualization of air traffic data from the Airspa ce Concept Evaluation System (ACES) simulation program. The three mai n challenges tackled by IV4D developers were: 1) determining how to d istill multiple NASA data formats into a few minimal dataset types; 2 ) creating an environment, consisting of a user interface, heuristic algorithms, and retained metadata, that facilitates easy setup and fa st visualization; and 3) maximizing the user?s ability to utilize the extended range of visualization available with AFRL?s existing 3D te chnologies. IV4D is currently being used by air traffic management re searchers at NASA?s Ames and Langley Research Centers to support data visualizations.

Murphy, James R.

DAISY: A Rapid Approach to Evaluating Marine Energy Converter Sound (Final Technical Report)

This project’s objective was to improve the quality of acoustic information about marine energy converters that could be collected from groups of drifting hydrophones, while reducing the costs of deployment and data analysis. This was achieved through technology development addressing four focus areas: (1) minimizing flow-noise and self-noise, (2) integrating metadata streams into a single data acquisition system, (3) developing post-processing routines to facilitate rapid data review, and (4) enabling objective identification of marine energy converter sound against a backdrop of ambient noise using time-delay-of-arrival localization.

16 TIDAL AND WAVE POWER

Expanding Biological Repository Data Available for Sharing and Knowledge Discovery

Biology has developed next-generation data science and alternative analytical approaches with methodologies which require principal investigator (PI) experimental assay data be re-used. This new approach involves mining multiple datasets at once from various hierarchical organizations of biological complexity, while concurrently evaluating how experimental factors affect endpoints of standard assays. The purpose of the NASA Ames Life Sciences Data Archive (ALSDA) is to collect, curate, and make findable, accessible, interoperable, and reusable (FAIR) all non-human space-relevant biological data. These data include mission metadata, subject metadata, assay metadata (parameters), raw and processed assay data, assay imagery, and subject-experienced telemetry (radiation, temperature, humidity, acoustics, vibrations). ALSDA has transformed to bring current biological repository data and all future collected data into this new scientific data mining reality. It has integrated into the ‘NASA Open Science’ group of projects to facilitate a suite of new tools and workflows to improve data accessibility and reusability by implementing data management plans, automating data submission agreements, and adopting the single-point-of-entry data submission portal, originally developed by NASA GeneLab. These systems required ALSDA to develop science assay configurations for the submission portal, capturing essential assay parameters according to established norms in each sub-field within biology. The submission portal expedites data collection by enhancing ease of PI data submission, providing a user interface and specificity for which data is to be submitted. ALSDA datasets are curated to maintain rich metadata, accuracy of datasets, data transparency, provenance, and additionally ensure data are machine-readable (e.g., R and Python languages). ALSDA integration with GeneLab and its analysis portals enable higher-order physiological-level datasets be mined in conjunction with -omics datasets. As ALSDA physiological-level datasets are published (micro-computed tomography, histology, intraocular pressure, hormonal assays, immunostaining, ultrasonography), the merging of hierarchical organizations of biological complexity from spaceflight will enable new knowledge discovery approaches.

Ryan T Scott

Model scripts associated with “Revisiting controls on hyporheic respiration with knowledge-guided machine learning at continental scale”

NOTE: The manuscript associated with this data package is currently in review. The data/scripts may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final scripts and additional metadata. This data package is associated with the publication “Revisiting controls on hyporheic respiration with knowledge-guided machine learning at continental scale” submitted to Environmental Science & Technology (Zheng et al. 2026). The project combines mechanistic process modeling with knowledge-guided machine learning (KGML) to evaluate how organic matter chemistry, microbial biomass, and physical substrate accessibility regulate realized respiration rates across river corridors. All data used in this paper have been previously published and can be accessed at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1729719 (Goldman et al., 2020). This data package contains 3 R-markdown (Rmd) preprocessing scripts for the previously published data and subsequent modelling workflows. The full workflow with input and output data can be found in the associated GitHub repository at https://github.com/jianqiuz/KGML-WHONDRS.

Biogeochemistry

Validation of ADAR System 5500 Digital Imagery: Delivery Task Order #1, Task Request #857 - Brookings, SD

This work was performed under NASA's Verification and Validation Program as an independent check of data supplied by Positive Systems, Inc. through the Earth Science Enterprise's Scientific Data Purchase (SDP) Program. This document serves as the basis for reporting results associated with validation of multispectral imagery according to the specifications of contract NAS 13-98049. The validation was performed under the Positive Systems Imaging System Validation Work Instruction CRSP-WI-28: Spectral registration, spatial resolution, endlaps, sidelaps, and image quality were evaluated. The validation was proceded by Shipment Verification, as described in the Work Instruction CRSP-WI-22: Every image was passed through an automatic ingest verification and thumbnail review process to identify omissions, problems with media integrity, and gross errors in data quality. Validation of metadata files is not within the scope of this report, but it was performed separately.

Blonski, Slawomir

Snow Depth Datasets for Snodgrass Catchment, Colorado, Water Year 2022-2023

This data package presents snow depths data from distributed temperature probes at 18 locations near Snodgrass catchment, Colorado. These data show that snow melt-out dates are approximately one or two weeks later under evergreen forests compared to other vegetation types even at the same elevation. These data were collected to understand how snowmelt heterogeneity impacts headwater hydrology, including streamflow and groundwater levels. They were also used to compare with process-based model simulations of snow depth to evaluate whether the model accurately represents snowmelt dynamics and their effects on headwater hydrology. Snow_DTPs_locations.csv includes all probes locations and their associated elevation and vegetation types. Snow_Depth_Snodgrass_WY2022_2023.csv includes processed snow depths datasets for Water Year (WY) 2022 and 2023. This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type. Several probes have recordings for WY 2021.

54 ENVIRONMENTAL SCIENCES

Remote sensing images, DEM, and point clouds associated with “Accuracy evaluation of cost-effective 3D reconstruction approaches for hydrobiogeochemical processes in non-perennial stream riverbeds”

This data package is associated with the publication “Accuracy evaluation of cost-effective 3D reconstruction approaches for hydrobiogeochemical processes in non-perennial stream riverbeds” published in Frontiers in Environmental Science, Environmental Informatics and Remote Sensing (Bao et al., 2026; doi: 10.3389/fenvs.2026.1725258). This data package includes the drone photos for a section of Umtanum Creek in Washington, Unted States. The photos were used to reconstruct the 3-dimensional (3D) digital elevation model (DEM) of the riverbed for the investigated stream section. The reconstruction results from four approaches are provided: (1) unoccupied aerial vehicle (UAV, colloquially known as drone) imagery-based Structure-from-Motion (SfM), (2) a machine learning-based 3D reconstruction model, Visual Geometry Grounded Deep Structure from Motion (VGGSfM), (3) Visual Geometry Grounded Transformer for long sequence of images (VGGT-Long), and (4) handheld smartphone LiDAR scanning. The ground truth measurements by tripod-mounted optical level kit and ground control points GPS locations for evaluating the accuracy of the four reconstruction approaches are also provided in this data package. A preliminary version of this data package was published in October 2025 at the time of manuscript submission. It was updated in March 2026, at the time of manuscript acceptance, to include additional metadata (this readme, data dictionary, and file level metadata). The data did not change. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) 8 folders; (2) the detailed flight configuration html files; (3) field metadata; (4) a readme; (5) a data dictionary; and (6) file-level metadata. The folders “2024_10_18_d01” and “2024_10_18_d02” contain the original drone photos for the two drone flights (d01 and d02) on October 18, 2024. The reconstruction results from each of the approaches are in the folders called “ODM_SfM”, “VGGSfM”, “VGGTLong”, and “LiDAR”. The ground truth measurements are in the folder called “optical_level_kit”. Lastly, results comparing the different approaches are in the folder called “comparisons”. All files are .csv, .html, .jpg, .obj, .txt, and .npy. For information on using the .obj and .npy files, see the readme files within the same folder as the files.

54 ENVIRONMENTAL SCIENCES

Leveraging Pre-Built Catalogs and Object-Level Scheduling to Eliminate I/O Bottlenecks in HPC Environments

Modern High-Performance Computing (HPC) environments face mounting challenges due to the shift from large to small file datasets, along with an increasing number of users and parallelized applications. As HPC systems rely on Parallel File Systems (PFS), such as Lustre for data processing, performance bottlenecks stemming from Object Storage Target (OST) contention have become a significant concern. Existing solutions, such as LADS with its object-level scheduling approach, fall short in large-scale HPC environments due to their inability to effectively address metadata I/O bottlenecks and the growing number of I/O processes. This study highlights the pressing need for a comprehensive solution that tackles both OST contention and metadata I/O challenges in diverse HPC workloads. To address these challenges, we propose SwiftLoad, an object-level I/O scheduling framework that leverages a metadata catalog to enhance the performance and efficiency of parallel HPC utilities. The adoption of the metadata catalog mitigates the metadata I/O bottlenecks that commonly occur in HPC utilities, a challenge that is particularly pronounced in object-level I/O scheduling. SwiftLoad addresses OST contention and the uneven distribution of I/O processes across different OSTs through mathematical modeling and incorporates a Loader Configuration Module to regulate the number of I/O processes. Evaluated with two representative utilities—data deduplication profiling and data augmentation—SwiftLoad achieved performance improvements of up to 5.63x and 11.0x, respectively, on a production supercomputer.

HPC

Benchmarking DAOS Filesystem on Aurora

We benchmark the DAOS filesystem on Argonne's Aurora supercomputer (127 nodes, 4,064 targets) using fio, IOR, mdtest, and IO500 to characterize I/O and metadata performance across the DFS API and DFuse+POSIX. Single-client fio shows POSIX bandwidth saturating at 1–2 MiB I/O sizes, with write-heavy workloads outperforming reads. Multi-node IOR shows DFS bandwidth scaling well up to ~32 tasks/node, with write latency growing faster than read latency. An 8-node IO500 evaluation shows DFS achieving ~5x higher bandwidth and ~190x higher IOPS than POSIX. Results indicate DAOS is well-suited to read-heavy workloads like AI training data loading, given appropriately sized transfers and concurrency.

George, Rebecca [College of William and Mary, Will

AIACHNE's contribution for Nuclear Energy Agency Working Party on International Nuclear Data Evaluation Co-operation Subgroup 50

The AIACHNE (AI/ML Informed cAlifornium CHi Nuclear data Experiment) project aims at designing an experiment for the 252 Cf Prompt Fission Neutron Spectrum (PFNS) that explores systematic biases in an experimental database retrieved from the EXFOR databases. To that end, machine learning (ML) methods were applied to pint-point measurement features likely related to bias. From that information, we selected a feature that should be explored by the AIACHNE experiment. Measurement features are metadata encapsulating all pertinent information about the physical measurement and analysis techniques. Examples are, for instance, what neutron and fission detectors were used for the physical metadata, and what background reduction techniques were employed for analysis techniques. Such metadata were retrieved both from EXFOR entries as well as the literature of data sets described in detail in Reference 2 (at the end of the article).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Model Card for WaveDenoiser

This study used STEAD to train and evaluate this model because STEAD is among the best benchmark datasets available for local to regional data. STEAD is a global dataset with over 1 million 60 s long seismic waveforms that originated from approximately 450,000 earthquakes and background noise captured by more than 2,500 seismic stations. Each waveform in STEAD was attached with metadata such as earthquake locations, station locations, and signal arrival times when available

58 GEOSCIENCES

Towards Next-Generation Urban Decision Support Systems through AI-Powered Construction of Scientific Ontology Using Large Language Models—A Case in Optimizing Intermodal Freight Transportation

The incorporation of Artificial Intelligence (AI) models into various optimization systems is on the rise. However, addressing complex urban and environmental management challenges often demands deep expertise in domain science and informatics. This expertise is essential for deriving data and simulation-driven insights that support informed decision-making. In this context, we investigate the potential of leveraging the pre-trained Large Language Models (LLMs) to create knowledge representations for supporting operations research. By adopting ChatGPT-4 API as the reasoning core, we outline an applied workflow that encompasses natural language processing, Methontology-based prompt tuning, and Generative Pre-trained Transformer (GPT), to automate the construction of scenario-based ontologies using existing research articles and technical manuals of urban datasets and simulations. From these ontologies, knowledge graphs can be derived using widely adopted formats and protocols, guiding various tasks towards data-informed decision support. The performance of our methodology is evaluated through a comparative analysis that contrasts our AI-generated ontology with the widely recognized pizza ontology, commonly used in tutorials for popular ontology software. We conclude with a real-world case study on optimizing the complex system of multi-modal freight transportation. Our approach advances urban decision support systems by enhancing data and metadata modeling, improving data integration and simulation coupling, and guiding the development of decision support strategies and essential software components.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION

Processed Soil Respiration at the TRACE experimental Warming project, Aug 2015 - Sep 2017, Sabana, Luquillo, Puerto Rico

This data package contains processed measurements of soil carbon dioxide (CO₂) efflux collected using LI-COR LI-8100 soil respiration chambers at the Tropical Responses to Altered Climate Experiment (TRACE) located at the Sabana Field Research Station near Luquillo, Puerto Rico. The TRACE site is a mature, closed-canopy tropical wet forest within the Luquillo Experimental Forest. These data quantify soil surface CO₂ fluxes from both ambient (control) and experimentally warmed plots to evaluate how long-term soil warming affects belowground carbon cycling in tropical ecosystems. The data files include time-series tables of CO₂ flux (µmol CO₂ m⁻² s⁻¹), soil temperature (°C), and ancillary environmental variables, stored in comma-separated values (CSV) format and viewable with any text editor, spreadsheet, or statistical software (e.g., R, Python, Excel). Associated metadata describe plot identifiers, measurement intervals, and processing steps. These data were generated to address the research question: How does sustained soil warming influence soil respiration and carbon flux dynamics in tropical wet forests?

54 ENVIRONMENTAL SCIENCES

Standardizing Interfaces for External Access to Data and Processing for the NASA Ozone Product Evaluation and Test Element (PEATE)

NASA's traditional science data processing systems have focused on specific missions, and providing data access, processing and services to the funded science teams of those specific missions. Recently NASA has been modifying this stance, changing the focus from Missions to Measurements. Where a specific Mission has a discrete beginning and end, the Measurement considers long term data continuity across multiple missions. Total Column Ozone, a critical measurement of atmospheric composition, has been monitored for'decades on a series of Total Ozone Mapping Spectrometer (TOMS) instruments. Some important European missions also monitor ozone, including the Global Ozone Monitoring Experiment (GOME) and SCIAMACHY. With the U.S.IEuropean cooperative launch of the Dutch Ozone Monitoring Instrument (OMI) on NASA Aura satellite, and the GOME-2 instrumental on MetOp, the ozone monitoring record has been further extended. In conjunction with the U.S. Department of Defense (DoD) and the National Oceanic and Atmospheric Administration (NOAA), NASA is now preparing to evaluate data and algorithms for the next generation Ozone Mapping and Profiler Suite (OMPS) which will launch on the National Polar-orbiting Operational Environmental Satellite System (NPOESS) Preparatory Project (NPP) in 2010. NASA is constructing the Science Data Segment (SDS) which is comprised of several elements to evaluate the various NPP data products and algorithms. The NPP SDS Ozone Product Evaluation and Test Element (PEATE) will build on the heritage of the TOMS and OM1 mission based processing systems. The overall measurement based system that will encompass these efforts is the Atmospheric Composition Processing System (ACPS). We have extended the system to include access to publically available data sets from other instruments where feasible, including non-NASA missions as appropriate. The heritage system was largely monolithic providing a very controlled processing flow from data.ingest of satellite data to the ultimate archive of specific operational data products. The ACPS allows more open access with standard protocols including HTTP, SOAPIXML, RSS and various REST incarnations. External entities can be granted access to various modules within the system, including an extended data archive, metadata searching, production planning and processing. Data access is provided with very fine grained access control. It is possible to easily designate certain datasets as being available to the public, or restricted to groups of researchers, or limited strictly to the originator. This can be used, for example, to release one's best validated data to the public, but restrict the "new version" of data processed with a new, unproven algorithm until it is ready. Similarly, the system can provide access to algorithms, both as modifiable source code (where possible) and fully integrated executable Algorithm Plugin Packages (APPs). This enables researchers to download publically released versions of the processing algorithms and easily reproduce the processing remotely, while interacting with the ACPS. The algorithms can be modified allowing better experimentation and rapid improvement. The modified algorithms can be easily integrated back into the production system for large scale bulk processing to evaluate improvements. The system includes complete provenance tracking of algorithms, data and the entire processing environment. The origin of any data or algorithms is recorded and the entire history of the processing chains are stored such that a researcher can understand the entire data flow. Provenance is captured in a form suitable for the system to guarantee scientific reproducability of any data product it distributes even in cases where the physical data products themselves have been deleted due to space constraints. We are currently working on Semantic Web ontologies for representing the various provenance information. A new web site focusing on consolidating informaon about the measurement, processing system, and data access has been established to encourage interaction with the overall scientific community. We will describe the system, its data processing capabilities, and the methods the community can use to interact with the standard interfaces of the system.

Tilmes, Curt A.