Search NASA⌕ Search

SEARCH · Search NASA

Results for “Science Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

The Vertebrate Breed Ontology: Toward Effective Breed Data Standardization

Abstract Background Limited universally-adopted data standards in veterinary medicine hinder data interoperability and therefore integration and comparison; this ultimately impedes the application of existing information-based tools to support advancement in diagnostics, treatments, and precision medicine. Hypothesis/Objectives A single, coherent, logic-based standard for documenting breed names in health, production, and research-related records will improve data use capabilities in veterinary and comparative medicine. Animals No live animals were used. Methods The Vertebrate Breed Ontology (VBO) was created from breed names and related information compiled from the Food and Agriculture Organization of the United Nations, breed registries, communities, and experts, using manual and computational approaches. Each breed is represented by a VBO term that includes breed information and provenance as metadata. VBO terms are classified using description logic to allow computational applications and Artificial Intelligence–readiness. Results VBO is an open, community-driven ontology representing over 19 500 livestock and companion animal breed concepts covering 49 species. Breeds are classified based on community and expert conventions (e.g., cattle breed) and supported by relations to the breed's genus and species indicated by National Center for Biotechnology Information (NCBI) Taxonomy terms. Relationships between VBO terms (e.g., relating breeds to their foundation stock) provide additional context to support advanced data analytics. VBO term metadata includes synonyms, breed identifiers/codes, and attributed cross-references to other databases. Conclusion and Clinical Importance The adoption of VBO as a standard for breed names in databases and veterinary electronic health records enhances veterinary data interoperability and computability, supporting precision medicine.

Veterinary Sciences↗

Human Host Cellular Response to HCoV-229E Infection Proteomics (ACS-JM-DP2)

The purpose of this experiment was to evaluate the human host cellular response to wild-type Human coronavirus strain 229E (HCoV-229E) infection. Sample data was obtained for mock and infected immortalized human lung epithelial cells (A549) (MOI 5) nuclear extracts, immortalized human lung fibroblasts cells (MRC5) (MOI5) nuclear extracts, and primary human airway epithelial (HAE) (MOI 3) cells from lung tissue and processed for proteome analysis. Processed datasets are openly accessible from the download button and contain secondary processed proteomic results files and supporting metadata materials. Experimental proteomics samples were prepared using Limited Proteolysis (LiP) methods for Label-free quantification (LFQ) and global proteomic evaluation. Sample data was acquired using a Q-Exactive HF-X mass spectrometer and was processed and compiled using MaxQuant software (v.1.6.17.0). Processed proteomic data downloads include a sample naming key, processed MaxQuant results/parameters, and protein annotated relative abundance files. See corresponding primary data accessions below and Viral Experiment LiP Analysis source code supporting data transparency and reuse. Experimental transcriptomics samples were collected in parallel and processed for RNA sequencing (RNA-Seq) as summarized under ACS-DP1 (https://data.pnnl.gov/group/nodes/dataset/34069).

59 BASIC BIOLOGICAL SCIENCES↗

Linking Spatiotemporal Biological Data to Predict Harmful Algal Blooms

Cyanobacterial Harmful Algal Blooms (cHABs) have significant impacts on an affected region’s economy, ecology, and human health. The blooms can release toxins that kill fish and poison water for people and animals. The global adverse effects of cHABs are exacerbated by the consequences of climate change and increased pollution. Though the phenomena are well documented, scientists’ efforts to mitigate the damage are hampered by insufficient predictive models and incomplete granular knowledge of cHAB community structure. With a goal of leveraging bioinformatics and machine learning tools to better understand and predict cHABs, we are first exploring water sample data sets. Using nearly four thousand samples from the National Center for Biotechnology Information Sequence Read Archive (NCBI-SRA) across 16 years with latitude and longitude embedded in the metadata, we mapped the location of the samples onto a Lake Erie shape file. We combined information about location, date, and community taxa in the NCBI samples to discover factors that determine cHAB features. The data are separated into three distinct zones, with the majority pooled at the southwest end of the lake and occurring in 2017. The samples are rich in biological data; our next steps are to carry out whole genome sequence analysis and use the community profiles as part of our predictive machine learning model.

59 BASIC BIOLOGICAL SCIENCES↗

RhizoGrid Indexed Sorghum Rhizosphere Multi-Omics

PerCon SFA project data dentification of spatially resolved biomarkers of drought in Sorghum bicolor rhizosphere molecular-microbe interactions using a novel root cartography "RhizoGrid" system for sampling plants under drought and control conditions across 10 equally sized root zone environments (4 quadrants each). Each quadrant was sampled and processed for 16S amplicon, metabolomics, and X-ray computed tomography (XCT). Data download includes experimental metadata and results files for 16S rRNA sequence analysis of microbial community assembly (processed data files), liquid chromatography mass spectrometry (LC-MS) metabolomics analysis of microbial community root exudates (processed data files), X-ray computed tomography (XCT) spatial gradient analysis (raw and processed data files) of microbial community composition, and related computational modeling outputs.

59 BASIC BIOLOGICAL SCIENCES↗

An Open Benchmark of One Million High-Fidelity Cislunar Trajectories

Cislunar space spans from geosynchronous altitudes to beyond the Moon and will underpin future exploration, science, and security operations. We describe and release an open dataset of one million numerically propagated cislunar trajectories generated with the open-source Space Situational Awareness Python package (SSAPy). The model includes high-degree Earth/Moon gravity, solar gravity, and Earth/Sun radiation pressure; other planetary gravities are omitted by design for computational efficiency. Initial conditions uniformly sample commonly used osculating-element ranges, and each trajectory is propagated for up to six years under a single, fixed start epoch. The dataset is intended as a reusable benchmark for method development (e.g., space domain awareness, navigation, and machine-learning pipelines), a reference library for statistical studies of orbit families, and a starting point for community-driven extensions (e.g., alternative epochs). We report empirically observed stability trends (e.g., a band near ~5 GEO and persistence of some co-orbital classes including L4/L5 librators) as dataset descriptors rather than new dynamical results. The chief contribution is the scale, fidelity, organization (CSV/HDF5 with full state time series and metadata), and open availability, which together lower the barrier to comparative and data-driven studies in the cislunar regime.

79 ASTRONOMY AND ASTROPHYSICS↗

Carbon dioxide, water vapor and methane soil efflux (soil respiration) in a Pinus palustris restoration site in Georgetown, SC

This dataset contains processed data from a combination of survey flux chambers and long-term automated flux chambers. Biweekly soil flux measurements were conducted from June 2023 through December 2025 at a longleaf pine restoration site in Georgetown, SC. Processed, QAQC’d data can be found in the file: 1_DATA_ESS_DOE_HR_RS_HB3_QAQC_Survey_Data_20260223.csv. Raw and working data files (.json, & .81x format) from LI-COR equipment are included for reference and can be accessed using SoilFluxPro software. CSV metadata files describe the raw data and modifications made using SoilFluxPro v5 and Matlab R2024b, as well as formatting and units for processed CSVs. Matlab code is included for reading in the processed CSVs. This research was performed as part of the project: “Improving models of stand and watershed carbon and water fluxes with more accurate representations of soil-plant-water dynamics in southern pine ecosystems”, which examines in part the effects hydraulic redistribution on soil efflux of carbon dioxide, water vapor and methane, as well as soil moisture and temperature in a southern pine ecosystem with sandy soils and high water table.

CARBON DIOXIDE FLUX↗

Towards Next-Generation Urban Decision Support Systems through AI-Powered Construction of Scientific Ontology Using Large Language Models—A Case in Optimizing Intermodal Freight Transportation

The incorporation of Artificial Intelligence (AI) models into various optimization systems is on the rise. However, addressing complex urban and environmental management challenges often demands deep expertise in domain science and informatics. This expertise is essential for deriving data and simulation-driven insights that support informed decision-making. In this context, we investigate the potential of leveraging the pre-trained Large Language Models (LLMs) to create knowledge representations for supporting operations research. By adopting ChatGPT-4 API as the reasoning core, we outline an applied workflow that encompasses natural language processing, Methontology-based prompt tuning, and Generative Pre-trained Transformer (GPT), to automate the construction of scenario-based ontologies using existing research articles and technical manuals of urban datasets and simulations. From these ontologies, knowledge graphs can be derived using widely adopted formats and protocols, guiding various tasks towards data-informed decision support. The performance of our methodology is evaluated through a comparative analysis that contrasts our AI-generated ontology with the widely recognized pizza ontology, commonly used in tutorials for popular ontology software. We conclude with a real-world case study on optimizing the complex system of multi-modal freight transportation. Our approach advances urban decision support systems by enhancing data and metadata modeling, improving data integration and simulation coupling, and guiding the development of decision support strategies and essential software components.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Carbon dioxide, water vapor and methane soil efflux (soil respiration) in a Pinus palustris root exclusion in Georgetown, SC

This dataset contains processed data from a combination of survey flux chambers and long-term automated flux chambers. Soil flux measurements were conducted from June 2023 through December 2025 in a mature longleaf pine forest in Georgetown, SC. Soil respiration measurements were conducted approximately biweekly for two and a half years, before and after a root exclusion that took place on May 5, 2024. Processed, QAQC’d data for the treatment (root exclusion) and control (roots intact) before and after the root exclusion can be found in the file: 1_DATA_ESS_DOE_HR_RS_HB2_QAQC_Survey_Data_20260223.csv. Two multiday deployments were also conducted prior to the root exclusion using long-term automated chambers to continuously monitor greenhouse gas soil efflux. Processed, QAQC’d data for both long-term deployments can be found in the file: 2_DATA_ESS_DOE_HR_RS_HB2_QAQC_Longterm_Data_20260209.csv. Raw and working data files (.json, .81x, & .82z format) from LI-COR equipment are included for reference and can be accessed using SoilFluxPro software. CSV metadata files describe the raw data and modifications made using SoilFluxPro v5 and Matlab R2024b, as well as formatting and units for processed CSVs. Matlab code is included for reading in the processed CSVs, with sample figures comparing treatment and control. This research was performed as part of the project: “Improving models of stand and watershed carbon and water fluxes with more accurate representations of soil-plant-water dynamics in southern pine ecosystems”, which examines in part the effects hydraulic redistribution on soil efflux of carbon dioxide, water vapor and methane, as well as soil moisture and temperature in a southern pine ecosystem with sandy soils and high water table.

CARBON DIOXIDE FLUX↗

Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification" Willard et al. (2025).

This data release provides all data and code used in the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025)" to model stream temperature, evaluate, and assess results. The associated manuscript explores the effect of different ensemble construction techniques across different common machine learning (ML) architectures for predictions in unmonitored basins. Modeling was done using long short-term memory (LSTM), gated recurrent unit (GRU), temporal convolution network (TCN), and extreme gradient boosting (XGBoost) models, and stream site coverage spans 1362 locations across the conterminous United States. The ensemble construction techniques investigated include ensemble by random weight initialization, differing hyperparameters, different random subsets of training data, different subselections of input features, different architectures, and Monte Carlo Dropout. The data is organized into these items items:Code repository and data for the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025).Code: stream_temp_ml_regionalization.zip contains the code repositoryData to run the code:- data_dir.zip -- contains all files that should be moved to the "DATA_DIR" variable defined in the "set_env_vars.sh" script in the code repository- metadata_dir.zip -- contains all files that should be moved to the "METADATA_DIR" variable defined in the "set_env_vars.sh" script in the code repositoryData produced by the code and used in the paper:- outputs_dir.zip - contains model output and results (outputs_dir/results), model weights (outputs_dir/models), and all other outputs used for the paper including feature importances.To cite this code, please use the following BibTeX or MLA entries:bibtex:@misc{willard2025streamensembles,author = {Jared Willard and Charuleka Varadharajan},title = {Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification"},year = {2024},doi = {10.15485/2527393},publisher = {ESS-DIVE Repository},url = {https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2527393}}MLA: Willard, Jared, et al. Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification". 2025. ESS-DIVE Repository, doi:10.15485/2448016.

54 ENVIRONMENTAL SCIENCES↗

Microbial community data from throughfall exclusion experiment: Metadata, SI, community composition, LefSe, and FunGuilR data tables from PARCHED Panama tropical forest soils, 2024-2025

Soil contains more carbon (C) than terrestrial vegetation and the atmosphere combined, with some of the largest terrestrial C stocks in tropical rainforests. Soil microbes decompose organic matter, playing a vital role in the storage or loss of soil C. With climate change, drought conditions are predicted to increase in many tropical regions, including both chronic drying and extended drought, potentially influencing these processes. This project explored the effects of chronic and seasonal drying on soil microbial communities across four distinct tropical forests in a long-term drying experiment. We investigated the effects of a chronic drying manipulation on soil microbial community abundance and variation across different forests and seasons. We also compared findings with previously published data from these forests after short-term drying. This project used soils from a long-term drying experiment established in 2018 across four seasonal lowland forests in Panama. Soils were collected from 0 – 10 cm depths during three seasonal periods in control and drying plots in 2024 and 2025 from a total of 32 plots (n = 4 per forest per treatment). The forests varied in baseline rainfall and soil fertility. We calculated alpha and beta diversity indices and compared taxonomic community composition. We found significant biogeographic variation in microbial diversity and taxonomy, with significant differences across the forests and significant effects of the drying treatment. Metadata and sample IDs are within Metadata_16S.csv and Metadata_ITS.csv. Relative abundance tables of every sample at every season are shown in the Excel workbooks 16S Relative Abundance.xlsx and ITS Relative Abundance.xlsx. They are then also shown in CSV files by each taxonomic level. Linear discriminant analysis effect size (LefSe) tables are shown for the full 16S and ITS datasets (n = 96), subsets for every site at every season (n = 8), and then for the forests with each plot merged by season (n = 8). FunGuildR data table of ITS data is uploaded.

Bacteria↗

BGC Atlas: a web resource for exploring the global chemical diversity encoded in bacterial genomes

Secondary metabolites are compounds not essential for an organism’s development, but provide significant ecological and physiological benefits. These compounds have applications in medicine, biotechnology and agriculture. Their production is encoded in biosynthetic gene clusters (BGCs), groups of genes collectively directing their biosynthesis. The advent of metagenomics has allowed researchers to study BGCs directly from environmental samples, identifying numerous previously unknown BGCs encoding unprecedented chemistry. Here, we present the BGC Atlas (https://bgc-atlas.cs.uni-tuebingen.de), a web resource that facilitates the exploration and analysis of BGC diversity in metagenomes. The BGC Atlas identifies and clusters BGCs from publicly available datasets, offering a centralized database and a web interface for metadata-aware exploration of BGCs and gene cluster families (GCFs). We analyzed over 35 000 datasets from MGnify, identifying nearly 1.8 million BGCs, which were clustered into GCFs. The analysis showed that ribosomally synthesized and post-translationally modified peptides are the most abundant compound class, with most GCFs exhibiting high environmental specificity. We believe that our tool will enable researchers to easily explore and analyze the BGC diversity in environmental samples, significantly enhancing our understanding of bacterial secondary metabolites, and promote the identification of ecological and evolutionary factors shaping the biosynthetic potential of microbial communities.

59 BASIC BIOLOGICAL SCIENCES↗

Stream Chemistry, Synoptic Surveys, East Fork Poplar Creek Watershed, TN, USA; April 2023 to February 2025

Impacts of developed land cover on stream chemistry can be difficult to discern from natural variability, particularly in carbonate watersheds where weathering of urban infrastructure and lithology generate similar signatures. We evaluated how spatial patterns of stream chemistry varied across perennial and non-perennial tributaries spanning an urban-to-forested gradient in a mid-order, carbonate-dominated watershed. This data package contains a processed and compiled summary of stream chemistry and properties obtained from 12 synoptic surveys of 54 stream sites across the East Fork Poplar Creek watershed located near Oak Ridge, TN, United States. The sites include non-perennial tributaries, perennial tributaries, and the main stem and span forested to urban (highly developed) land cover gradients. The data package includes the processed and flagged chemical data (WaDE_SynopticSummary_FinalChemistry), metadata describing data flagging and analysis (WaDE_SynopticSummary_Metadata), information about each site and its contributing subcatchment (WaDE_SynopticSummary_SiteInformation), and a comparison of instrument and field detection limits used to determine method detection limits for the study (WaDE_SynopticSummary_DetectionLimitComparison). Stream chemistry includes stream parameters measured in situ using multiparameter probes (dissolved oxygen, pH, specific conductance, temperature) and solutes including nutrients (nitrate, ammonium, soluble reactive phosphorus), dissolved organic carbon, dissolved inorganic carbon, major cations (calcium, magnesium, potassium, sodium), major anions (chloride, sulfate), and a broad suite of minor and trace elements.

EARTH SCIENCE > TERRESTRIAL HYDROSPHERE > SURFACE ↗

Strategies for community-sourced biocuration in bioinformatics: a case study on MIBiG 4.0

Biocuration is essential to transform molecular sequence data into standardized, machine-readable resources. Such curated datasets enable comparative analysis, predictive modeling, and data integration across bioinformatics platforms. While professional biocuration is resource-intensive and usually limited to institutional settings, community-driven approaches can mobilize large-scale annotation of specialized datasets and are more resilient to disruptions in scientific funding. Here, we present a model for community-powered curation applied to the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) repository. Through a framework of workflows for metadata capture, annotation validation, and contributor coordination, the MIBiG 4.0 initiative recruited 267 scientists across 178 institutions from 33 countries, volunteering an estimated 4000 h of work. These efforts expanded the MIBiG repository by 22% and enhanced its usability in downstream molecular data analyses in comparative genomic analyses, natural product discovery, and machine learning applications. We provide strategies and actionable lessons for adopting this model, supporting the sustainability of curated bioinformatics resources central to nucleic acid research and related fields.

biocuration↗

Leveraging Pre-Built Catalogs and Object-Level Scheduling to Eliminate I/O Bottlenecks in HPC Environments

Modern High-Performance Computing (HPC) environments face mounting challenges due to the shift from large to small file datasets, along with an increasing number of users and parallelized applications. As HPC systems rely on Parallel File Systems (PFS), such as Lustre for data processing, performance bottlenecks stemming from Object Storage Target (OST) contention have become a significant concern. Existing solutions, such as LADS with its object-level scheduling approach, fall short in large-scale HPC environments due to their inability to effectively address metadata I/O bottlenecks and the growing number of I/O processes. This study highlights the pressing need for a comprehensive solution that tackles both OST contention and metadata I/O challenges in diverse HPC workloads. To address these challenges, we propose SwiftLoad, an object-level I/O scheduling framework that leverages a metadata catalog to enhance the performance and efficiency of parallel HPC utilities. The adoption of the metadata catalog mitigates the metadata I/O bottlenecks that commonly occur in HPC utilities, a challenge that is particularly pronounced in object-level I/O scheduling. SwiftLoad addresses OST contention and the uneven distribution of I/O processes across different OSTs through mathematical modeling and incorporates a Loader Configuration Module to regulate the number of I/O processes. Evaluated with two representative utilities—data deduplication profiling and data augmentation—SwiftLoad achieved performance improvements of up to 5.63x and 11.0x, respectively, on a production supercomputer.

HPC↗

Automated pipeline processing X-ray diffraction data from dynamic compression experiments on the Extreme Conditions Beamline of PETRA III

Presented and discussed here is the implementation of a software solution that provides prompt X-ray diffraction data analysis during fast dynamic compression experiments conducted within the dynamic diamond anvil cell technique. It includes efficient data collection, streaming of data and metadata to a high-performance cluster (HPC), fast azimuthal data integration on the cluster, and tools for controlling the data processing steps and visualizing the data using the DIOPTAS software package. This data processing pipeline is invaluable for a great number of studies. The potential of the pipeline is illustrated with two examples of data collected on ammonia–water mixtures and multiphase mineral assemblies under high pressure. The pipeline is designed to be generic in nature and could be readily adapted to provide rapid feedback for many other X-ray diffraction techniques, e.g. large-volume press studies, in situ stress/strain studies, phase transformation studies, chemical reactions studied with high-resolution diffraction etc.

97 MATHEMATICS AND COMPUTING↗

Observational Data for Next-Generation Climate Model Evaluation: Requirements, Considerations, and Best Practices

Climate model simulations are an important source of information about our planet’s climate system and also enable informed decision-making under different future scenarios. As a new archive of results from the next generation of climate models is anticipated to become available with the Coupled Model Intercomparison Project phase 7 (CMIP7), the need to develop efficient and robust methods to evaluate models is paramount. Observations are an integral part of model evaluation, providing a means to quantify and understand the degree to which climate models can faithfully reproduce Earth system processes. Such analysis is critical for constraining climate projections, identifying areas of focus for model development, and assisting analysts in deciphering the utility of models for specific applications. Observations of Earth system come from a diversity of sources, span different space–time domains, and are produced by different communities, and each dataset features different data structures and formats, metadata standards, and its own unique uncertainties. Uncertainties in an observational dataset may stem from gaps in temporal and spatial coverage, instrumentation errors, or assumptions in retrieval and processing methods. How then does one ensure that observational data are ready for use and utilized in the most appropriate way for robust, rapid, and routine climate model evaluation? The CMIP7 Model Benchmarking Task Team with input from the broader climate modeling, model evaluation, and observational data communities present a vision and considerations for best practices toward the optimal and appropriate use of observational data to support next-generation climate model evaluation.

Climate models↗

Instantiation of the Damara Tern Platform for Advanced Materials and Manufacturing Technologies (AMMT) Program Collaborative Data Management

This work package focused on deploying an instance of the Damara Tern platform to support AMMT collaborative research activities. The objectives were to provide selected AMMT collaborators with access to a shared environment for capturing operations, trackables, and associated metadata, and to implement data entry functionalities that reflect site-specific procedures. Key activities included creating configurable, schema-driven entry forms and validating the data collection process. The report details the deployment process, the platform infrastructure, and the implemented data entry workflows, providing a reference for end users and establishing a foundation for future production-scale deployments.

36 MATERIALS SCIENCE↗