Search NASA⌕ Search

SEARCH · Search NASA

Results for “metabolic database”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

The secondary metabolism collaboratory: a database and web discussion portal for secondary metabolite biosynthetic gene clusters

Secondary metabolites are small molecules produced by all corners of life, often with specialized bioactive functions with clinical and environmental relevance. Secondary metabolite biosynthetic gene clusters (BGCs) can often be identified within DNA sequences by various sequence similarity tools, but determining the exact functions of genes in the pathway and predicting their chemical products can often only be done by careful, manual comparative analysis. To facilitate this, we report the first release of the secondary metabolism collaboratory (SMC), which aims to provide a comprehensive, tool-agnostic repository of BGC sequence data drawn from all publicly available and user-submitted bacterial and archaeal genome and contig sources. On the website, users are provided a searchable catalog of putative BGCs identified from each source, along with visualizations of gene and domain annotations derived from multiple sequence analysis tools. SMC’s data is also available through publicly-accessible application programming interface (API) endpoints to facilitate programmatic access. Users are encouraged to share their findings (and search for others’) through comment posts on BGC and source pages. At the time of writing, SMC is the largest repository of BGC information, holding 13.1M BGC regions from 1.3M source sequences and growing, and can be found at https://smc.jgi.doe.gov.

59 BASIC BIOLOGICAL SCIENCES↗

Plant Metabolic Network 16: expansion of underrepresented plant groups and experimentally supported enzyme data

Abstract The Plant Metabolic Network (PMN) is a free online database of plant metabolism available at https://plantcyc.org. The latest release, PMN 16, provides metabolic databases representing >1200 metabolic pathways, 1.3 million enzymes, >8000 metabolites, >10 000 reactions and >15 000 citations for 155 plant and green algal genomes, as well as a pan-plant reference database called PlantCyc. This release contains 29 additional genomes compared with PMN 15, including species listed by the African Orphan Crop Consortium and nonflowering plant species. Furthermore, 52 new enzymes with experimentally supported function information have been included in this release. The single-species databases contain a combination of experimental information from the literature and computationally predicted information obtained through PMN’s database generation pipeline for a single species, while PlantCyc contains only experimental information but for any species within Viridiplantae. PMN is a comprehensive resource for querying, visualizing, analyzing and interpreting omics data with metabolic knowledge. It also serves as a useful and interactive tool for teaching plant metabolism.

Hawkins, Charles (ORCID:0000000312849047)↗

A comparative bear model for immobility-induced osteopenia

The National Institutes of Health (NIH) and the National Aeronautics and Space Administration (NASA) are seeking solutions to the human problem of osteopenia, or immobility-induced bone loss. Bears, during winter dormancy, appear uniquely exempted from the debilitating effects of immobility osteopenia. NIH and ESA, Inc. are creating a large database of metabolic information on human ambulatory and bedrest plasma samples for comparison with metabolic data obtained from bear plasma samples collected in different seasons. The database generated from NASA's HR113 human bedrest study showed a clear difference between plasma samples of ambulatory and immobile subjects through cluster analysis using compounds determined by high performance liquid chromatography with coulometric electrochemical array detection (HPLC-EC). We collected plasma samples from black bears (Ursus americanus) across 4 seasons and from 3 areas and subjected them to similar analysis, with particular attention to compounds that changed significantly in the NASA human study. We found seasonal differences in 28 known compounds and 33 unknown compounds. A final database contained 40 known and 120 unknown peaks that were reliably assayed in all bear and human samples; these were the primary data set for interspecies comparison. Six unidentified compounds changed significantly but differentially in wintering bears and immobile humans. The data are discussed in light of current theories regarding dormancy, starvation, and anabolic metabolism. Work is in progress by ESA Laboratories on a larger database to confirm these findings prior to a chemical isolation and identification effort. This research could lead to new pharmaceuticals or dietary interventions for the treatment of immobility osteopenia.

NASA Discipline Musculoskeletal↗

Automating methods for estimating metabolite volatility

The volatility of metabolites can influence their biological roles and inform optimal methods for their detection. Yet, volatility information is not readily available for the large number of described metabolites, limiting the exploration of volatility as a fundamental trait of metabolites. Here, we adapted methods to estimate vapor pressure from the functional group composition of individual molecules (SIMPOL.1) to predict the gas-phase partitioning of compounds in different environments. We implemented these methods in a new open pipeline called volcalc that uses chemoinformatic tools to automate these volatility estimates for all metabolites in an extensive and continuously updated pathway database: the Kyoto Encyclopedia of Genes and Genomes (KEGG) that connects metabolites, organisms, and reactions. We first benchmark the automated pipeline against a manually curated data set and show that the same category of volatility (e.g., nonvolatile, low, moderate, high) is predicted for 93% of compounds. We then demonstrate how volcalc might be used to generate and test hypotheses about the role of volatility in biological systems and organisms. Specifically, we estimate that 3.4 and 26.6% of compounds in KEGG have high volatility depending on the environment (soil vs. clean atmosphere, respectively) and that a core set of volatiles is shared among all domains of life (30%) with the largest proportion of kingdom-specific volatiles identified in bacteria. With volcalc , we lay a foundation for uncovering the role of the volatilome using an approach that is easily integrated with other bioinformatic pipelines and can be continually refined to consider additional dimensions to volatility. The volcalc package is an accessible tool to help design and test hypotheses on volatile metabolites and their unique roles in biological systems.

59 BASIC BIOLOGICAL SCIENCES↗

FatPlants: a comprehensive information system for lipid-related genes and metabolic pathways in plants

Abstract FatPlants, an open-access, web-based database, consolidates data, annotations, analysis results, and visualizations of lipid-related genes, proteins, and metabolic pathways in plants. Serving as a minable resource, FatPlants offers a user-friendly interface for facilitating studies into the regulation of plant lipid metabolism and supporting breeding efforts aimed at increasing crop oil content. This web resource, developed using data derived from our own research, curated from public resources, and gleaned from academic literature, comprises information on known fatty-acid-related proteins, genes, and pathways in multiple plants, with an emphasis on Glycine max, Arabidopsis thaliana, and Camelina sativa. Furthermore, the platform includes machine-learning based methods and navigation tools designed to aid in characterizing metabolic pathways and protein interactions. Comprehensive gene and protein information cards, a Basic Local Alignment Search Tool search function, similar structure search capacities from AphaFold, and ChatGPT-based query for protein information are additional features. Database URL: https://www.fatplants.net/

59 BASIC BIOLOGICAL SCIENCES↗

GeneLab

GeneLab collects and enables analysis of spaceflight and ground-based spaceflight simulation genomic data, RNA and protein expression, and metabolic profiles. It interfaces with other existing databases containing spaceflight omic data. The 2011 National Research Council (NRC) Decadal Survey on NASA Life and Physical Sciences called for increased opportunities for multi-investigator spaceflight opportunities and greater use of genomic approaches to meet the needs of NASA researchers. To address these recommendations of the NRC Decadal Survey, the Space Life and Physical Sciences Research and Applications Division of NASA's Human Exploration and Operations Mission Directorate has initiated a transition to an Open Science architecture to increase research opportunities, and has developed the GeneLab Platform based on highly leveraged and integrated bioinformatics analytics. GeneLab is an interactive, open-access resource where scientists can upload, download, store, search, share, transfer, and analyze omics data from spaceflight and corresponding analogue experiments. Users can explore GeneLab datasets in the Data Repository, analyze data using the Analysis Platform, visualize high-order data and create collaborative projects using the Collaborative Workspace. Our primary goal is to maximize the utilization of the valuable biological research conducted aboard the International Space Station (ISS) by collecting genomic, transcriptomic, proteomic, and metabolomics data known as “omics”. By providing a portal linking processed data to flight parameters, GeneLab enables exploration of the molecular network responses of terrestrial biology to the space environment. This allows researchers to understand the complex responses of biological systems to the space environment. This technology development activity was transferred from the Human Exploration and Operations Mission Directorate to the Science Mission Directorate Division of Biological and Physical Sciences (BPS) in October 2020.

GeneLab↗

CyanoCyc cyanobacterial web portal

CyanoCyc is a web portal that integrates an exceptionally rich database collection of information about cyanobacterial genomes with an extensive suite of bioinformatics tools. It was developed to address the needs of the cyanobacterial research and biotechnology communities. The 277 annotated cyanobacterial genomes currently in CyanoCyc are supplemented with computational inferences including predicted metabolic pathways, operons, protein complexes, and orthologs; and with data imported from external databases, such as protein features and Gene Ontology (GO) terms imported from UniProt. Five of the genome databases have undergone manual curation with input from more than a dozen cyanobacteria experts to correct errors and integrate information from more than 1,765 published articles. CyanoCyc has bioinformatics tools that encompass genome, metabolic pathway and regulatory informatics; omics data analysis; and comparative analyses, including visualizations of multiple genomes aligned at orthologous genes, and comparisons of metabolic networks for multiple organisms. CyanoCyc is a high-quality, reliable knowledgebase that accelerates scientists’ work by enabling users to quickly find accurate information using its powerful set of search tools, to understand gene function through expert mini-reviews with citations, to acquire information quickly using its interactive visualization tools, and to inform better decision-making for fundamental and applied research.

59 BASIC BIOLOGICAL SCIENCES↗

pnnl-predictive-phenomics/csc031cyc

Organism-specific Pathway/Genome databases enable the analysis, visualization and interrogation of metabolism, regulation, and genetics. Licensed under the CC-BY-4.0 license

Zucker, Jeremy [Pacific Northwest National Laborat↗

pnnl-predictive-phenomics/csc043cyc

Organism-specific Pathway/Genome databases enable the analysis, visualization and interrogation of metabolism, regulation, and genetics. Licensed under the CC-BY-4.0 license

Zucker, Jeremy [Pacific Northwest National Laborat↗

DancePartner: Python Package to Mine Multiomics Relationship Networks from Literature and Databases

A goal of multi-omics experiments is to understand how mechanistic molecular biology is altered between conditions, typically a control group and experimental groups. Oftentimes this involves studying changes in biomolecule relationships (e.g. interactions, metabolic relationships) of several types of biomolecules (e.g. proteins, lipids, metabolites). Though several databases contain relationships between biomolecules, understudied species may have little to no relationship information in databases and thus must be mined from literature. There are several challenges to literature mining, including automated full-text extraction, duplicate biomolecule term collapsing, and implementing complex machine learning tools. To make relationship extraction more accessible to the community, a python package called DancePartner was developed to allow for the extraction of relationships from literature and databases, with functions to map biomolecule synonyms to standardized identifiers and visualize and characterize the resulting multi-omics network. Here, in this study, an example dataset involving Caenorhabditis elegans is presented, where relationships are mined from 1443 publications using DancePartner. These relationships are combined with relationships from KEGG, WikiPathways, UniProt, and LipidMaps, and visualized.

BERT↗

BIGEL analysis of gene expression in HL60 cells exposed to X rays or 60 Hz magnetic fields

We screened a panel of 1,920 randomly selected cDNAs to discover genes that are differentially expressed in HL60 cells exposed to 60 Hz magnetic fields (2 mT) or X rays (5 Gy) compared to unexposed cells. Identification of these clones was accomplished using our two-gel cDNA library screening method (BIGEL). Eighteen cDNAs differentially expressed in X-irradiated compared to control HL60 cells were recovered from a panel of 1,920 clones. Differential expression in experimental compared to control cells was confirmed independently by Northern blotting of paired total RNA samples hybridized to each of the 18 clone-specific cDNA probes. DNA sequencing revealed that 15 of the 18 cDNA clones produced matches with the database for genes related to cell growth, protein synthesis, energy metabolism, oxidative stress or apoptosis (including MYC, neuroleukin, copper zinc-dependent superoxide dismutase, TC4 RAS-like protein, peptide elongation factor 1alpha, BNIP3, GATA3, NF45, cytochrome c oxidase II and triosephosphate isomerase mRNAs). In contrast, BIGEL analysis of the same 1,920 cDNAs revealed no differences greater than 1.5-fold in expression levels in magnetic-field compared to sham-exposed cells. Magnetic-field-exposed and control samples were analyzed further for the presence of mRNA encoding X-ray-responsive genes by hybridization of the 18 specific cDNA probes to RNA from exposed and control HL60 cells. Our results suggest that differential gene expression is induced in approximately 1% of a random pool of cDNAs by ionizing radiation but not by 60 Hz magnetic fields under the present experimental conditions.

Non-NASA Center↗

ThermoBase: A Database of the Phylogeny and Physiology of Thermophilic and Hyperthermophilic Organisms

Thermophiles and hyperthermophiles are those organisms which grow at high temperature (> 40°C). The unusual properties of these organisms have received interest in multiple fields of biological research, and have found applications in biotechnology, especially in industrial processes. However, there are few listings of thermophilic and hyperthermophilic organisms and their relevant environmental and physiological data. Such repositories can be used to standardize definitions of thermophile and hyperthermophile limits and tolerances and would mitigate the need for extracting organism data from diverse literature sources across multiple, sometimes loosely related, research fields. Therefore, we have developed ThermoBase, a web-based and freely available database which currently houses comprehensive descriptions for 1238 thermophilic or hyperthermophilic organisms. ThermoBase reports taxonomic, metabolic, environmental, experimental, and physiological information in addition to literature resources. This includes parameters such as coupling ions for chemiosmosis, optimal pH and range, optimal temperature and range, optimal pressure, and optimal salinity. The database interface allows for search features and sorting of parameters. As such, it is the goal of ThermoBase to facilitate and expedite hypothesis generation, literature research, and understanding relating to thermophiles and hyperthermophiles within the scientific community in an accessible and centralized repository. ThermoBase is freely available online at the Astrobiology Habitable Environments Database (AHED; https://ahed.nasa.gov), at the Database Center for Life Science (TogoDB; http://togodb.org/db/thermobase), and in the S1 File.

Thermophiles↗

An in silico assessment of gene function and organization of the phenylpropanoid pathway metabolic networks in Arabidopsis thaliana and limitations thereof

The Arabidopsis genome sequencing in 2000 gave to science the first blueprint of a vascular plant. Its successful completion also prompted the US National Science Foundation to launch the Arabidopsis 2010 initiative, the goal of which is to identify the function of each gene by 2010. In this study, an exhaustive analysis of The Institute for Genomic Research (TIGR) and The Arabidopsis Information Resource (TAIR) databases, together with all currently compiled EST sequence data, was carried out in order to determine to what extent the various metabolic networks from phenylalanine ammonia lyase (PAL) to the monolignols were organized and/or could be predicted. In these databases, there are some 65 genes which have been annotated as encoding putative enzymatic steps in monolignol biosynthesis, although many of them have only very low homology to monolignol pathway genes of known function in other plant systems. Our detailed analysis revealed that presently only 13 genes (two PALs, a cinnamate-4-hydroxylase, a p-coumarate-3-hydroxylase, a ferulate-5-hydroxylase, three 4-coumarate-CoA ligases, a cinnamic acid O-methyl transferase, two cinnamoyl-CoA reductases) and two cinnamyl alcohol dehydrogenases can be classified as having a bona fide (definitive) function; the remaining 52 genes currently have undetermined physiological roles. The EST database entries for this particular set of genes also provided little new insight into how the monolignol pathway was organized in the different tissues and organs, this being perhaps a consequence of both limitations in how tissue samples were collected and in the incomplete nature of the EST collections. This analysis thus underscores the fact that even with genomic sequencing, presumed to provide the entire suite of putative genes in the monolignol-forming pathway, a very large effort needs to be conducted to establish actual catalytic roles (including enzyme versatility), as well as the physiological function(s) for each member of the (multi)gene families present and the metabolic networks that are operative. Additionally, one key to identifying physiological functions for many of these (and other) unknown genes, and their corresponding metabolic networks, awaits the development of technologies to comprehensively study molecular processes at the single cell level in particular tissues and organs, in order to establish the actual metabolic context.

NASA Program Fundamental Space Biology↗

Targeted curation of the gut microbial gene content modulating human cardiovascular disease

Despite the promise of the gut microbiome to predict human health, few studies expose the molecular-scale processes underpinning such forecasts. We mined over 200,000 gut-derived genomes from cultivated and uncultivated microbial lineages to inventory the gut microorganisms and their gene content that control trimethylamine-induced cardiovascular disease. We assigned an atherosclerotic profile to the 6,341 microbial genomes that encoded metabolisms associated with heart disease, creating the Methylated Amine Gene Inventory of Catabolism database (MAGICdb). From microbiome gene expression data sets, we demonstrate that MAGICdb enhanced the recovery of disease-relevant genes and identified the most active microorganisms, unveiling future therapeutic targets. From the feces of healthy and diseased subjects, we show that MAGICdb predicted cardiovascular disease status as effectively as traditional lipid blood tests. This functional microbiome catalog is a public, exploitable resource, designed to enable a new era of microbiota-based therapeutics and diagnostics

metatranscriptomics↗

Speeding genomic island discovery through systematic design of reference database composition

Background Genomic islands (GIs) are mobile genetic elements that integrate site-specifically into bacterial chromosomes, bearing genes that affect phenotypes such as pathogenicity and metabolism. GIs typically occur sporadically among related bacterial strains, enabling comparative genomic approaches to GI identification. For a candidate GI in a query genome, the number of reference genomes with a precise deletion of the GI serves as a support value for the GI. Our comparative software for GI identification was slowed by our original use of large reference genome databases (DBs). Here we explore smaller species-focused DBs. Results With increasing DB size, recovery of our reliable prophage GI calls reached a plateau, while recovery of less reliable GI calls (FPs) increased rapidly as DB sizes exceeded ~500 genomes; i.e., overlarge DBs can increase FP rates. Paradoxically, relative to prophages, FPs were both more frequently supported only by genomes outside the species and more frequently supported only by genomes inside the species; this may be due to their generally lower support values. Setting a DB size limit for our SMA ll R anked T ailored (SMART) DB design speeded runtime ~65-fold. Strictly intra-species DBs would tend to lower yields of prophages for small species (with few genomes available); simulations with large species showed that this could be partially overcome by reaching outside the species to closely related taxa, without an FP burden. Employing such taxonomic outreach in DB design generated redundancy in the DB set; as few as 2984 DBs were needed to cover all 47894 prokaryotic species. Conclusions Runtime decreased dramatically with SMART DB design, with only minor losses of prophages. We also describe potential utility in other comparative genomics projects.

59 BASIC BIOLOGICAL SCIENCES↗

DNA parts and gene constructs for plant biodesign

Plant biodesign requires the knowledge of DNA parts (e.g., genes, promoters, terminators), along with their combinations (as gene constructs) linked to engineered traits. DNA parts with validated or predicted functions in plants have been deposited in various online databases. However, these existing databases focus on basic biological functions of individual DNA parts, leaving a gap between basic knowledge and bioengineering applications. To fill this knowledge gap, we have created a user-friendly, open-ended database as a knowledge graph linking DNA parts to gene constructs to traits. This database contains experimentally validated DNA parts and gene constructs documented in peer-reviewed publications. The DNA parts include 1) molecular components with biological functions, such as genes involved in various biological processes (e.g., metabolic and signal transduction pathways) and 2) molecular components with technical functions, such as gene expression, genome engineering and sequence splicing. The gene constructs deposited in this database include both single-gene and multi-gene constructs. This database allows users to submit DNA parts and gene construct compositions linked to engineered traits described in peer-reviewed publications, providing a public digital repository for sharing the biodesign information among the researchers in the fields of plant biotechnology and plant synthetic biology.

plant biodesign synthetic biology gene constructs ↗

Metagenomic clustering links specific metabolic functions to globally relevant ecosystems

ABSTRACT Metagenomic sequencing has advanced our understanding of biogeochemical processes by providing an unprecedented view into the microbial composition of different ecosystems. While the amount of metagenomic data has grown rapidly, simple-to-use methods to analyze and compare across studies have lagged behind. Thus, tools expressing the metabolic traits of a community are needed to broaden the utility of existing data. Gene abundance profiles are a relatively low-dimensional embedding of a metagenome’s functional potential and are, thus, tractable for comparison across many samples. Here, we compare the abundance of KEGG Ortholog Groups (KOs) from 6,539 metagenomes from the Joint Genome Institute’s Integrated Microbial Genomes and Metagenomes (JGI IMG/M) database. We find that samples cluster into terrestrial, aquatic, and anaerobic ecosystems with marker KOs reflecting adaptations to these environments. For instance, functional clusters were differentiated by the metabolism of antibiotics, photosynthesis, methanogenesis, and surprisingly GC content. Using this functional gene approach, we reveal the broad-scale patterns shaping microbial communities and demonstrate the utility of ortholog abundance profiles for representing a rapidly expanding body of metagenomic data. IMPORTANCE Metagenomics, or the sequencing of DNA from complex microbiomes, provides a view into the microbial composition of different environments. Metagenome databases were created to compile sequencing data across studies, but it remains challenging to compare and gain insight from these large data sets. Consequently, there is a need to develop accessible approaches to extract knowledge across metagenomes. The abundance of different orthologs (i.e., genes that perform a similar function across species) provides a simplified representation of a metagenome’s metabolic potential that can easily be compared with others. In this study, we cluster the ortholog abundance profiles of thousands of metagenomes from diverse environments and uncover the traits that distinguish them. This work provides a simple to use framework for functional comparison and advances our understanding of how the environment shapes microbial communities.

54 ENVIRONMENTAL SCIENCES↗

From Field to Laboratory: A New Database Approach for Linking Microbial Field Ecology with Laboratory Studies

The Ames Exobiology Culture Collection Database (AECC-DB) has been developed as a collaboration between microbial ecologists and information technology specialists. It allows for extensive web-based archiving of information regarding field samples to document microbial co-habitation of specific ecosystem micro-environments. Documentation and archiving continues as pure cultures are isolated, metabolic properties determined, and DNA extracted and sequenced. In this way metabolic properties and molecular sequences are clearly linked back to specific isolates and the location of those microbes in the ecosystem of origin. Use of this database system presents a significant advancement over traditional bookkeeping wherein there is generally little or no information regarding the environments from which microorganisms were isolated. Generally there is only a general ecosystem designation (i.e., hot-spring). However within each of these there are a myriad of microenvironments with very different properties and determining exactly where (which microenvironment) a given microbe comes from is critical in designing appropriate isolation media and interpreting physiological properties. We are currently using the database to aid in the isolation of a large number of cyanobacterial species and will present results by PI's and students demonstrating the utility of this new approach.

Bebout, Leslie↗