Search NASA⌕ Search

SEARCH · Search NASA

Results for “metabolic database”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

The secondary metabolism collaboratory: a database and web discussion portal for secondary metabolite biosynthetic gene clusters

Secondary metabolites are small molecules produced by all corners of life, often with specialized bioactive functions with clinical and environmental relevance. Secondary metabolite biosynthetic gene clusters (BGCs) can often be identified within DNA sequences by various sequence similarity tools, but determining the exact functions of genes in the pathway and predicting their chemical products can often only be done by careful, manual comparative analysis. To facilitate this, we report the first release of the secondary metabolism collaboratory (SMC), which aims to provide a comprehensive, tool-agnostic repository of BGC sequence data drawn from all publicly available and user-submitted bacterial and archaeal genome and contig sources. On the website, users are provided a searchable catalog of putative BGCs identified from each source, along with visualizations of gene and domain annotations derived from multiple sequence analysis tools. SMC’s data is also available through publicly-accessible application programming interface (API) endpoints to facilitate programmatic access. Users are encouraged to share their findings (and search for others’) through comment posts on BGC and source pages. At the time of writing, SMC is the largest repository of BGC information, holding 13.1M BGC regions from 1.3M source sequences and growing, and can be found at https://smc.jgi.doe.gov.

59 BASIC BIOLOGICAL SCIENCES↗

Plant Metabolic Network 16: expansion of underrepresented plant groups and experimentally supported enzyme data

Abstract The Plant Metabolic Network (PMN) is a free online database of plant metabolism available at https://plantcyc.org. The latest release, PMN 16, provides metabolic databases representing >1200 metabolic pathways, 1.3 million enzymes, >8000 metabolites, >10 000 reactions and >15 000 citations for 155 plant and green algal genomes, as well as a pan-plant reference database called PlantCyc. This release contains 29 additional genomes compared with PMN 15, including species listed by the African Orphan Crop Consortium and nonflowering plant species. Furthermore, 52 new enzymes with experimentally supported function information have been included in this release. The single-species databases contain a combination of experimental information from the literature and computationally predicted information obtained through PMN’s database generation pipeline for a single species, while PlantCyc contains only experimental information but for any species within Viridiplantae. PMN is a comprehensive resource for querying, visualizing, analyzing and interpreting omics data with metabolic knowledge. It also serves as a useful and interactive tool for teaching plant metabolism.

Hawkins, Charles (ORCID:0000000312849047)↗

Automating methods for estimating metabolite volatility

The volatility of metabolites can influence their biological roles and inform optimal methods for their detection. Yet, volatility information is not readily available for the large number of described metabolites, limiting the exploration of volatility as a fundamental trait of metabolites. Here, we adapted methods to estimate vapor pressure from the functional group composition of individual molecules (SIMPOL.1) to predict the gas-phase partitioning of compounds in different environments. We implemented these methods in a new open pipeline called volcalc that uses chemoinformatic tools to automate these volatility estimates for all metabolites in an extensive and continuously updated pathway database: the Kyoto Encyclopedia of Genes and Genomes (KEGG) that connects metabolites, organisms, and reactions. We first benchmark the automated pipeline against a manually curated data set and show that the same category of volatility (e.g., nonvolatile, low, moderate, high) is predicted for 93% of compounds. We then demonstrate how volcalc might be used to generate and test hypotheses about the role of volatility in biological systems and organisms. Specifically, we estimate that 3.4 and 26.6% of compounds in KEGG have high volatility depending on the environment (soil vs. clean atmosphere, respectively) and that a core set of volatiles is shared among all domains of life (30%) with the largest proportion of kingdom-specific volatiles identified in bacteria. With volcalc , we lay a foundation for uncovering the role of the volatilome using an approach that is easily integrated with other bioinformatic pipelines and can be continually refined to consider additional dimensions to volatility. The volcalc package is an accessible tool to help design and test hypotheses on volatile metabolites and their unique roles in biological systems.

59 BASIC BIOLOGICAL SCIENCES↗

FatPlants: a comprehensive information system for lipid-related genes and metabolic pathways in plants

Abstract FatPlants, an open-access, web-based database, consolidates data, annotations, analysis results, and visualizations of lipid-related genes, proteins, and metabolic pathways in plants. Serving as a minable resource, FatPlants offers a user-friendly interface for facilitating studies into the regulation of plant lipid metabolism and supporting breeding efforts aimed at increasing crop oil content. This web resource, developed using data derived from our own research, curated from public resources, and gleaned from academic literature, comprises information on known fatty-acid-related proteins, genes, and pathways in multiple plants, with an emphasis on Glycine max, Arabidopsis thaliana, and Camelina sativa. Furthermore, the platform includes machine-learning based methods and navigation tools designed to aid in characterizing metabolic pathways and protein interactions. Comprehensive gene and protein information cards, a Basic Local Alignment Search Tool search function, similar structure search capacities from AphaFold, and ChatGPT-based query for protein information are additional features. Database URL: https://www.fatplants.net/

59 BASIC BIOLOGICAL SCIENCES↗

CyanoCyc cyanobacterial web portal

CyanoCyc is a web portal that integrates an exceptionally rich database collection of information about cyanobacterial genomes with an extensive suite of bioinformatics tools. It was developed to address the needs of the cyanobacterial research and biotechnology communities. The 277 annotated cyanobacterial genomes currently in CyanoCyc are supplemented with computational inferences including predicted metabolic pathways, operons, protein complexes, and orthologs; and with data imported from external databases, such as protein features and Gene Ontology (GO) terms imported from UniProt. Five of the genome databases have undergone manual curation with input from more than a dozen cyanobacteria experts to correct errors and integrate information from more than 1,765 published articles. CyanoCyc has bioinformatics tools that encompass genome, metabolic pathway and regulatory informatics; omics data analysis; and comparative analyses, including visualizations of multiple genomes aligned at orthologous genes, and comparisons of metabolic networks for multiple organisms. CyanoCyc is a high-quality, reliable knowledgebase that accelerates scientists’ work by enabling users to quickly find accurate information using its powerful set of search tools, to understand gene function through expert mini-reviews with citations, to acquire information quickly using its interactive visualization tools, and to inform better decision-making for fundamental and applied research.

59 BASIC BIOLOGICAL SCIENCES↗

pnnl-predictive-phenomics/csc031cyc

Organism-specific Pathway/Genome databases enable the analysis, visualization and interrogation of metabolism, regulation, and genetics. Licensed under the CC-BY-4.0 license

Zucker, Jeremy [Pacific Northwest National Laborat↗

pnnl-predictive-phenomics/csc043cyc

Organism-specific Pathway/Genome databases enable the analysis, visualization and interrogation of metabolism, regulation, and genetics. Licensed under the CC-BY-4.0 license

Zucker, Jeremy [Pacific Northwest National Laborat↗

DancePartner: Python Package to Mine Multiomics Relationship Networks from Literature and Databases

A goal of multi-omics experiments is to understand how mechanistic molecular biology is altered between conditions, typically a control group and experimental groups. Oftentimes this involves studying changes in biomolecule relationships (e.g. interactions, metabolic relationships) of several types of biomolecules (e.g. proteins, lipids, metabolites). Though several databases contain relationships between biomolecules, understudied species may have little to no relationship information in databases and thus must be mined from literature. There are several challenges to literature mining, including automated full-text extraction, duplicate biomolecule term collapsing, and implementing complex machine learning tools. To make relationship extraction more accessible to the community, a python package called DancePartner was developed to allow for the extraction of relationships from literature and databases, with functions to map biomolecule synonyms to standardized identifiers and visualize and characterize the resulting multi-omics network. Here, in this study, an example dataset involving Caenorhabditis elegans is presented, where relationships are mined from 1443 publications using DancePartner. These relationships are combined with relationships from KEGG, WikiPathways, UniProt, and LipidMaps, and visualized.

BERT↗

Targeted curation of the gut microbial gene content modulating human cardiovascular disease

Despite the promise of the gut microbiome to predict human health, few studies expose the molecular-scale processes underpinning such forecasts. We mined over 200,000 gut-derived genomes from cultivated and uncultivated microbial lineages to inventory the gut microorganisms and their gene content that control trimethylamine-induced cardiovascular disease. We assigned an atherosclerotic profile to the 6,341 microbial genomes that encoded metabolisms associated with heart disease, creating the Methylated Amine Gene Inventory of Catabolism database (MAGICdb). From microbiome gene expression data sets, we demonstrate that MAGICdb enhanced the recovery of disease-relevant genes and identified the most active microorganisms, unveiling future therapeutic targets. From the feces of healthy and diseased subjects, we show that MAGICdb predicted cardiovascular disease status as effectively as traditional lipid blood tests. This functional microbiome catalog is a public, exploitable resource, designed to enable a new era of microbiota-based therapeutics and diagnostics

metatranscriptomics↗

Speeding genomic island discovery through systematic design of reference database composition

Background Genomic islands (GIs) are mobile genetic elements that integrate site-specifically into bacterial chromosomes, bearing genes that affect phenotypes such as pathogenicity and metabolism. GIs typically occur sporadically among related bacterial strains, enabling comparative genomic approaches to GI identification. For a candidate GI in a query genome, the number of reference genomes with a precise deletion of the GI serves as a support value for the GI. Our comparative software for GI identification was slowed by our original use of large reference genome databases (DBs). Here we explore smaller species-focused DBs. Results With increasing DB size, recovery of our reliable prophage GI calls reached a plateau, while recovery of less reliable GI calls (FPs) increased rapidly as DB sizes exceeded ~500 genomes; i.e., overlarge DBs can increase FP rates. Paradoxically, relative to prophages, FPs were both more frequently supported only by genomes outside the species and more frequently supported only by genomes inside the species; this may be due to their generally lower support values. Setting a DB size limit for our SMA ll R anked T ailored (SMART) DB design speeded runtime ~65-fold. Strictly intra-species DBs would tend to lower yields of prophages for small species (with few genomes available); simulations with large species showed that this could be partially overcome by reaching outside the species to closely related taxa, without an FP burden. Employing such taxonomic outreach in DB design generated redundancy in the DB set; as few as 2984 DBs were needed to cover all 47894 prokaryotic species. Conclusions Runtime decreased dramatically with SMART DB design, with only minor losses of prophages. We also describe potential utility in other comparative genomics projects.

59 BASIC BIOLOGICAL SCIENCES↗

DNA parts and gene constructs for plant biodesign

Plant biodesign requires the knowledge of DNA parts (e.g., genes, promoters, terminators), along with their combinations (as gene constructs) linked to engineered traits. DNA parts with validated or predicted functions in plants have been deposited in various online databases. However, these existing databases focus on basic biological functions of individual DNA parts, leaving a gap between basic knowledge and bioengineering applications. To fill this knowledge gap, we have created a user-friendly, open-ended database as a knowledge graph linking DNA parts to gene constructs to traits. This database contains experimentally validated DNA parts and gene constructs documented in peer-reviewed publications. The DNA parts include 1) molecular components with biological functions, such as genes involved in various biological processes (e.g., metabolic and signal transduction pathways) and 2) molecular components with technical functions, such as gene expression, genome engineering and sequence splicing. The gene constructs deposited in this database include both single-gene and multi-gene constructs. This database allows users to submit DNA parts and gene construct compositions linked to engineered traits described in peer-reviewed publications, providing a public digital repository for sharing the biodesign information among the researchers in the fields of plant biotechnology and plant synthetic biology.

plant biodesign synthetic biology gene constructs ↗

Metagenomic clustering links specific metabolic functions to globally relevant ecosystems

ABSTRACT Metagenomic sequencing has advanced our understanding of biogeochemical processes by providing an unprecedented view into the microbial composition of different ecosystems. While the amount of metagenomic data has grown rapidly, simple-to-use methods to analyze and compare across studies have lagged behind. Thus, tools expressing the metabolic traits of a community are needed to broaden the utility of existing data. Gene abundance profiles are a relatively low-dimensional embedding of a metagenome’s functional potential and are, thus, tractable for comparison across many samples. Here, we compare the abundance of KEGG Ortholog Groups (KOs) from 6,539 metagenomes from the Joint Genome Institute’s Integrated Microbial Genomes and Metagenomes (JGI IMG/M) database. We find that samples cluster into terrestrial, aquatic, and anaerobic ecosystems with marker KOs reflecting adaptations to these environments. For instance, functional clusters were differentiated by the metabolism of antibiotics, photosynthesis, methanogenesis, and surprisingly GC content. Using this functional gene approach, we reveal the broad-scale patterns shaping microbial communities and demonstrate the utility of ortholog abundance profiles for representing a rapidly expanding body of metagenomic data. IMPORTANCE Metagenomics, or the sequencing of DNA from complex microbiomes, provides a view into the microbial composition of different environments. Metagenome databases were created to compile sequencing data across studies, but it remains challenging to compare and gain insight from these large data sets. Consequently, there is a need to develop accessible approaches to extract knowledge across metagenomes. The abundance of different orthologs (i.e., genes that perform a similar function across species) provides a simplified representation of a metagenome’s metabolic potential that can easily be compared with others. In this study, we cluster the ortholog abundance profiles of thousands of metagenomes from diverse environments and uncover the traits that distinguish them. This work provides a simple to use framework for functional comparison and advances our understanding of how the environment shapes microbial communities.

54 ENVIRONMENTAL SCIENCES↗

Metabolic interactions underpinning high methane fluxes across terrestrial freshwater wetlands

Current estimates of wetland contributions to the global methane budget carry high uncertainty, particularly in accurately predicting emissions from high methane-emitting wetlands. Microorganisms drive methane cycling, but little is known about their conservation across wetlands. To address this, we integrate 16S rRNA amplicon datasets, metagenomes, metatranscriptomes, and annual methane flux data across 9 wetlands, creating the Multi-Omics for Understanding Climate Change (MUCC) v2.0.0 database. This resource is used to link microbiome composition to function and methane emissions, focusing on methane-cycling microbes and the networks driving carbon decomposition. We identify eight methane-cycling genera shared across wetlands and show wetland-specific metabolic interactions in marshes, revealing low connections between methanogens and methanotrophs in high-emitting wetlands. Methanoregula emerged as a hub methanogen across networks and is a strong predictor of methane flux. In these wetlands it also displays the functional potential for methylotrophic methanogenesis, highlighting the importance of this pathway in these ecosystems. Collectively, our findings illuminate trends between microbial decomposition networks and methane flux while providing an extensive publicly available database to advance future wetland research.

54 ENVIRONMENTAL SCIENCES↗

Evolutionary flexibility and rigidity in the bacterial methylerythritol phosphate (MEP) pathway

Terpenoids are a diverse class of compounds with wide-ranging uses including as industrial solvents, pharmaceuticals, and fragrances. Efforts to produce terpenoids sustainably by engineering microbes for fermentation are ongoing, but industrial production still largely relies on nonrenewable sources. The methylerythritol phosphate (MEP) pathway generates terpenoid precursor molecules and includes the enzyme Dxs and two iron–sulfur cluster enzymes: IspG and IspH. IspG and IspH are rate limiting-enzymes of the MEP pathway but are challenging for metabolic engineering because they require iron–sulfur cluster biogenesis and an ongoing supply of reducing equivalents to function. Therefore, identifying novel alternatives to IspG and IspH has been an on-going effort to aid in metabolic engineering of terpenoid biosynthesis. We report here an analysis of the evolutionary diversity of terpenoid biosynthesis strategies as a resource for exploration of alternative terpenoid biosynthesis pathways. Using comparative genomics, we surveyed a database of 4,400 diverse bacterial species and found that some may have evolved alternatives to the first enzyme in the pathway, Dxs making it evolutionarily flexible. In contrast, we found that IspG and IspH are evolutionarily rigid because we could not identify any species that appear to have enzymatic routes that circumvent these enzymes. The ever-growing repository of sequenced bacterial genomes has great potential to provide metabolic engineers with alternative metabolic pathway solutions. With the current state of knowledge, we found that enzymes IspG and IspH are evolutionarily indispensable which informs both metabolic engineering efforts and our understanding of the evolution of terpenoid biosynthesis pathways.

59 BASIC BIOLOGICAL SCIENCES↗

Transcriptomic Analysis of Arachidonic Acid Pathway Genes Provides Mechanistic Insight into Multi-Organ Inflammatory and Vascular Diseases

Arachidonic acid (AA) metabolites have been associated with several diseases across various organ systems, including the cardiovascular, pulmonary, and renal systems. Lipid mediators generated from AA oxidation have been studied to control macrophages, T-cells, cytokines, and fibroblasts, and regulate inflammatory mediators that induce vascular remodeling and dysfunction. AA is metabolized by cyclooxygenase (COX), lipoxygenase (LOX), and cytochrome P450 (CYP) to generate anti-inflammatory, pro-inflammatory, and pro-resolutory oxidized lipids. As comorbid states such as diabetes, hypertension, and obesity become more prevalent in cardiovascular disease, studying the expression of AA pathway genes and their association with these diseases can provide unique pathophysiological insights. In addition, the AA pathway of oxidized lipids exhibits diverse functions across different organ systems, where a lipid can be both anti-inflammatory and pro-inflammatory depending on the location of metabolic activity. Therefore, we aimed to characterize the gene expression of these lipid enzymes and receptors throughout multi-organ diseases via a transcriptomic meta-analysis using the Gene Expression Omnibus (GEO) Database. In our study, we found that distinct AA pathways were expressed in various comorbid conditions, especially those with prominent inflammatory risk factors. Comorbidities, such as hypertension, diabetes, and obesity appeared to contribute to elevated expression of pro-inflammatory lipid mediator genes. Our results demonstrate that expression of inflammatory AA pathway genes may potentiate and attenuate disease; therefore, we suggest further exploration of these pathways as therapeutic targets to improve outcomes.

59 BASIC BIOLOGICAL SCIENCES↗

MjCyc: Rediscovering the pathway-genome landscape of the first sequenced archaeon, Methanocaldococcus (Methanococcus) jannaschii

The genome of Methanocaldococcus (Methanococcus) jannaschii DSM 2661 was the first Archaeal genome to be sequenced in 1996. Subsequent sequence-based annotation cycles led to its first metabolic reconstruction in 2005. Leveraging new experimental results and function assignments, we have now re-annotated M. jannaschii, creating an updated resource with novel information and testable predictions in a pathway-genome database available at BioCyc.org. This reannotation effort has resulted in 652 function assignments with enzyme roles, accounting for a third of the total protein-coding entries for this genome. The updated resource includes 883 reactions, 540 enzymes, and 142 individual pathways. Despite notable progress in computational genomics, more than a third of the genome remains functionally uncharacterized. The publicly available MjCyc pathway-genome database holds great potential for the wider community to conduct research on the biology of methanogenic Archaea.

59 BASIC BIOLOGICAL SCIENCES↗

Meta-virus resource (MetaVR): expanding the frontiers of viral diversity with 24 million uncultivated virus genomes

Viruses are ubiquitous in all environments and impact host metabolism, evolution, and ecology, although our knowledge of their biodiversity is still extremely limited. Viral diversity from genomic and metagenomic datasets has led to an explosion of uncultivated virus genomes (UViGs) and the development of specialized databases to catalog this viral diversity, though many lack comprehensive integration. Here, we introduce meta-virus resource (MetaVR), the successor of the IMG/VR database, designed to overcome previous limitations such as large-scale querying and programmatic access. Drawing on the increase of publicly available genomes and metagenomes, MetaVR significantly expands viral diversity, now comprising 24,435,662 UViGs, a 57.6% increase from its predecessor, organized into over 12 million viral operational taxonomic units. Key enhancements include the integration of curated eukaryotic host information, the integration of protein clusters and predicted structures for comparative studies, and an API for programmatic data access. Furthermore, MetaVR features an updated taxonomic framework based on ICTV release 39, assignment to Baltimore classes, and enhanced host assignment through novel computational tools like iPHoP. These advancements position MetaVR as a unique resource for exploring viral diversity, evolution, and host interactions across diverse environments. MetaVR can be freely accessed at https://www.meta-virome.org/.

Fiamenghi, Mateus B↗

Untargeted GC-MS Metabolic Profiling of Anaerobic Gut Fungi Reveals Putative Terpenoids and Strain-Specific Metabolites

Background/Objectives: Anaerobic gut fungi (Neocallimastigomycota) are biotechnologically relevant, lignocellulose-degrading microbes with under-explored biosynthetic potential for secondary metabolites. Untargeted metabolomic profiling with gas chromatography–mass spectrometry (GC-MS) was applied to two gut fungal strains, Anaeromyces robustus and Caecomyces churrovis, to establish a foundational metabolomic dataset to identify metabolites and provide insights into gut fungal metabolic capabilities. Methods: Gut fungi were cultured anaerobically in rumen-fluid-based media with a soluble substrate (cellobiose), and metabolites were extracted using the Metabolite, Protein, and Lipid Extraction (MPLEx) method, enabling metabolomic and proteomic analysis from the same cell samples. Samples were derivatized and analyzed via GC-MS, followed by compound identification by spectral matching to reference databases, molecular networking, and statistical analyses. Results: Distinct metabolites were identified between A. robustus and C. churrovis, including 2,3-dihydroxyisovaleric acid produced by A. robustus and maltotriitol, maltotriose, and melibiose produced by C. churrovis. C. churrovis may polymerize maltotriose to form an extracellular polysaccharide, like pullulan. GC-MS profiling potentially captured sufficiently volatile products of proteomically detected, putative non-ribosomal peptide synthetases and polyketide synthases of A. robustus and C. churrovis. The triterpene squalene and triterpenoid tetrahymanol were putatively identified in A. robustus and C. churrovis. Their conserved, predicted biosynthetic genes—squalene synthase and squalene tetrahymanol cyclase—were identified in A. robustus, C. churrovis, and other anaerobic gut fungal genera. Conclusions: This study provides a foundational, untargeted metabolomic dataset to unmask gut fungal metabolic pathways and biosynthetic potential and to prioritize future efforts for compound isolation and identification.

Biochemistry & Molecular Biology↗