Search NASA⌕ Search

SEARCH · Search NASA

Results for “Bioinformatics Analysis Pipeline”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Exabiome: Advancing Microbial Science through Exascale Computing

The Exabiome project seeks to improve the understanding of microbiomes through the development of methods for accelerating metagenomic science using exascale computing. This article gives an overview of scientific impact of the three components of the project: metagenome assembly, protein family detection, and comparative analysis of metagenomes. Exabiome developed MetaHipMer, the only metagenome assembler capable of scaling to full exascale systems. MetaHipMer has enabled ground-breaking assemblies on the Frontier supercomputer, with many scientific benefits, such as the discovery of rare species and viral genomes. To investigate protein families, Exabiome developed two exascale tools, PASTIS and HipMCL. Together, these can utilize exascale resources to understand the functional diversity of billions of dark matter proteins and novel protein families. For comparative analysis, Exabiome developed kmerprof, a tool that can be used to compare huge metagenomes for many different scientific purposes, for example, grouping human microbiomes according to body location.

59 BASIC BIOLOGICAL SCIENCES↗

Demographic drivers of gut microbiome diversity

Abstract The gut microbiome plays a central role in orchestrating metabolic, immune, and neurological functions essential for human health. While extensive research has explored the effects of diseases and pathological conditions on gut microbiome composition, the influence of demographic factors remains underexplored, limiting our understanding of microbiome variations in disease states. This study addresses this gap by investigating the impact of demographic variables, including age, sex, and geography, on gut microbiome diversity in healthy individuals. Using the American Gut Project’s extensive dataset and the QIIME2 bioinformatics pipeline, we conducted a comprehensive analysis of microbial profiles across diverse demographic groups. Our results revealed significant age-related shifts in microbial richness and composition, and geographic location strongly influenced phylogenetic diversity. In contrast, sex exhibited limited impact on microbial diversity within healthy BMI ranges. These findings highlight the critical role of demographic factors in shaping gut microbiome diversity, providing a foundational framework to better contextualize disease-related microbiome variations and advance personalized healthcare approaches.

Biotechnology & Applied Microbiology↗

Results from a multi-laboratory ocean metaproteomic intercomparison: effects of LC-MS acquisition and data analysis procedures

Metaproteomics is an increasingly popular methodology that provides information regarding the metabolic functions of specific microbial taxa and has potential for contributing to ocean ecology and biogeochemical studies. A blinded multi-laboratory intercomparison was conducted to assess comparability and reproducibility of taxonomic and functional results and their sensitivity to methodological variables. Euphotic zone samples from the Bermuda Atlantic Time-series Study (BATS) in the North Atlantic Ocean collected by in situ pumps and the autonomous underwater vehicle (AUV) Clio were distributed with a paired metagenome, and one-dimensional (1D) liquid chromatographic data-dependent acquisition mass spectrometry analysis was stipulated. Analysis of mass spectra from seven laboratories through a common bioinformatic pipeline identified a shared set of 1056 proteins from 1395 shared peptide constituents. Quantitative analyses showed good reproducibility: pairwise regressions of spectral counts between laboratories yielded R 2 values averaged 0.62±0.11, and a Sørensen similarity analysis of the top 1000 proteins revealed 70 %–80 % similarity between laboratory groups. Taxonomic and functional assignments showed good coherence between technical replicates and different laboratories. A bioinformatic intercomparison study, involving 10 laboratories using eight software packages, successfully identified thousands of peptides within the complex metaproteomic datasets, demonstrating the utility of these software tools for ocean metaproteomic research. Lessons learned and potential improvements in methods were described. Future efforts could examine reproducibility in deeper metaproteomes, examine accuracy in targeted absolute quantitation analyses, and develop standards for data output formats to improve data interoperability. Together, these results demonstrate the reproducibility of metaproteomic analyses and their suitability for microbial oceanography research, including integration into global-scale ocean surveys and ocean biogeochemical models.

59 BASIC BIOLOGICAL SCIENCES↗

An FPGA-based hardware accelerator supporting sensitive sequence homology filtering with profile hidden Markov models

Abstract Background Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). Here, we describe an FPGA hardware accelerator, called HAVAC, that targets a key bottleneck step (SSV) in the analysis pipeline of the popular pHMM alignment tool, HMMER. Results The HAVAC kernel calculates the SSV matrix at 1739 GCUPS on a $$\sim$$ ∼ $3000 Xilinx Alveo U50 FPGA accelerator card, $$\sim$$ ∼ 227× faster than the optimized SSV implementation in nhmmer . Accounting for PCI-e data transfer data processing, HAVAC is 65× faster than nhmmer’s SSV with one thread and 35× faster than nhmmer with four threads, and uses $$\sim$$ ∼ 31% the energy of a traditional high end Intel CPU. Conclusions HAVAC demonstrates the potential offered by FPGA hardware accelerators to produce dramatic speed gains in sequence annotation and related bioinformatics applications. Because these computations are performed on a co-processor, the host CPU remains free to simultaneously compute other aspects of the analysis pipeline.

59 BASIC BIOLOGICAL SCIENCES↗

MVP: a modular viromics pipeline to identify, filter, cluster, annotate, and bin viruses from metagenomes

While numerous computational frameworks and workflows are available for recovering prokaryote and eukaryote genomes from metagenome data, only a limited number of pipelines are designed specifically for viromics analysis. With many viromics tools developed in the last few years alone, it can be challenging for scientists with limited bioinformatics experience to easily recover, evaluate quality, annotate genes, dereplicate, assign taxonomy, and calculate relative abundance and coverage of viral genomes using state-of-the-art methods and standards. Here, we describe Modular Viromics Pipeline (MVP) v.1.0, a user-friendly pipeline written in Python and providing a simple framework to perform standard viromics analyses. MVP combines multiple tools to enable viral genome identification, characterization of genome quality, filtering, clustering, taxonomic and functional annotation, genome binning, and comprehensive summaries of results that can be used for downstream ecological analyses. Overall, MVP provides a standardized and reproducible pipeline for both extensive and robust characterization of viruses from large-scale sequencing data including metagenomes, metatranscriptomes, viromes, and isolate genomes. As a typical use case, we show how the entire MVP pipeline can be applied to a set of 20 metagenomes from wetland sediments using only 10 modules executed via command lines, leading to the identification of 11,656 viral contigs and 8,145 viral operational taxonomic units (vOTUs) displaying a clear beta-diversity pattern. Further, acting as a dynamic wrapper, MVP is designed to continuously incorporate updates and integrate new tools, ensuring its ongoing relevance in the rapidly evolving field of viromics. MVP is available at https://gitlab.com/ccoclet/mvp and as versioned packages in PyPi and Conda.

59 BASIC BIOLOGICAL SCIENCES↗

Creation of an Acyltransferase Toolbox for Plant Biomass Engineering (Final Report)

The major goal of this project was to expand our understanding of acyl‐CoA ligases and BAHD acyltransferases and their utility in plant engineering. We combined bioinformatic analysis of genes and transcripts with functional fingerprinting of synthesized genes produced by JGI. Best candidates from this experimental pipeline were transferred into bioenergy plants to study their effects on lignin composition. We found combinations of ligase and transferase genes encoding enzymes with interesting catalytic specificities. Our work demonstrated the feasibility of use of acyl-CoA ligases and BAHD acyltransferases to alter the composition of plant cell walls without deleterious effects on the modified plant.

59 BASIC BIOLOGICAL SCIENCES↗

OrthoPhyl—streamlining large-scale, orthology-based phylogenomic studies of bacteria at broad evolutionary scales

Abstract There are a staggering number of publicly available bacterial genome sequences (at writing, 2.0 million assemblies in NCBI's GenBank alone), and the deposition rate continues to increase. This wealth of data begs for phylogenetic analyses to place these sequences within an evolutionary context. A phylogenetic placement not only aids in taxonomic classification but informs the evolution of novel phenotypes, targets of selection, and horizontal gene transfer. Building trees from multi-gene codon alignments is a laborious task that requires bioinformatic expertise, rigorous curation of orthologs, and heavy computation. Compounding the problem is the lack of tools that can streamline these processes for building trees from large-scale genomic data. Here we present OrthoPhyl, which takes bacterial genome assemblies and reconstructs trees from whole genome codon alignments. The analysis pipeline can analyze an arbitrarily large number of input genomes (>1200 tested here) by identifying a diversity-spanning subset of assemblies and using these genomes to build gene models to infer orthologs in the full dataset. To illustrate the versatility of OrthoPhyl, we show three use cases: E. coli/Shigella, Brucella/Ochrobactrum and the order Rickettsiales. We compare trees generated with OrthoPhyl to trees generated with kSNP3 and GToTree along with published trees using alternative methods. We show that OrthoPhyl trees are consistent with other methods while incorporating more data, allowing for greater numbers of input genomes, and more flexibility of analysis.

59 BASIC BIOLOGICAL SCIENCES↗

GenomeDepot v1.0

GenomeDepot is a web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of web-sites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, BLAST search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

GenomeDepot: data management system for microbial comparative genomics

Summary GenomeDepot is an open-source web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of websites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, Basic Local Alignment Search Tool (BLAST) search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools. Availability and implementation GenomeDepot is open source and distributed under the GNU General Public License via GitHub (https://github.com/aekazakov/genome-depot). GenomeDepot is implemented in Python and was tested in Ubuntu Linux. Full installation instructions and documentation are available at https://aekazakov.github.io/genome-depot/. GenomeDepot demo server is freely accessible at https://iseq.lbl.gov/demogd/.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

A large-scale screening campaign of putative carbohydrate-active enzymes reveals a novel xylanase from anaerobic gut fungi

The genomes of anaerobic gut fungi (AGF) encode a diverse array of carbohydrate-active enzymes (CAZymes), yet exceedingly few of these enzymes have been experimentally validated or expressed in heterologous systems. Here, we developed a predictive bioinformatic pipeline to annotate novel putative CAZymes from anaerobic fungi and validate their activity through large-scale heterologous expression in Escherichia coli. A total of 173 fungal proteins from Piromyces finnis associated with biomass degradation were synthesized and expressed in E. coli, and 9.8% were soluble with expression levels exceeding 5% of the total proteome using high-throughput proteomic screening. Among these 17 heterologously expressed proteins, analysis with AlphaFold and FoldSeek predicted 13 multi-functional proteins containing catalytic domains fused with repetitive fungal dockerins, and half of the substrate predictions were experimentally validated. One promising enzyme, celsome_012, exhibited robust and specific activity against beechwood xylan at 37°C and pH 6.4, with titers that were also fivefold higher than those of other recombinant proteins screened here. Both Michaelis-Menten kinetics and the linearized Lineweaver-Burk equation yielded consistent values for K m , and its activation energy was estimated at 51.9 kJ/mol based on the Arrhenius model. This work supports the industrial translation of anaerobic fungal CAZymes due to their robust lignocellulolytic activity and provides a framework for prioritizing AGF proteins for efficient E. coli heterologous expression.

59 BASIC BIOLOGICAL SCIENCES↗

CAHS: Context-Aware Homology Search

Protein homology search is foundational to bioinformatics: it supports annotation transfer, structure/function inference, and evolutionary analysis over rapidly expanding sequence repositories (e.g., UniProtKB). Profile hidden Markov models (pHMMs), as implemented in HMMER, remain the most widely trusted approach because they provide statistically calibrated E-values; however, their gap behavior is fixed once a profile is trained, despite biological evidence that insertion/deletion tolerance varies across flexible loops and intrinsically disordered regions. We present CAHS (Context-Aware Homology Search), a lightweight query-time adapter for pHMM search that incorporates learned and biologically motivated signals without changing HMMER's downstream search pipeline or its calibrated E-value reporting. Given a query sequence, CAHS computes per-residue representations from a protein language model and a disorder predictor, maps these to profile coordinates, and modulates only match-state transition rows (gap-open and gap-extension probabilities) while preserving Plan7 constraints. We comprehensively evaluate CAHS across six structurally diverse protein families and multi-domain architectures against a 570k-sequence target corpus. CAHS expands detection capability, retrieving thousands of additional remote homologs at relaxed thresholds by maintaining alignment quality through flexible regions. For multi-domain proteins, context-aware modulation resolves 94% of fragmented alignments. Crucially, CAHS preserves hit-set invariance at stringent operating points (E<10-10), demonstrating increased statistical confidence without inflating false positives. Furthermore, sharper statistical distinction between homologs and background noise during early filter stages yields up to a 3.87× acceleration in end-to-end wall-clock time on high-performance computing clusters. Overall, CAHS illustrates a practical AI-for-science design pattern: augmenting a trusted probabilistic model with query-specific learned signals to improve interpretable, reproducible inference in data-rich biology.

Bhattaram, Swethasree [Georgia Institute of Techno↗

Hypermut 3: identifying specific mutational patterns in a defined nucleotide context that allows multistate characters

Abstract Motivation The detection of APOBEC3F- and APOBEC3G-induced mutations in virus sequences is useful for identifying hypermutated sequences. These sequences are not representative of viral evolution and can therefore alter the results of downstream sequence analyses if included. We previously published the software Hypermut, which detects hypermutation events in sequences relative to a reference. Two versions of this method are available as a webtool. Neither of these methods consider multistate characters or gaps in the sequence alignment. Results Here, we present an updated, user-friendly web and command-line version of Hypermut with functionality to handle multistate characters and gaps in the sequence alignment. This tool allows for straightforward integration of hypermutation detection into sequence analysis pipelines. As with the previous tool, while the main purpose is to identify G to A hypermutation events, any mutational pattern and context can be specified. Availability and implementation Hypermut 3 is written in Python 3. It is available as a command-line tool at https://github.com/MolEvolEpid/hypermut3 and as a webtool at https://www.hiv.lanl.gov/content/sequence/HYPERMUT/hypermutv3.html.

59 BASIC BIOLOGICAL SCIENCES↗

Cryptic cycling by electroactive bacterioplankton in Trout Bog Lake

The potential for extracellular electron transfer (EET) is a prevailing genomic feature of humic lake bacterioplankton. However, there has been little evidence for the substantial ecological contribution predicted by genetics. We hypothesized that anoxygenic phototrophic electrotrophs and accompanying heterotrophic electrogens cycle dissolved organic matter (DOM) between oxidized and reduced states. We predicted that such bacterioplankton would exhibit diel-scale oscillations due to the light dependency of photosynthesis. Using Trout Bog Lake in Wisconsin, USA, as our model ecosystem, we profiled the water column with depth-discrete metagenomic, physiochemical, and electrochemical analyses. We observed variation in oxidation reduction potential (ORP) in response to sunlight, initiating at depths populated by anoxygenic phototrophs with EET genes. We developed an automated buoy to measure electric current flow between many pairs of electrodes simultaneously, observing correlation in electron consumption to sunlight. Our results, combined with published metatranscriptomic analysis, indicate the occurrence of electron cycling between phototrophic oxidation (electrotrophic metabolism) by Chlorobium and anaerobic respiration (electrogenic metabolism) by Geothrix, involving DOM. We also repeatedly observed gradual seasonal increases in hypolimnion ORP throughout summer. These diel and seasonal patterns imply that electroactive DOM mediates the ecology of electroactive bacteria in lakes, controlling humic lake methane emissions.IMPORTANCEWe investigated the physical, chemical, and redox characteristics of a bog lake and electrodes hung therein to test the hypothesis that dissolved organic matter is being cycled between oxidized and reduced states by electroactive bacterioplankton powered by phototrophy. To do so, we performed field-based analyses on multiple timescales using both established and novel instrumentation. We paired these analyses with recently developed bioinformatics pipelines for metagenomics data to investigate genes that enable electroactive metabolism and accompanying metabolisms. Our results are consistent with our hypothesis and yet upend some of our other expectations. Our findings have implications for understanding greenhouse gas emissions from lakes, including electroactivity as an integral part of lake metabolism throughout more of the anoxic parts of lakes and for a longer portion of the summer than expected. Our results also give a sense of what electroactivity occurs at given depths and provide a strong basis for future studies.

carbon emissions↗

Bioelectrocatalytic conversion of CO₂ to PHA bioplastics using engineered methylotrophs

The sustainable generation of biodegradable plastics represents an opportunity to capture atmospheric CO 2 while reducing plastic waste accumulation in the environment. This study implements an integrated platform for bioelectrocatalytic CO 2 conversion to medium-chain-length polyhydroxyalkanoates (mcl-PHAs). Immobilizing cobalt phthalocyanine electrocatalysts on a covalent-organic framework in a gas recirculation electrolyzer enabled CO 2 -to-methanol conversion with a carbon conversion efficiency of 98%. Integration of polymer biosynthesis pathways enabled Methylotuvimicrobium alcaliphilum 20Z R to produce ~20% mcl-PHA of the dry cell weight with a CO 2 -to-bioproducts carbon conversion efficiency of 50%. This cell line was adapted to high sodium bicarbonate media, eliminating costly intermediate separation steps while improving economic potential. Transcriptomic analysis revealed sulfate transporters and peptidoglycan biosynthesis as key pathways involved in sodium bicarbonate halotolerance. Altogether, this research presents a foundation for integrating divergent chemical and biological processes into a transformative electrobiomanufacturing platform, addressing the need for alternative pipelines for generating valuable plastics and chemicals.

CO2 utilization↗