Search NASA⌕ Search

SEARCH · Search NASA

Results for “bioinformatics tool”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

AlgaeOrtho, a bioinformatics tool for processing ortholog inference results in algae

Introduction: Microalgae constitute a prominent feedstock for producing biofuels and biochemicals by virtue of their prolific reproduction, high bioproduct accumulation, and the ability to grow in brackish and saline water. However, naturally occurring wild type algal strains are rarely optimal for industrial use; therefore, bioengineering of algae is necessary to generate superior performing strains that can address production challenges in industrial settings, particularly the bioenergy and bioproduct sectors. One of the crucial steps in this process is deciding on a bioengineering target: namely, which gene/protein to differentially express. These targets are often orthologs which are defined as genes/proteins originating from a common ancestor in divergent species. Although bioinformatics tools for the identification of protein orthologs already exist, processing the output from such tools is nontrivial, especially for a researcher with little or no bioinformatics experience. Methods: The present study introduces AlgaeOrtho, a user-friendly tool that builds upon the SonicParanoid orthology inference tool (based on an algorithm that identifies potential protein orthologs based on amino acid sequences) and the PhycoCosm database from JGI (Joint Genome Institute) to help researchers identify orthologs of their proteins of interest in multiple diverse algal species. Results: The output of this application includes a table of the putative orthologs of their protein of interest, a heatmap showing sequence similarity (%), and an unrooted tree of the putative protein orthologs. Notably, the tool would be instrumental in identifying novel bioengineering targets in different algal strains, including targets in not-fully annotated algal species, since it does not depend on existing protein annotations. We tested AlgaeOrtho using three case studies, for which orthologs of proteins relevant to bioengineering targets, were identified from diverse algal species, demonstrating its ease of use and utility for bioengineering researchers. Discussion: This tool is unique in the protein ortholog identification space as it can visualize putative orthologs, as desired by the user, across several algal species.

09 BIOMASS FUELS↗

Benchmark datasets for SARS-CoV-2 surveillance bioinformatics

Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), the cause of coronavirus disease 2019 (COVID-19), has spread globally and is being surveilled with an international genome sequencing effort. Surveillance consists of sample acquisition, library preparation, and whole genome sequencing. This has necessitated a classification scheme detailing Variants of Concern (VOC) and Variants of Interest (VOI), and the rapid expansion of bioinformatics tools for sequence analysis. These bioinformatic tools are means for major actionable results: maintaining quality assurance and checks, defining population structure, performing genomic epidemiology, and inferring lineage to allow reliable and actionable identification and classification. Additionally, the pandemic has required public health laboratories to reach high throughput proficiency in sequencing library preparation and downstream data analysis rapidly. However, both processes can be limited by a lack of a standardized sequence dataset. We identified six SARS-CoV-2 sequence datasets from recent publications, public databases and internal resources. In addition, we created a method to mine public databases to identify representative genomes for these datasets. Using this novel method, we identified several genomes as either VOI/VOC representatives or non-VOI/VOC representatives. To describe each dataset, we utilized a previously published datasets format, which describes accession information and whole dataset information. Additionally, a script from the same publication has been enhanced to download and verify all data from this study.

60 APPLIED LIFE SCIENCES↗

Uncovering novel mutational signatures by de novo extraction with SigProfilerExtractor

Mutational signature analysis is commonly performed in cancer genomic studies. Here, we present SigProfilerExtractor, an automated tool for de novo extraction of mutational signatures, and benchmark it against another 13 bioinformatics tools by using 34 scenarios encompassing 2,500 simulated signatures found in 60,000 synthetic genomes and 20,000 synthetic exomes. For simulations with 5% noise, reflecting high-quality datasets, SigProfilerExtractor outperforms other approaches by elucidating between 20% and 50% more true-positive signatures while yielding 5-fold less false-positive signatures. Applying SigProfilerExtractor to 4,643 whole-genome- and 19,184 whole-exome-sequenced cancers reveals four novel signatures. Two of the signatures are confirmed in independent cohorts, and one of these signatures is associated with tobacco smoking. In summary, this report provides a reference tool for analysis of mutational signatures, a comprehensive benchmarking of bioinformatics tools for extracting signatures, and several novel mutational signatures, including one putatively attributed to direct tobacco smoking mutagenesis in bladder tissues.

59 BASIC BIOLOGICAL SCIENCES↗

CyanoCyc cyanobacterial web portal

CyanoCyc is a web portal that integrates an exceptionally rich database collection of information about cyanobacterial genomes with an extensive suite of bioinformatics tools. It was developed to address the needs of the cyanobacterial research and biotechnology communities. The 277 annotated cyanobacterial genomes currently in CyanoCyc are supplemented with computational inferences including predicted metabolic pathways, operons, protein complexes, and orthologs; and with data imported from external databases, such as protein features and Gene Ontology (GO) terms imported from UniProt. Five of the genome databases have undergone manual curation with input from more than a dozen cyanobacteria experts to correct errors and integrate information from more than 1,765 published articles. CyanoCyc has bioinformatics tools that encompass genome, metabolic pathway and regulatory informatics; omics data analysis; and comparative analyses, including visualizations of multiple genomes aligned at orthologous genes, and comparisons of metabolic networks for multiple organisms. CyanoCyc is a high-quality, reliable knowledgebase that accelerates scientists’ work by enabling users to quickly find accurate information using its powerful set of search tools, to understand gene function through expert mini-reviews with citations, to acquire information quickly using its interactive visualization tools, and to inform better decision-making for fundamental and applied research.

59 BASIC BIOLOGICAL SCIENCES↗

Comprehensive characterization of N- and O- glycosylation of SARS-CoV-2 human receptor angiotensin converting enzyme 2

The emergence of the coronavirus disease 2019 (COVID-19) pandemic caused by severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) has created the need for development of new therapeutic strategies. Understanding the mode of viral attachment, entry and replication has become a key aspect of such interventions. The coronavirus surface features a trimeric spike (S) protein that is essential for viral attachment, entry and membrane fusion. The S protein of SARS-CoV-2 binds to human angiotensin converting enzyme 2 (hACE2) for entry. Herein, we describe glycomic and glycoproteomic analysis of hACE2 expressed in HEK293 cells. We observed high glycan occupancy (73.2 to 100%) at all seven possible N-glycosylation sites and surprisingly detected one novel O-glycosylation site. To deduce the detailed structure of glycan epitopes on hACE2 that may be involved in viral binding, we have characterized the terminal sialic acid linkages, the presence of bisecting GlcNAc and the pattern of N-glycan fucosylation. We have conducted extensive manual interpretation of each glycopeptide and glycan spectrum, in addition to using bioinformatics tools to validate the hACE2 glycosylation. Our elucidation of the site-specific glycosylation and its terminal orientations on the hACE2 receptor, along with the modeling of hACE2 glycosylation sites can aid in understanding the intriguing virus-receptor interactions and assist in the development of novel therapeutics to prevent viral entry. Here, the relevance of studying the role of ACE2 is further increased due to some recent reports about the varying ACE2 dependent complications with regard to age, sex, race and pre-existing conditions of COVID-19 patients.

59 BASIC BIOLOGICAL SCIENCES↗

Integrated Phage-Host Prediction tool (iPHoP) v1.0.0

iPHoP is a bioinformatic tools that uses a set of approaches to predict the potential host of novel bacteriophages (viruses infecting bacteria) that are only known by their genome sequence, and not cultivated in the laboratory. Existing technologies typically rely on a single method, and the main advantage of iPHoP is its ability to integrate the results from multiple methods into a single prediction. This is of interest for microbial ecology researchers, as they often analyze novel bacteriophage genomes that they were able to assemble from metagenomes, but they don't know which bacteria these phages infect.

Roux, Simon↗

Bioscience COVID Rapid Response Report

The COVID-19 disease outbreak and its impact on global health and economies have highlighted the national security threat posed by pathogens with pandemic potential and the need for rapid development of effective diagnostics and medical countermeasures. The Bioscience IA selected for funding rapid COVID LDRD project proposals that addressed critical R&D gaps in pandemic response that could be accomplished in 1-3 months with the requested funding. In total, the Bioscience IA funded nine rapid projects that addressed 1) rapid and accurate methods for SARS-CoV-2 RNA detection, 2) modeling tools to help prioritize populations for diagnostic testing, 3) bioinformatic tools to track SARS-CoV-2 genomic sequence changes over time, 4) molecular inhibitors of SARS-CoV-2 cellular infection, and 5) method for rapid staging of COVID19 disease to enable administration of more effective treatments. In addition, LDRD funded one larger project to be completed in FY21 that leverages Sandia capabilities to address the need for platform diagnostics and therapeutics that can be rapidly tailored against emerging pathogen targets.

59 BASIC BIOLOGICAL SCIENCES↗

KBase Narrative - SOMATA - Motivating Example

A systems view in the biological context is a useful approach to study complex and potentially interacting systems. In this narrative (accompanying an associated manuscript), we take a systems view of software tools themselves. Instead of using software as a single tool, we demonstrate how users can approach software with the same mindset as the organisms they study; as a complex system that can be optimized based on one’s environment and goals. We demonstrate this idea with an example KBase narrative showing how application parameters in KBase tools can change how a researcher perceives and interacts with bioinformatics tools while conducting a common experimental scenario. In this scenario we are trying to understand how different chemical compounds in a growth media change the metabolic pathways utilized in Escherichia coli.

Cashman, Mikaela↗

Expanding the Scope of Genomic Security: Targeted Genome Editing within Microbiomes through Designer Bacteriophage Vectors

The ability to engineer the genome of a bacterial strain, not as an isolate, but while present among other microbes in a microbiome, would open new technological possibilities in the areas of medicine, energy and biomanufacturing. Our approach is to develop sets of phages (bacterial viruses) active on the target strain and themselves engineered to act not as killers but as vectors for gene delivery. This approach is rooted in our bioinformatic tools that map prophages accurately within bacterial genomes. We present new bioinformatic results in cross-contig search, design of phage genome assemblies, satellites that embed within prophages, alignment of large numbers of biological sequences, and improvement of reference databases for prophage discovery. We targeted a Pseudomonas putida strain within a lignin-degrading microbiome, but were unable to obtain active phages, and turned toward a defined microbiome of the mouse gut.

59 BASIC BIOLOGICAL SCIENCES↗

Introducing the Bacterial and Viral Bioinformatics Resource Center (BV-BRC): a resource combining PATRIC, IRD and ViPR

The National Institute of Allergy and Infectious Diseases (NIAID) established the Bioinformatics Resource Center (BRC) program to assist researchers with analyzing the growing body of genome sequence and other omics-related data. In this report, we describe the merger of the PAThosystems Resource Integration Center (PATRIC), the Influenza Research Database (IRD) and the Virus Pathogen Database and Analysis Resource (ViPR) BRCs to form the Bacterial and Viral Bioinformatics Resource Center (BV-BRC) https://www.bv-brc.org/. The combined BV-BRC leverages the functionality of the bacterial and viral resources to provide a unified data model, enhanced web-based visualization and analysis tools, bioinformatics services, and a powerful suite of command line tools that benefit the bacterial and viral research communities.

59 BASIC BIOLOGICAL SCIENCES↗

HIResist: a database of HIV-1 resistance to broadly neutralizing antibodies

Changing the course of the human immunodeficiency virus type I (HIV-1) pandemic is a high public health priority with approximately 39 million people currently living with HIV-1 (PLWH) and about 1.5 million new infections annually worldwide. Broadly neutralizing antibodies (bnAbs) typically target highly conserved sites on the HIV-1 envelope glycoproteins (Envs), which mediate viral entry, and block the infection of diverse HIV-1 strains. But different mechanisms of HIV-1 resistance to bnAbs prevent robust application of bnAbs for therapeutic and preventive interventions. Here we report the development of a new database that provides data and computational tools to aid the discovery of resistant features and may assist in analysis of HIV-1 resistance to bnAbs. Bioinformatic tools allow identification of specific patterns in Env sequences of resistant strains and development of strategies to elucidate the mechanisms of HIV-1 escape; comparison of resistant and sensitive HIV-1 strains for each bnAb; identification of resistance and sensitivity signatures associated with specific bnAbs or groups of bnAbs; and visualization of antibody pairs on cross-sensitivity plots. The database has been designed with a particular focus on user-friendly and interactive interface. Our database is a valuable resource for the scientific community and provides opportunities to investigate patterns of HIV-1 resistance and to develop new approaches aimed to overcome HIV-1 resistance to bnAbs.

60 APPLIED LIFE SCIENCES↗

Systematic benchmarking demonstrates large language models have not reached the diagnostic accuracy of traditional rare-disease decision support tools

Large language models (LLMs) show promise in supporting differential diagnosis, but their performance is challenging to evaluate due to the unstructured nature of their responses, and their accuracy compared to existing diagnostic tools is not well characterized. To assess the current capabilities of LLMs to diagnose genetic diseases, we benchmarked these models on 5213 previously published case reports using the Phenopacket Schema, the Human Phenotype Ontology and Mondo disease ontology. Prompts generated from each phenopacket were sent to seven LLMs, including four generalist models and three LLMs specialized for medical applications. The same phenopackets were used as input to a widely used diagnostic tool, Exomiser, in phenotype-only mode. The best LLM ranked the correct diagnosis first in 23.6% of cases, whereas Exomiser did so in 35.5% of cases. While the performance of LLMs for supporting differential diagnosis has been improving, it has not reached the level of commonly used traditional bioinformatics tools. Future research is needed to determine the best approach to incorporate LLMs into diagnostic pipelines.

Reese, Justin T. [Lawrence Berkeley National Labor↗

Interpreting the Lipidome: Bioinformatic Approaches to Embrace the Complexity

Background Improvements in mass spectrometry (MS) technologies coupled with bioinformatics developments have allowed considerable advancement in the measurement and interpretation of lipidomics data in recent years. Since research areas employing lipidomics are rapidly increasing, there is a great need for bioinformatic tools that capture and utilize the complexity of the data. Currently, the diversity and complexity within the lipidome is often concealed by summing over or averaging individual lipids up to (sub)class-based descriptors, losing valuable information about biological function and interactions with other distinct lipids molecules, proteins and/or metabolites. Aim of review To address this gap in knowledge, novel bioinformatics methods are needed to improve identification, quantification, integration and interpretation of lipidomics data. The purpose of this mini-review is to summarize exemplary methods to explore the complexity of the lipidome. Key scientific concepts of review Here we describe six approaches that capture three core focus areas for lipidomics: (1) lipidome annotation including a resolvable database identifier, (2) interpretation via pathway- and enrichment-based methods, and (3) understanding complex interactions to emphasize specific steps in the analytical process and highlight challenges in analyses associated with the complexity of lipidome data.

Kyle, Jennifer E.↗

Challenges and Opportunities in State‐of‐the‐Art Proteomics Analysis for Biomarker Development From Plasma Extracellular Vesicles

Extracellular vesicles (EVs) are membrane-bound particles secreted by cells, playing crucial roles in intercellular communication. The composition of EVs can undergo changes in response to stress and disease conditions, making them excellent biomarker candidates. However, extracting protein information from EVs can be challenging due to their low abundance in complex biofluids and copurification with contaminant proteins and particles. Techniques to enrich EVs have their strengths and limitations, without one being able to purify EVs to complete homogeneity. This can lead to compromised recovery rates and increased complexity, making data interpretation difficult. In this viewpoint article, we explore the concept that better characterization of EV composition, followed by quantification of EV proteins in complex samples, might be a more viable route for biomarker development. Mass spectrometers can provide reproducible deep coverage of the EV proteome, despite sample impurities. This paradigm shift presents opportunities to integrate advanced bioinformatics tools to refine the EV proteome landscape, identify novel biomarkers, and streamline validation processes in biomarker development. By focusing on leveraging technology rather than achieving absolute purity, this approach can transform current practices and open opportunities for robust biomarker discovery. Herein, we highlight not only such opportunities but also challenges to implement this concept.

Dakup, Panshak P. [Pacific Northwest National Labo↗

Interpreting omics data with pathway enrichment analysis

Pathway enrichment analysis is indispensable for interpreting omics datasets and generating hypotheses. However, the foundations of enrichment analysis remain elusive to many biologists. Here, in this study, we discuss best practices in interpreting different types of omics data using pathway enrichment analysis and highlight the importance of considering intrinsic features of various types of omics data. We further explain major components that influence the outcomes of a pathway enrichment analysis, including defining background sets and choosing reference annotation databases. To improve reproducibility, we describe how to standardize reporting methodological details in publications. This article aims to serve as a primer for biologists to leverage the wealth of omics resources and motivate bioinformatics tool developers to enhance the power of pathway enrichment analysis.

60 APPLIED LIFE SCIENCES↗

PathTracer Comprehensively Identifies Hypoxia-Induced Dormancy Adaptations in Mycobacterium tuberculosis

Mining large-scale data to discover biologically relevant information remains a challenge despite the rapid development of bioinformatics tools. Here, we have developed a new tool, PathTracer, to identify biologically relevant information flows by mining genome-wide protein–protein interaction networks following integration of gene expression data. PathTracer successfully mines interactions between genes and traces the most perturbed paths of perceived activities under the conditions of the study. Here, we further demonstrated the utility of this tool by identifying adaptation mechanisms of hypoxia-induced dormancy in Mycobacterium tuberculosis (Mtb).

59 BASIC BIOLOGICAL SCIENCES↗

Redefining fundamental concepts of transcription initiation in bacteria

Despite enormous progress in understanding the fundamentals of bacterial gene regulation, our knowledge remains limited when compared with the number of bacterial genomes and regulatory systems to be discovered. Derived from a small number of initial studies, classic definitions for concepts of gene regulation have evolved as the number of characterized promoters has increased. Together with discoveries made using new technologies, this knowledge has led to revised generalizations and principles. In this Expert Recommendation, we suggest precise, updated definitions that support a logical, consistent conceptual framework of bacterial gene regulation, focusing on transcription initiation. The resulting concepts can be formalized by ontologies for computational modelling, laying the foundation for improved bioinformatics tools, knowledge-based resources and scientific communication. Furthermore, this work will help researchers construct better predictive models, with different formalisms, that will be useful in engineering, synthetic biology, microbiology and genetics.

59 BASIC BIOLOGICAL SCIENCES↗

GenomeDepot: data management system for microbial comparative genomics

Summary GenomeDepot is an open-source web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of websites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, Basic Local Alignment Search Tool (BLAST) search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools. Availability and implementation GenomeDepot is open source and distributed under the GNU General Public License via GitHub (https://github.com/aekazakov/genome-depot). GenomeDepot is implemented in Python and was tested in Ubuntu Linux. Full installation instructions and documentation are available at https://aekazakov.github.io/genome-depot/. GenomeDepot demo server is freely accessible at https://iseq.lbl.gov/demogd/.

Kazakov, Alexey [Lawrence Berkeley National Labora↗