Search NASA⌕ Search

SEARCH · Search NASA

Results for “Sequence annotation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Gaia: An AI-enabled genomic context–aware platform for protein sequence annotation

Protein sequence similarity search is fundamental to biology research, but current methods are typically not able to consider crucial genomic context information indicative of protein function, especially in microbial systems. Here, we present Gaia (Genomic AI Annotator), a sequence annotation platform that enables rapid, context-aware protein sequence search across genomic datasets. Gaia leverages gLM2, a mixed-modality genomic language model trained on both amino acid sequences and their genomic neighborhoods to generate embeddings that integrate sequence-structure-context information. This approach allows for the identification of functionally and/or evolutionarily related genes that are found in conserved genomic contexts, which may be missed by traditional sequence- or structure-based search alone. Gaia enables real-time search of a curated database comprising more than 85 million protein clusters from 131,744 microbial genomes. We compare the homolog retrieval performance of Gaia search against other embedding and alignment-based approaches. We provide Gaia as a web-based, freely available tool.

Jha, Nishant↗

An FPGA-based hardware accelerator supporting sensitive sequence homology filtering with profile hidden Markov models

Abstract Background Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). Here, we describe an FPGA hardware accelerator, called HAVAC, that targets a key bottleneck step (SSV) in the analysis pipeline of the popular pHMM alignment tool, HMMER. Results The HAVAC kernel calculates the SSV matrix at 1739 GCUPS on a $$\sim$$ ∼ $3000 Xilinx Alveo U50 FPGA accelerator card, $$\sim$$ ∼ 227× faster than the optimized SSV implementation in nhmmer . Accounting for PCI-e data transfer data processing, HAVAC is 65× faster than nhmmer’s SSV with one thread and 35× faster than nhmmer with four threads, and uses $$\sim$$ ∼ 31% the energy of a traditional high end Intel CPU. Conclusions HAVAC demonstrates the potential offered by FPGA hardware accelerators to produce dramatic speed gains in sequence annotation and related bioinformatics applications. Because these computations are performed on a co-processor, the host CPU remains free to simultaneously compute other aspects of the analysis pipeline.

59 BASIC BIOLOGICAL SCIENCES↗

RCSB protein data Bank: Next‐generation advanced search for exploration of experimental structures and computed structure models

Abstract The Protein Data Bank (PDB), established in 1971, is the primary global, open‐access archive for experimentally determined 3D macromolecular structures (proteins, RNA, DNA). The research‐focused RCSB.org web‐portal provides access to these data alongside more than one million machine‐learning‐predicted structure models, greatly expanding the available structural landscape. Rapid growth of both experimental and computational structures has increased the need for powerful yet accessible search tools that serve a broad and diverse scientific community. Herein, we describe a redesigned RCSB Protein Data Bank RCSB.org Advanced Search capability that supports intuitive discovery of 3D structures through a unified interface. This interface integrates annotation‐, sequence‐, and 3D structure‐based searches, embeds an interactive 3D viewer, and incorporates curated biological knowledge, such as catalytic site definitions from Mechanism and Catalytic Site Atlas and ligand‐guided structural motifs, for constructing geometry‐driven queries. A new Chemical Search tool allows definition of chemical queries via an integrated drawing tool or standard identifiers, seamlessly combining them with annotation filters. By allowing query definition directly within spatial and chemical contexts, these search interfaces reduce the need for detailed knowledge of residue numbering, chain identifiers, or external cheminformatics software. This capability enables efficient exploration of structures, chemical diversity, and structure–function relationships across all life domains. The redesigned interfaces can be accessed directly at rcsb.org/search/advanced for Advanced Search and rcsb.org/search/chemical for Chemical Search.

Rose, Yana [Research Collaboratory for Structural ↗

Genomic and morphological characterization of Knufia obscura isolated from the Mars 2020 spacecraft assembly facility

Members of the family Trichomeriaceae, belonging to the Chaetothyriales order and the Ascomycota phylum, are known for their capability to inhabit hostile environments characterized by extreme temperatures, oligotrophic conditions, drought, or presence of toxic compounds. The genus Knufia encompasses many polyextremophilic species. In this report, the genomic and morphological features of the strain FJI-L2-BK-P2 presented, which was isolated from the Mars 2020 mission spacecraft assembly facility located at the Jet Propulsion Laboratory in Pasadena, California. The identification is based on sequence alignment for marker genes, multi-locus sequence analysis, and whole genome sequence phylogeny. The morphological features were studied using a diverse range of microscopic techniques (bright field, phase contrast, differential interference contrast and scanning electron microscopy). The phylogenetic marker genes of the strain FJI-L2-BK-P2 exhibited highest similarities with type strain of Knufia obscura (CBS 148926 T ) that was isolated from the gas tank of a car in Italy. To validate the species identity, whole genomes of both strains (FJI-L2-BK-P2 and CBS 148926 T ) were sequenced, annotated, and strain FJI-L2-BK-P2 was confirmed as K. obscura. The morphological analysis and description of the genomic characteristics of K. obscura FJI-L2-BK-P2 may contribute to refining the taxonomy of Knufia species. Key morphological features are reported in this K. obscura strain, resembling microsclerotia and chlamydospore-like propagules. These features known to be characteristic features in black fungi which could potentially facilitate their adaptation to harsh environments.

59 BASIC BIOLOGICAL SCIENCES↗

Biological Parts Search Portal (BioParts) v1.0.0

BioParts is a web based search portal for biological parts available in the public domain. It combines the ease and convenience of modern web search engines with the capabilities of bioinformatics search tools such as BLAST. This portal, available at bioparts.org, allows anyone to search for publicly accessible biological part information (e.g., NCBI, iGEM, SynBioHub, Addgene), including parts publicly accessible through ICE Registries. Additionally, the portal offers a REST API that enables third-party applications and tools to access the portal's functionality programmatically. While there are several standalone biological part repositories, there doesn't exist an application that indexes these publicly available parts and enables features such as keyword and BLAST searches along with automatic sequence annotation.

Plahar, Hector↗

VIBES: a workflow for annotating and visualizing viral sequences integrated into bacterial genomes

Abstract Bacteriophages are viruses that infect bacteria. Many bacteriophages integrate their genomes into the bacterial chromosome and become prophages. Prophages may substantially burden or benefit host bacteria fitness, acting in some cases as parasites and in others as mutualists. Some prophages have been demonstrated to increase host virulence. The increasing ease of bacterial genome sequencing provides an opportunity to deeply explore prophage prevalence and insertion sites. Here we present VIBES (Viral Integrations in Bacterial genomES), a workflow intended to automate prophage annotation in complete bacterial genome sequences. VIBES provides additional context to prophage annotations by annotating bacterial genes and viral proteins in user-provided bacterial and viral genomes. The VIBES pipeline is implemented as a Nextflow-driven workflow, providing a simple, unified interface for execution on local, cluster and cloud computing environments. For each step of the pipeline, a container including all necessary software dependencies is provided. VIBES produces results in simple tab-separated format and generates intuitive and interactive visualizations for data exploration. Despite VIBES’s primary emphasis on prophage annotation, its generic alignment-based design allows it to be deployed as a general-purpose sequence similarity search manager. We demonstrate the utility of the VIBES prophage annotation workflow by searching for 178 Pf phage genomes across 1072 Pseudomonas spp. genomes.

59 BASIC BIOLOGICAL SCIENCES↗

ULTRA-effective labeling of tandem repeats in genomic sequence

In the age of long read sequencing, genomics researchers now have access to accurate repetitive DNA sequence (including satellites) that, due to the limitations of short read-sequencing, could previously be observed only as unmappable fragments. Tools that annotate repetitive sequence are now more important than ever, so that we can better understand newly uncovered repetitive sequences, and also so that we can mitigate errors in bioinformatic software caused by those repetitive sequences. To that end, we introduce the 1.0 release of our tool for identifying and annotating locally repetitive sequence, ULTRA Locates Tandemly Repetitive Areas (ULTRA). ULTRA is fast enough to use as part of an efficient annotation pipeline, produces state-of-the-art reliable coverage of repetitive regions containing many mutations, and provides interpretable statistics and labels for repetitive regions.

59 BASIC BIOLOGICAL SCIENCES↗

Structure-aware annotation of leucine-rich repeat domains

Protein domain annotation is typically done by predictive models such as HMMs trained on sequence motifs. However, sequence-based annotation methods are prone to error, particularly in calling domain boundaries and motifs within them. These methods are limited by a lack of structural information accessible to the model. With the advent of deep learning-based protein structure prediction, existing sequenced-based domain annotation methods can be improved by taking into account the geometry of protein structures. We develop dimensionality reduction methods to annotate repeat units of the Leucine Rich Repeat solenoid domain. The methods are able to correct mistakes made by existing machine learning-based annotation tools and enable the automated detection of hairpin loops and structural anomalies in the solenoid. The methods are applied to 127 predicted structures of LRR-containing intracellular innate immune proteins in the model plant Arabidopsis thaliana and validated against a benchmark dataset of 172 manually-annotated LRR domains.

Xu, Boyan↗

metagRoot: a comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71 091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, hidden Markov models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org.

Chasapi, Maria N↗

A large sequenced mutant library – valuable reverse genetic resource that covers 98% of sorghum genes

SUMMARY Mutant populations are crucial for functional genomics and discovering novel traits for crop breeding. Sorghum , a drought and heat‐tolerant C4 species, requires a vast, large‐scale, annotated, and sequenced mutant resource to enhance crop improvement through functional genomics research. Here, we report a sorghum large‐scale sequenced mutant population with 9.5 million ethyl methane sulfonate (EMS)‐induced mutations that covered 98% of sorghum's annotated genes using inbred line BTx623. Remarkably, a total of 610 320 mutations within the promoter and enhancer regions of 18 000 and 11 790 genes, respectively, can be leveraged for novel research of cis ‐regulatory elements. A comparison of the distribution of mutations in the large‐scale mutant library and sorghum association panel (SAP) provides insights into the influence of selection. EMS‐induced mutations appeared to be random across different regions of the genome without significant enrichment in different sections of a gene, including the 5′ UTR, gene body, and 3′‐UTR. In contrast, there were low variation density in the coding and UTR regions in the SAP. Based on the K a / K s value, the mutant library (~1) experienced little selection, unlike the SAP (0.40), which has been strongly selected through breeding. All mutation data are publicly searchable through SorbMutDB ( https://www.depts.ttu.edu/igcast/sorbmutdb.php ) and SorghumBase ( https://sorghumbase.org/ ). This current large‐scale sequence‐indexed sorghum mutant population is a crucial resource that enriched the sorghum gene pool with novel diversity and a highly valuable tool for the Poaceae family, that will advance plant biology research and crop breeding.

59 BASIC BIOLOGICAL SCIENCES↗

ContScout: sensitive detection and removal of contamination from annotated genomes

Contamination of genomes is an increasingly recognized problem affecting several downstream applications, from comparative evolutionary genomics to metagenomics. Here we introduce ContScout, a precise tool for eliminating foreign sequences from annotated genomes. It achieves high specificity and sensitivity on synthetic benchmark data even when the contaminant is a closely related species, outperforms competing tools, and can distinguish horizontal gene transfer from contamination. A screen of 844 eukaryotic genomes for contamination identified bacteria as the most common source, followed by fungi and plants. Furthermore, we show that contaminants in ancestral genome reconstructions lead to erroneous early origins of genes and inflate gene loss rates, leading to a false notion of complex ancestral genomes. Taken together, we offer here a tool for sensitive removal of foreign proteins, identify and remove contaminants from diverse eukaryotic genomes and evaluate their impact on phylogenomic analyses.

59 BASIC BIOLOGICAL SCIENCES↗

Near-complete genome sequence of Lipomyces tetrasporous NRRL Y-64009, an oleaginous yeast capable of growing on lignocellulosic hydrolysates

ABSTRACT Lipomyces tetrasporous is an oleaginous yeast that can utilize a variety of plant-based sugars. It accumulates lipids during growth on lignocellulosic biomass hydrolysates. We present the annotated genome sequence of L. tetrasporous NRRL Y-64009 to aid in its development as a platform organism for producing lipids and lipid-based bioproducts.

59 BASIC BIOLOGICAL SCIENCES↗

Structural Models and Sequence Alignment Results of the Desulfovibrio vulgaris Proteome

This dataset contains the structural models for the primary transcripts of the Desulfovibrio vulgaris proteome as well as sequence alignment results for a subset of the encoded proteins. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the D. vulgaris proteome to those available in the AlphaFold Protein Structure Database (AFDB). This is a bit more complicated since the proteins reporting in the AFDB originate from an outdated form of the D. vulgaris sequence. The different versions of the D. vulgaris gene annotation are collected in the Chronology subdirectory; further consideration of these changes on the structural space of the proteome are currently underway. For proteins that have been annotated as hypothetical, sequence alignment results from the HHblits and SAdLSA alignment methods are provided. These methods are often more capable to resolve sequence homology than other methods. Therefore, the results from both HHblits and SAdLSA are provided to identify possible homologs for these challenging proteins. Numerous sequence databases are utilized for these alignments. References AlphaFold v2 Multimer: https://doi.org/10.1101/2021.10.04.463034. References HHblits: hhtps://doi.org/10.1186/s12859-019-3019-7. References SAdLSA: hhtps://doi.org/10.3389/fbinf.2021.689960.

59 BASIC BIOLOGICAL SCIENCES↗

Semi-supervised permutation invariant particle-level anomaly detection

The development of analysis methods to distinguish potential beyond the Standard Model phenomena in a model-agnostic way can significantly enhance the discovery reach in collider experiments. However, the typical machine learning (ML) algorithms employed for this task require fixed length and ordered inputs that break the natural permutation invariance in collision events. To address this, a semi-supervised anomaly detection tool is presented that takes a variable number of particle-level inputs and leverages a signal model to encode this information into a permutation invariant, event-level representation via supervised training with a Particle Flow Network (PFN). Data events are then encoded into this representation and given as input to an autoencoder for unsupervised ANomaly deTEction on particLe flOw latent sPacE (ANTELOPE), classifying anomalous events based on a low-level and permutation invariant input modeling. Performance of the ANTELOPE architecture is evaluated on simulated samples of hadronic processes in a high energy collider experiment, showing good capability to distinguish disparate models of new physics.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

MjCyc: Rediscovering the pathway-genome landscape of the first sequenced archaeon, Methanocaldococcus (Methanococcus) jannaschii

The genome of Methanocaldococcus (Methanococcus) jannaschii DSM 2661 was the first Archaeal genome to be sequenced in 1996. Subsequent sequence-based annotation cycles led to its first metabolic reconstruction in 2005. Leveraging new experimental results and function assignments, we have now re-annotated M. jannaschii, creating an updated resource with novel information and testable predictions in a pathway-genome database available at BioCyc.org. This reannotation effort has resulted in 652 function assignments with enzyme roles, accounting for a third of the total protein-coding entries for this genome. The updated resource includes 883 reactions, 540 enzymes, and 142 individual pathways. Despite notable progress in computational genomics, more than a third of the genome remains functionally uncharacterized. The publicly available MjCyc pathway-genome database holds great potential for the wider community to conduct research on the biology of methanogenic Archaea.

59 BASIC BIOLOGICAL SCIENCES↗

The Factors Governing Metal Dependence of an Emergent Superfamily of Bimetallic Oxygenases

Metalloenzyme superfamilies are typically defined by their protein scaffolds and active sites. Owing to the high tunability of protein structures, members of a single superfamily can catalyze diverse reactions with the same metallocofactor. Some superfamilies, such as amidohydrolase-related dinuclear oxygenases (AROs), display further versatility by utilizing multiple metallocofactors. We have shown that certain AROs catalyze monooxygenation reactions with diiron, dimanganese, and/or mixed manganese−iron cofactors, but the molecular factors governing the selection of a particular cofactor remain unknown, and the extent of this superfamily in biology is unclear. Here, we report bioinformatic analyses that expand the ARO superfamily to approximately 17,000 unique UniProt sequences, far exceeding the number of previously characterized enzymes. Through the integration of structural, spectroscopic, and thermodynamic analyses of representative proteins with a bioinformatic pipeline that identifies key secondary- and tertiary-sphere residues, we can predict in silico the metal preference for the majority of reported ARO sequences. These annotations were validated via the characterization of multiple new AROs, including ones implicated in key oxidative steps of natural product biosyntheses. This study establishes the key structure−function relationships governing metal preferences in AROs and highlights their vastly underappreciated role in myriad biological processes.

Liu, Chang [University of California, Berkeley, CA↗

Compilation and utilization of a sorghum transcriptome compendium for gene regulatory network analysis and crop trait engineering

Sorghum bicolor (Sorghum) is a drought and heat tolerant C4 grass crop used to produce grain, forage, biofuels, and other bioproducts. Genetic improvement of sorghum hybrid crops is aided by a large and diverse germplasm, sorghum's diploid inbreeding genetics, and a relatively small genome that has facilitated genomic research. Over the past 20 years, the sorghum research community characterized the cytogenetic and recombinant landscapes of sorghum's 10 chromosomes, sequenced and annotated the sorghum genome, and used that information to identify genes/alleles that modulate flowering time, plant height, seed shattering, and other important traits. More recently, >1000 RNA-seq transcriptome profiles were collected from 15 sorghum genotypes to help understand the genetic basis of variation in growth and development of sorghum stems, tillers, roots, and leaves, and the regulation of biosynthetic pathways that produce epicuticular wax, dhurrin, and RFOs, compounds that contribute to sorghum's resilience. Transcriptome studies were designed to identify differentially expressed genes that are co-expressed during development or in response to a treatment to enable construction of gene regulatory networks. Co-expression and network analysis identified transcription factors and their cognate binding sites in target gene promoters and signaling pathways that modulate gene regulatory networks providing gene editing targets for further trait optimization. RNA-seq data from >20 experiments targeting sorghum organs, tissues, cell types, developmental stages, and responses to environmental conditions (i.e., diel, day-length, shading, water-deficit, temperature) has been compiled in a sorghum transcriptome compendium. The goal of this resource paper is to describe compendium content, accessibility, and a compendium data analysis pipeline and to illustrate the types of information that can be derived from the compendium with a focus on the elucidation of gene regulatory networks useful for guiding the improvement of sorghum traits through gene editing.

RNA-seq↗

Inventory of Composable Elements (ICE) v6.0.0

The Inventory of Composable Elements (ICE) is an open source registry software platform for managing information about biological parts. It is capable of recording information about plasmids, microbial host strains and seeds, as well as DNA parts. Includes features such as DNA sequence visualization, editing and annotation, auto-aligning sequencing trace files against reference templates, SBOL XML/RDF support, and web-of-registries functionality. The web of registries functionality provides strong support for distributed interconnected use and enables sharing and transfer of biological parts across various independent ICE instances. ICE adopts modern software development principles, leveraging component-base frameworks, offering a REST API for convenient third-party integration and emphasizing scalability, security, and service integrations for dynamic content availability. The source code is hosted at https://github.com/JBEI/ice. A public instance is available at public-registry.jbei.org, where users can try out features, upload parts or simply use it for their projects.

Plahar, Hector↗