Search NASASearch

SEARCH · Search NASA

Results for “Protein function predictions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Interactive tools for functional annotation of bacterial genomes

Automated annotations of protein functions are error-prone because of our lack of knowledge of protein functions. For example, it is often impossible to predict the correct substrate for an enzyme or a transporter. Furthermore, much of the knowledge that we do have about the functions of proteins is missing from the underlying databases. We discuss how to use interactive tools to quickly find different kinds of information relevant to a protein’s function. Many of these tools are available via PaperBLAST (http://papers.genomics.lbl.gov). Combining these tools often allows us to infer a protein’s function. Ideally, accurate annotations would allow us to predict a bacterium’s capabilities from its genome sequence, but in practice, this remains challenging. We describe interactive tools that infer potential capabilities from a genome sequence or that search a genome to find proteins that might perform a specific function of interest.

59 BASIC BIOLOGICAL SCIENCES

Artificial Intelligence Transforming Post-Translational Modification Research

Post-Translational Modifications (PTMs) are covalent changes to amino acids that occur after protein synthesis, including covalent modifications on side chains and peptide backbones. Many PTMs profoundly impact cellular and molecular functions and structures, and their significance extends to evolutionary studies as well. In light of these implications, we have explored how artificial intelligence (AI) can be utilized in researching PTMs. Initially, rationales for adopting AI and its advantages in understanding the functions of PTMs are discussed. Then, various deep learning architectures and programs, including recent applications of language models, for predicting PTM sites on proteins and the regulatory functions of these PTMs are compared. Finally, our high-throughput PTM-data-generation pipeline, which formats data suitably for AI training and predictions is described. We hope this review illuminates areas where future AI models on PTMs can be improved, thereby contributing to the field of PTM bioengineering.

59 BASIC BIOLOGICAL SCIENCES

Simple Math is Enough: Two Examples of Inferring Functional Associations from Genomic Data

Non-random features in the genomic data are usually biologically meaningful. The key is to choose the feature well. Having a p-value based score prioritizes the findings. If two proteins share a unusually large number of common interaction partners, they tend to be involved in the same biological process. We used this finding to predict the functions of 81 un-annotated proteins in yeast.

Liang, Shoudan

Birth of protein folds and functions in the virome

The rapid evolution of viruses generates proteins that are essential for infectivity and replication but with unknown functions, due to extreme sequence divergence. Here, using a database of 67,715 newly predicted protein structures from 4,463 eukaryotic viral species, we found that 62% of viral proteins are structurally distinct and lack homologues in the AlphaFold database. Among the remaining 38% of viral proteins, many have non-viral structural analogues that revealed surprising similarities between human pathogens and their eukaryotic hosts. Structural comparisons suggested putative functions for up to 25% of unannotated viral proteins, including those with roles in the evasion of innate immunity. In particular, RNA ligase T-like phosphodiesterases were found to resemble phage-encoded proteins that hydrolyse the host immune-activating cyclic dinucleotides 3',3'- and 2',3'-cyclic GMP-AMP (cGAMP). Experimental analysis showed that RNA ligase T homologues encoded by avian poxviruses similarly hydrolyse cGAMP, showing that RNA ligase T-mediated targeting of cGAMP is an evolutionarily conserved mechanism of immune evasion that is present in both bacteriophage and eukaryotic viruses. Together, the viral protein structural database and analyses presented here afford new opportunities to identify mechanisms of virus–host interactions that are common across the virome.

59 BASIC BIOLOGICAL SCIENCES

Finding the missing pieces: filling gaps that impede the translation of omics data into models

High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.

59 BASIC BIOLOGICAL SCIENCES

Maize Rough Endosperm6 (rgh6) Encodes A Predicted Dead-Box RNA Helicase and Affects Mirna Processing in Endosperm Development

Maize rough endosperm (rgh) mutants have defective kernels with a rough, etched, or pitted endosperm surface. Molecular genetic analysis of this mutant class has identified multiple RNA processing proteins critical to endosperm development. Here, we report on the developmental and molecular function of the rgh6 locus. The rgh6 mutant was isolated from the UniformMu transposon tagging population. Mutant kernels have reduced endosperm size and defective embryos that develop in a more apical position than typical for defective embryos. TB translocation crosses revealed that rgh6 mutant endosperm inhibits normal embryo development. Positional cloning of the rgh6 locus found that it encodes a predicted DEAD-box RNA helicase. Consistent with a predicted function for RNA processing, transient expression of a RGH6-GFP fusion protein is localized to nucleolus and nuclear speckles in Nicotiana benthamiana leaves. Rgh6 transcripts are highly expressed in endosperm epidermal cell types such as the aleurone, basal endosperm cell layer, embryo surrounding region, and endosperm adjacent to scutellum. Markers of these cell types show increased levels in rgh6 mutant kernels. Mutant endosperm tissues have increased precursor microRNA (pre-miRNA) and decreased mature miRNA relative to normal sibling endosperm, indicating that rgh6 is required for miRNA processing. The transcript levels for most miRNA target genes accumulate to a higher level in rgh6 mutant tissue. These results suggest that miRNA processing and regulation of miRNA target genes are required for normal endosperm development.

Plant Sciences

Energy metric prediction for double insertion mutants via the RoseNet deep learning framework

Studying the structural and functional implications of protein mutations is an important task in computational biology and bioinformatics. We leverage our previously proposed RoseNet neural network architecture to predict energy metrics of proteins with double amino acid insertions or deletions (InDels). We train models on previously generated benchmark datasets containing the exhaustive double InDel mutations for three proteins, as well as an additional three proteins for which ∼145k random mutants, each with two InDels, have been generated. We expand on our previous work by evaluating three additional proteins and analyzing domain features that impact the prediction capabilities of RoseNet. These features include InDels into secondary structures and the solvent accessible surface area (SASA) scores of the residues. We uncover further evidence to support that RoseNet has a higher proficiency of generalizing to unseen residue combinations than unseen insertion positions. We also observe that RoseNet produces higher-quality predictions when inserting into a β-sheet over an α-helix. Additionally, when the insertions fall in an area of high SASA, RoseNet often displays better performance than inserting into areas of low SASA.

59 BASIC BIOLOGICAL SCIENCES

Genes encoding calmodulin-binding proteins in the Arabidopsis genome

Analysis of the recently completed Arabidopsis genome sequence indicates that approximately 31% of the predicted genes could not be assigned to functional categories, as they do not show any sequence similarity with proteins of known function from other organisms. Calmodulin (CaM), a ubiquitous and multifunctional Ca(2+) sensor, interacts with a wide variety of cellular proteins and modulates their activity/function in regulating diverse cellular processes. However, the primary amino acid sequence of the CaM-binding domain in different CaM-binding proteins (CBPs) is not conserved. One way to identify most of the CBPs in the Arabidopsis genome is by protein-protein interaction-based screening of expression libraries with CaM. Here, using a mixture of radiolabeled CaM isoforms from Arabidopsis, we screened several expression libraries prepared from flower meristem, seedlings, or tissues treated with hormones, an elicitor, or a pathogen. Sequence analysis of 77 positive clones that interact with CaM in a Ca(2+)-dependent manner revealed 20 CBPs, including 14 previously unknown CBPs. In addition, by searching the Arabidopsis genome sequence with the newly identified and known plant or animal CBPs, we identified a total of 27 CBPs. Among these, 16 CBPs are represented by families with 2-20 members in each family. Gene expression analysis revealed that CBPs and CBP paralogs are expressed differentially. Our data suggest that Arabidopsis has a large number of CBPs including several plant-specific ones. Although CaM is highly conserved between plants and animals, only a few CBPs are common to both plants and animals. Analysis of Arabidopsis CBPs revealed the presence of a variety of interesting domains. Our analyses identified several hypothetical proteins in the Arabidopsis genome as CaM targets, suggesting their involvement in Ca(2+)-mediated signaling networks.

NASA Discipline Plant Biology

Cholesterol modulates membrane elasticity via unified biophysical laws

Cholesterol and lipid unsaturation underlie a balance of opposing forces that features prominently in adaptive cell responses to diet and environmental cues. These competing factors have resulted in contradictory observations of membrane elasticity across different measurement scales, requiring chemical specificity to explain incompatible structural and elastic effects. Here, we demonstrate that – unlike macroscopic observations – lipid membranes exhibit a unified elastic behavior in the mesoscopic regime between molecular and macroscopic dimensions. Using nuclear spin techniques and computational analysis, we find that mesoscopic bending moduli follow a universal dependence on the lipid packing density regardless of cholesterol content, lipid unsaturation, or temperature. Our observations reveal that compositional complexity can be explained by simple biophysical laws that directly map membrane elasticity to molecular packing associated with biological function, curvature transformations, and protein interactions. The obtained scaling laws closely align with theoretical predictions based on conformational chain entropy and elastic stress fields. These findings provide unique insights into the membrane design rules optimized by nature and unlock predictive capabilities for guiding the functional performance of lipid-based materials in synthetic biology and real-world applications.

Kumarage, Teshani [Virginia Polytechnic Inst. and

Shedding Light on Microbial Dark Matter with A Universal Language of Life

The majority of microbial genomes have yet to be cultured, and most proteins predicted from microbial genomes or sequenced from the environment cannot be functionally annotated. As a result, current computational approaches to describe microbial systems rely on incomplete reference databases that cannot adequately capture the full functional diversity of the microbial tree of life, limiting our ability to model high-level features of biological sequences. The scientific community needs a means to capture the functionally and evolutionarily relevant features underlying biology, independent of our incomplete reference databases. Such a model can form the basis for transfer learning tasks, enabling downstream applications in environmental microbiology, medicine, and bioengineering. Here we present LookingGlass, a deep learning model capturing a “universal language of life”. LookingGlass encodes contextually-aware, functionally and evolutionarily relevant representations of short DNA reads, distinguishing reads of disparate function, homology, and environmental origin. We demonstrate the ability of LookingGlass to be fine-tuned to perform a range of diverse tasks: to identify novel oxidoreductases, to predict enzyme optimal temperature, and to recognize the reading frames of DNA sequence fragments. LookingGlass is the first contextually-aware, general purpose pre-trained “biological language” representation model for short-read DNA sequences. LookingGlass enables functionally relevant representations of otherwise unknown and unannotated sequences, shedding light on the microbial dark matter that dominates life on Earth.

A Hoarfrost

High-throughput protein characterization by complementation using DNA barcoded fragment libraries

Abstract Our ability to predict, control, or design biological function is fundamentally limited by poorly annotated gene function. This can be particularly challenging in non-model systems. Accordingly, there is motivation for new high-throughput methods for accurate functional annotation. Here, we used co mplementation of aux otrophs and DNA barcode seq uencing (Coaux-Seq) to enable high-throughput characterization of protein function. Fragment libraries from eleven genetically diverse bacteria were tested in twenty different auxotrophic strains of Escherichia coli to identify genes that complement missing biochemical activity. We recovered 41% of expected hits, with effectiveness ranging per source genome, and observed success even with distant E. coli relatives like Bacillus subtilis and Bacteroides thetaiotaomicron . Coaux-Seq provided the first experimental validation for 53 proteins, of which 11 are less than 40% identical to an experimentally characterized protein. Among the unexpected function identified was a sulfate uptake transporter, an O-succinylhomoserine sulfhydrylase for methionine synthesis, and an aminotransferase. We also identified instances of cross-feeding wherein protein overexpression and nearby non-auxotrophic strains enabled growth. Altogether, Coaux-Seq’s utility is demonstrated, with future applications in ecology, health, and engineering.

59 BASIC BIOLOGICAL SCIENCES

Comparison of Two Bioinformatics Tools Used to Characterize the Microbial Diversity and Predictive Functional Attributes of Microbial Mats from Lake Obersee, Antarctica

In this study, using NextGen sequencing of the collective 16S rRNA genes obtained from two sets of samples collected from Lake Obersee, Antarctica, we compared and contrasted two bioinformatics tools, PICRUSt and Tax4Fun. We then developed an R script to assess the taxonomic and predictive functional profiles of the microbial communities within the samples. Taxa such as Pseudoxanthomonas, Planctomycetaceae, Cyanobacteria Subsection III, Nitrosomonadaceae, Leptothrix, and Rhodobacter were exclusively identified by Tax4Fun that uses SILVA database; whereas PICRUSt that uses Greengenes database uniquely identified Pirellulaceae, Gemmatimonadetes A1-B1, Pseudanabaena, Salinibacterium and Sinibacteraceae. Predictive functional profiling of the microbial communities using Tax4Fun and PICRUSt separately revealed common metabolic capabilities, while also showing specific functional IDs not shared between the two approaches. Combining these functional predictions using a customized R script revealed a more inclusive metabolic profile, such as hydrolases, oxidoreductases, transferases; enzymes involved in carbohydrate and amino acid metabolisms; and membrane transport proteins known for nutrient uptake from the surrounding environment. Our results present the first molecular-phylogenetic characterization and predictive functional profiles of the microbial mat communities in Lake Obersee, while demonstrating the efficacy of combining both the taxonomic assignment information and functional IDs using the R script created in this study for a more streamlined evaluation of predictive functional profiles of microbial communities.

Hyunmin Koo

A non-canonical fungal peroxisome PTS-1 signal, SYM, and its evolutionary aspects

Abstract Proteins localized to peroxisomes, particularly those expressed under specific conditions or in low abundance, are often undetected by routine proteomics methods due to detection sensitivity limits. In silico identification and experimental validation of peroxisomal targeting signals (PTSs) offer a reliable alternative. We demonstrate that SYM, a non-canonical plant PTS-1 signal, functions similarly inAspergillus nidulans, as GFP tagged with a SYM C-terminal tripeptide localizes to peroxisomes. One of two nativeA. nidulansproteins with C-terminal SYM tripeptide shows weak peroxisomal localization alongside cytoplasmic presence, indicating that only a subset of proteins with non-canonical signals access peroxisomes.In silicoanalysis of 1,010 fungal genomes identified diverse SYM-proteins with variable functions, suggesting that non-canonical PTS-1 signals may evolve spontaneously. Two-thirds of SYM-proteins are predicted to localize to specific intracellular compartments other than the peroxisome. We propose that despite their predicted localization, these proteins possessing SYM as a non-canonical peroxisomal signal might also have peroxisomal presence. Among SYM-proteins, pectinesterases, known plant pathogen virulence factors, were frequent. Notably, 25% of fungal pectinesterases harbor non-canonical PTS-1 signals, suggesting that partial peroxisomal localization of pectinesterases has evolved convergently. This suggests that partial peroxisomal localization may enhance protein functional flexibility, contributing to the organism’s adaptability.

Science & Technology - Other Topics

Knowledge Graph of RB-Tnseq Data from Fitness Browser (KP-DP1)

Motivation: Predicting microbial gene fitness across environmental conditions remains a central challenge for predictive phenomics and autonomous experimentation. Fitness assays generate large volumes of genotype–phenotype measurements difficult to integrate with experimental metadata and biological function in a form that supports mechanistic reasoning. Knowledge graphs offer a semantic framework for unifying modalities and enabling context-aware inference. Results: We build GIMME (Graph Inference for Microbial Metabolism Exploration), a semantically grounded knowledge graph that unifies gene fitness measurements spanning 10 Pseudomonas species with experimental metadata and biological context. Media are decomposed into chemical components and experiments carry structured links to natural-language descriptions. The resulting graph supports two inference modes: (1) symbolic graph traversal to surface candidate gene–environment and gene–chemical associations, and (2) learned inference using heterogeneous graph neural networks that propagate information across neighborhoods. We formulate link regression over (gene, media, experiment) triplets, combining learned gene embeddings with pretrained LLM sourced text embeddings of node descriptions to predict gene fitness. We then augment a baseline MLP with an auxiliary message-passing encoder (GraphSAGE/GAT) that propagates information over gene–protein–function and media–chemical subgraphs, and fuse the two pathways with a gated residual connection. This approach produces strong agreement with held-out fitness measurements (GraphSAGE Pearson r 0.74) while also highlighting inference challenges in extreme-fitness regimes. We aggregate GAT edge-attention weights by relation type and layer to estimate which biological and environmental relations most influence fitness predictions. Conclusion: This work explores using knowledge graphs as “context graphs” for microbial phenotype prediction. They provide a rich substrate which enables explainable retrieval of supporting evidence, and provides a natural bridge to autonomous workflows that prioritize the next experiment.

59 BASIC BIOLOGICAL SCIENCES

Spatial top-down proteomics for the functional characterization of human kidney

Background: The Human Proteome Project has credibly detected nearly 93% of the roughly 20,000 proteins which are predicted by the human genome. However, the proteome is enigmatic, where alterations in amino acid sequences from polymorphisms and alternative splicing, errors in translation, and post-translational modifications result in a proteome depth estimated at several million unique proteoforms. Recently mass spectrometry has been demonstrated in several landmark efforts mapping the human proteoform landscape in bulk analyses. Herein, we developed an integrated workflow for characterizing proteoforms from human tissue in a spatially resolved manner by coupling laser capture microdissection, nanoliter-scale sample preparation, and mass spectrometry imaging. Results: Using healthy human kidney sections as the case study, we focused our analyses on the major functional tissue units including glomeruli, tubules, and medullary rays. After laser capture microdissection, these isolated functional tissue units were processed with microPOTS (microdroplet processing in one-pot for trace samples) for sensitive top-down proteomics measurement. This provided a quantitative database of 616 proteoforms that was further leveraged as a library for mass spectrometry imaging with near-cellular spatial resolution over the entire section. Notably, several mitochondrial proteoforms were found to be differentially abundant between glomeruli and convoluted tubules, and further spatial contextualization was provided by mass spectrometry imaging confirming unique differences identified by microPOTS, and further expanding the field-of-view for unique distributions such as enhanced abundance of a truncated form (1-74) of ubiquitin within cortical regions. Conclusions: We developed an integrated workflow to directly identify proteoforms and reveal their spatial distributions. Where of the 20 differentially abundant proteoforms identified as discriminate between tubules and glomeruli by microPOTS, the vast majority of tubular proteoforms were of mitochondrial origin (8 of 10) where discriminate proteoforms in glomeruli were primarily hemoglobin subunits (9 of 10). These trends were also identified within ion images demonstrating spatially resolved characterization of proteoforms that has the potential to reshape discovery-based proteomics because the proteoforms are the ultimate effector of cellular functions. Applications of this technology have the potential to unravel etiology and pathophysiology of disease states, informing on biologically active proteoforms, which remodel the proteomic landscape in chronic and acute disorders.

59 BASIC BIOLOGICAL SCIENCES

Genetic Transfer in Action: Uncovering DNA Flow in an Extremophilic Microbial Community

ABSTRACT Horizontal genetic transfer (HGT) is a significant driver of genomic novelty in all domains of life. HGT has been investigated in many studies however, the focus has been on conspicuous protein‐coding DNA transfers that often prove to be adaptive in recipient organisms and are therefore fixed longer‐term in lineages. These results comprise a subclass of HGTs and do not represent exhaustive (coding and non‐coding) DNA transfer and its impact on ecology. Uncovering exhaustive HGT can provide key insights into the connectivity of genomes in communities and how these transfers may occur. In this study, we use the term frequency‐inverse document frequency (TF‐IDF) technique, that has been used successfully to mine DNA transfers within real and simulated high‐quality prokaryote genomes, to search for exhaustive HGTs within an extremophilic microbial community. We establish a pipeline for validating transfers identified using this approach. We find that most DNA transfers are within‐domain and involve non‐coding DNA. A relatively high proportion of the predicted protein‐coding HGTs appear to encode transposase activity, restriction‐modification system components, and biofilm formation functions. Our study demonstrates the utility of the TF‐IDF approach for HGT detection and provides insights into the mechanisms of recent DNA transfer.

Microbiology

African Swine Fever Virus Protein–Protein Interaction Prediction

The African swine fever virus (ASFV) is an often deadly disease in swine and poses a threat to swine livestock and swine producers. With its complex genome containing more than 150 coding regions, developing effective vaccines for this virus remains a challenge due to a lack of basic knowledge about viral protein function and protein–protein interactions between viral proteins and between viral and host proteins. In this work, we identified ASFV-ASFV protein–protein interactions (PPIs) using artificial intelligence-powered protein structure prediction tools. We benchmarked our PPI identification workflow on the Vaccinia virus, a widely studied nucleocytoplasmic large DNA virus, and found that it could identify gold-standard PPIs that have been validated in vitro in a genome-wide computational screening. We applied this workflow to more than 18,000 pairwise combinations of ASFV proteins and were able to identify seventeen novel PPIs, many of which have corroborating experimental or bioinformatic evidence for their protein–protein interactions, further validating their relevance. Two protein–protein interactions, I267L and I8L, I267L__I8L, and B175L and DP79L, B175L__DP79L, are novel PPIs involving viral proteins known to modulate host immune response.

59 BASIC BIOLOGICAL SCIENCES

metagRoot: a comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71 091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, hidden Markov models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org.

Chasapi, Maria N