Search NASA⌕ Search

SEARCH · Search NASA

Results for “Protein Sequences”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A Multiplexed Quantitative Analysis of Germline Single Amino Acid Variants by Targeted Proteomics in Nondepleted Human Plasma

Single amino acid variants (SAAVs) in protein sequences are often a direct result of single-nucleotide polymorphisms (SNPs). Certain germline SAAVs have shown biological relevance in different disease conditions but lack precise quantification in circulation, which could hinder functional investigations and progress in biomarker development. Here, we have developed a multiplexed liquid chromatography-selected reaction monitoring (LC-SRM) assay that monitors 5 wild-type and variant peptide pairs (Complement Factor B: CFB-R32Q/R32W, Clusterin: CLU-N317H, Fetuin B: FETUB-K360R, and Kininogen: KNG1-L212P) in nondepleted human plasma. The assay was optimized for imprecision, linearity, stability, and calibration assessments with CVs of under 20%. The wild-type and variant peptide pairs were characterized in a set of healthy individual plasma samples. These target identifications were also validated by SNP genotyping with more than 99% accuracy. For all protein targets, we observed significantly lower concentrations of WT species in the presence variant peptides. In CFB, the concentration of R32Q was significantly lower than its counterpart R32W variant and WT species. Furthermore, our results distinguished phenotypes of homozygosity and heterozygosity of the SAAV presence through direct concentration level characterization. These findings provide some insights into how SAAVs affect quantitative assessments of target peptides. The assay demonstrates a platform for proteogenomic analyses with potential applications in both research and clinical settings.

genetics↗

UnigeneFinder: An Automated Pipeline for Gene Calling From Transcriptome Assemblies Without a Reference Genome

ABSTRACT For most species, transcriptome data are much more readily available than genome data. Without a reference genome, gene calling is cumbersome and inaccurate because of the high degree of redundancy in de novo transcriptome assemblies. To simplify and increase the accuracy of de novo transcriptome assembly in the absence of a reference genome, we developed UnigeneFinder. Combining several clustering methods, UnigeneFinder substantially reduces the redundancy typical of raw transcriptome assemblies. This pipeline offers an effective solution to the problem of inflated transcript numbers, achieving a closer representation of the actual underlying genome. UnigeneFinder performs comparably or better, compared with existing tools, on plant species with varying genome complexities. UnigeneFinder is the only available transcriptome redundancy solution that fully automates the generation of primary transcript, coding region, and protein sequences, analogous to those available for high‐quality reference genomes. These features, coupled with the pipeline’s cross‐platform implementation, focus on automation, and an accessible, user‐friendly interface, make UnigeneFinder a useful tool for many downstream sequence‐based analyses in nonmodel organisms lacking a reference genome, including differential gene expression analysis, accurate ortholog identification, functional enrichments, and evolutionary analyses. UnigeneFinder also runs efficiently both on high‐performance computing (HPC) systems and personal computers, further reducing barriers to use.

Xue, Bo [Plant Resilience Institute Michigan State↗

Different chemical scaffolds bind to L-phe site in Mycobacterium tuberculosis Phe-tRNA synthetase

Tuberculosis (TB), caused by Mycobacterium tuberculosis (Mt), is one of the deadliest infectious diseases. The rise of multidrug-resistant strains represents a major public health threat, requiring new therapeutic options. Bacterial aminoacyl-tRNA synthetases (aaRS) have been shown to be highly promising drug targets, including for TB treatment. These enzymes play an essential role in translating the DNA gene code into protein sequence by attaching specific amino acid to their cognate tRNAs. They have multiple binding sites that can be targeted for inhibitor discovery: amino acid binding pocket, ATP binding pocket, tRNA binding site and an editing domain. Recently we reported several high-resolution structures of M. tuberculosis phenylalanyl-tRNA synthetase (MtPheRS) complexed with tRNA Phe and either L-Phe or a nonhydrolyzable phenylalanine adenylate analog. Here, in this study, using Nucleic Magnetic Resonance (NMR) and Surface Plasmon Resonance (SPR) we identified fragments that bind to MtPheRS and we determined crystal structures of their complexes with MtPheRS/tRNA Phe . All the binders interact with the L-Phe amino acid binding site. The analysis of interactions of the new compounds combined with adenylate analog structure provides insights for the rational design of antituberculosis drugs. The 3 ' arm of the tRNA Phe in all the structures was disordered with exception of one complex with D-735 compound. In this structure the 3' CCA end of the acceptor stem is observed in the editing domain of MtPheRS providing insights regarding the post-transfer editing activity of class II aaRS.

Gade, Priyanka [Univ. of Chicago, IL (United State↗

Evolutionary trajectory of transcription factors and selection of targets for metabolic engineering

Transcription factors (TFs) provide potentially powerful tools for plant metabolic engineering as they often control multiple genes in a metabolic pathway. However, selecting the best TF for a particular pathway has been challenging, and the selection often relies significantly on phylogenetic relationships. Here, we offer examples where evolutionary relationships have facilitated the selection of the suitable TFs, alongside situations where such relationships are misleading from the perspective of metabolic engineering. We argue that the evolutionary trajectory of a particular TF might be a better indicator than protein sequence homology alone in helping decide the best targets for plant metabolic engineering efforts. This article is part of the theme issue ‘The evolution of plant metabolism’.

Life Sciences & Biomedicine - Other Topics↗

Exabiome: Advancing Microbial Science through Exascale Computing

The Exabiome project seeks to improve the understanding of microbiomes through the development of methods for accelerating metagenomic science using exascale computing. This article gives an overview of scientific impact of the three components of the project: metagenome assembly, protein family detection, and comparative analysis of metagenomes. Exabiome developed MetaHipMer, the only metagenome assembler capable of scaling to full exascale systems. MetaHipMer has enabled ground-breaking assemblies on the Frontier supercomputer, with many scientific benefits, such as the discovery of rare species and viral genomes. To investigate protein families, Exabiome developed two exascale tools, PASTIS and HipMCL. Together, these can utilize exascale resources to understand the functional diversity of billions of dark matter proteins and novel protein families. For comparative analysis, Exabiome developed kmerprof, a tool that can be used to compare huge metagenomes for many different scientific purposes, for example, grouping human microbiomes according to body location.

59 BASIC BIOLOGICAL SCIENCES↗

ProteinTuneRL

ProteinTuneRL is a framework designed to harness the power of reinforcement learning for advanced protein design. The project enables fine-tuning of generative models to explore and optimize protein sequences with tailored structural and functional properties.

Landajuela Larma, Mikel [Lawrence Livermore Nation↗

Activation Domain Hunter (ADhunter) v2.0

ADhunter is a software program that enables accurate identification and quantification of transcriptional activation domains. Unlike previous software, ADhunter uses protein representations from a pre-trained protein language model, model ensembling, and a training dataset from a diverse sampling of protein sequence space for state-of-the-art performance. These advantages enable improved perception of transcriptional activation domains across sequence space that can be used for mapping natural genetic circuits and engineering synthetic genetic circuits. In particular, ADhunter enables fine-tuned control of gene expression through synthetic transcription factors that can be used for complex control of cellular programs.

Waldburger, Lucas [Lawrence Berkeley National Labo↗

PRIME: Protein Representation Inference for Mutation Evaluation

Protein language machine learning models built upon existing ESM-2 model developed by Evolutionary Scale (evolutionaryscale.ai) and an in-house protein language model based on the BERT model developed by Google. The code also includes model training scripts and saved checkpoints from our own training using publicly available SARS-CoV-2 protein sequences.

Gibson, Kaetlyn [Los Alamos National Lab]↗

Biosynthesis of bioprivileged, linear molecules via novel carboligase reactions

Over the award period, we made progress on the three aims. We screened twenty-five carboligases for activity coupling twenty-one possible -keto acids (Aim 1). The carboligases were selected across a diverse set of protein sequences. Using Q-Exactive UHPLC-MS, we tested a total of 210 coupled products per enzyme and generated a dataset of 5250 enzyme-substrate activity relationships. We identified multiple enzymes that had activity for synthesizing suberic acid and heptanoic acid (Aim 2). We built a random forest model for predicting the activity of each enzyme toward substrates on which it was not tested using the data from Aim 1. Finally, we evaluated growth defects that occurred due to expression of different carboligases in E. coli (Aim 3). We were able to identify specific metabolites and putative pathways that, when supplemented in the media, recovered the growth defect associated with the presence of specific carboligases. We are in the process of publishing two manuscript describing the methods for high-throughput screening of enzyme promiscuity, using machine learning to predict activity on untested substrates, and enzyme activity data we collected. This project has produced enabling data for biosynthesis of a range of new-to-nature compounds to support biomanufacturing.

60 APPLIED LIFE SCIENCES↗

Rational Design of Lanmodulin Variants for Size-Based Selectivity of Individual Rare Earth Elements

Rare earth elements (REEs) are essential to modern technologies, yet their high physical and chemical similarity makes separation of individual REEs difficult and environmentally taxing. Metalloproteins offer a promising alternative for selective REE binding, as they tend to have high metal ion affinity and specificity. Lanmodulin (LanM), in particular, has arisen as a potential candidate for REE separation as it exhibits picomolar affinity for elements in the REE family. Prior work has shown that the single point mutation D9N can shift LanM’s preference away from lanthanides toward actinides, motivating efforts to tune selectivity of LanM through targeted mutagenesis. Here, we tested the hypothesis that introducing selective aspartic acid to glutamic acid substitutions in the metal coordinating EF hands of LanM would impose steric constraints that would drive LanM affinity away from larger ions, such as La3+, to smaller ions, such as Y3+. To test this hypothesis, a combination of computational and experimental approaches were employed to evaluate the signal mutations LanM D5E and LanM D3E and the double mutants LanM D1ED5E and LanM D3ED9E. Surprisingly, increasing the number of mutations within the metal center did not enhance affinity for smaller REEs, or decrease affinity for larger ions. Only the single point mutation LanM D5E weakened La3+ binding by one order of magnitude relative to LanM wild type (WT), and pairing it with a second mutation to produce LanM D1ED5E drove La3+ affinity to be stronger than that seen for LanM WT. The D3E mutation alone prevented proper expression and folding, but paring it with D9E to produce LanM D3ED9E rescued expression and yielded La3+ affinities comparable to LanM WT. All variants that expressed (LanM D5E, LanM D1ED5E, LanM D3ED9E) displayed Y3+ affinities comparable to LanM WT. Overall, these results highlight the tunability of LanM’s metal-binding environment but also expose current limitations in predicting structural responses to point mutations within a protein sequence. This work establishes a foundation that can be used for refining computational and experimental strategies to engineer metalloproteins with tailored REE selectivity.

Close, Emily [Pacific Northwest National Laborator↗

Evolution of ribonuclease in relation to polypeptide folding mechanisms.

Comparisons of the N-terminal region of pancreatic RNAase in seven species are presented, taking into account cow, bison, deer, rat, pig, kangaroo, and turtle. The available limited evidence on hypervariable regions indicates that there is still an evolutionary constraint on them. It is proposed that there is a selection pressure acting on all regions of a protein sequence in evolution. Mutations that tend to obstruct the folding process can lead to various intensities of selection pressure.

Barnard, E. A.↗

Origin of the Eumetazoa: testing ecological predictions of molecular clocks against the Proterozoic fossil record

Molecular clocks have the potential to shed light on the timing of early metazoan divergences, but differing algorithms and calibration points yield conspicuously discordant results. We argue here that competing molecular clock hypotheses should be testable in the fossil record, on the principle that fundamentally new grades of animal organization will have ecosystem-wide impacts. Using a set of seven nuclear-encoded protein sequences, we demonstrate the paraphyly of Porifera and calculate sponge/eumetazoan and cnidarian/bilaterian divergence times by using both distance [minimum evolution (ME)] and maximum likelihood (ML) molecular clocks; ME brackets the appearance of Eumetazoa between 634 and 604 Ma, whereas ML suggests it was between 867 and 748 Ma. Significantly, the ME, but not the ML, estimate is coincident with a major regime change in the Proterozoic acritarch record, including: (i) disappearance of low-diversity, evolutionarily static, pre-Ediacaran acanthomorphs; (ii) radiation of the high-diversity, short-lived Doushantuo-Pertatataka microbiota; and (iii) an order-of-magnitude increase in evolutionary turnover rate. We interpret this turnover as a consequence of the novel ecological challenges accompanying the evolution of the eumetazoan nervous system and gut. Thus, the more readily preserved microfossil record provides positive evidence for the absence of pre-Ediacaran eumetazoans and strongly supports the veracity, and therefore more general application, of the ME molecular clock.

NASA Discipline Evolutionary Biology↗

Automated Miniaturized Instrument for Space Biology Applications and the Monitoring of the Astronauts Health Onboard the ISS

Human space travelers experience a unique environment that affects homeostasis and physiologic adaptation. The spacecraft environment subjects the traveler to noise, chemical and microbiological contaminants, increased radiation, and variable gravity forces. As humans prepare for long-duration missions to the International Space Station (ISS) and beyond, effective measures must be developed, verified and implemented to ensure mission success. Limited biomedical quantitative capabilities are currently available onboard the ISS. Therefore, the development of versatile instruments to perform space biological analysis and to monitor astronauts' health is needed. We are developing a fully automated, miniaturized system for measuring gene expression on small spacecraft in order to better understand the influence of the space environment on biological systems. This low-cost, low-power, multi-purpose instrument represents a major scientific and technological advancement by providing data on cellular metabolism and regulation. The current system will support growth of microorganisms, extract and purify the RNA, hybridize it to the array, read the expression levels of a large number of genes by microarray analysis, and transmit the measurements to Earth. The system will help discover how bacteria develop resistance to antibiotics and how pathogenic bacteria sometimes increase their virulence in space, facilitating the development of adequate countermeasures to decrease risks associated with human spaceflight. The current stand-alone technology could be used as an integrated platform onboard the ISS to perform similar genetic analyses on any biological systems from the tree of life. Additionally, with some modification the system could be implemented to perform real-time in-situ microbial monitoring of the ISS environment (air, surface and water samples) and the astronaut's microbiome using 16SrRNA microarray technology. Furthermore, the current system can be enhanced substantially by combining it with other technologies for automated, miniaturized, high-throughput biological measurements, such as fast sequencing, protein identification (proteomics) and metabolite profiling (metabolomics). Thus, the system can be integrated with other biomedical instruments in order to support and enhance telemedicine capability onboard ISS. NASA's mission includes sustained investment in critical research leading to effective countermeasures to minimize the risks associated with human spaceflight, and the use of appropriate technology to sustain space exploration at reasonable cost. Our integrated microarray technology is expected to fulfill these two critical requirements and to enable the scientific community to better understand and monitor the effects of the space environment on microorganisms and on the astronaut, in the process leveraging current capabilities and overcoming present limitations.

Human space travelers↗

Ah, Sweet Mystery of Life, OSIRIS-REx May Find You

The nature of the origin of life is a topic that has engaged people since ancient times. Where did we come from? What was the first life? How are we related? Are we alone? The study of biologic remains and environments preserved in rocks (fossils) and biochemical pathways and structures found across organisms (molecular fossils) can address these questions. Molecular evidence shows that all life on Earth is related fundamentally, biology shares a genetic language, related molecular machinery, and common chemistry By looking at the details of genetic and protein sequences more detailed relationships can be determined for modern organisms.

origin of life↗

Evolution of proteins.

The amino acid sequences of proteins from living organisms are dealt with. The structure of proteins is first discussed; the variation in this structure from one biological group to another is illustrated by the first halves of the sequences of cytochrome c, and a phylogenetic tree is derived from the cytochrome c data. The relative geological times associated with the events of this tree are discussed. Errors which occur in the duplication of cells during the evolutionary process are examined. Particular attention is given to evolution of mutant proteins, globins, ferredoxin, and transfer ribonucleic acids (tRNA's). Finally, a general outline of biological evolution is presented.

Dayhoff, M. O.↗

Evolutionary connections of biological kingdoms based on protein and nucleic acid sequence evidence

Prokaryotic and eukaryotic evolutionary trees are developed from protein and nucleic-acid sequences by the methods of numerical taxonomy. Trees are presented for bacterial ferredoxins, 5S ribosomal RNA, c-type cytochromes , cytochromes c2 and c', and 5.8S ribosomal RNA; the implications for early evolution are discussed; and a composite tree showing the branching of the anaerobes, aerobes, archaebacteria, and eukaryotes is shown. Single lines are found for all oxygen-evolving photosynthetic forms and for the salt-loving and high-temperature forms of archaebacteria. It is argued that the eukaryote mitochondria, chloroplasts, and cytoplasmic host material are descended from free-living prokaryotes that formed symbiotic associations, with more than one symbiotic event involved in the evolution of each organelle.

Dayhoff, M. O.↗

Library Screening, In Vivo Confirmation, and Structural and Bioinformatic Analysis of Pentapeptide Sequences as Substrates for Protein Farnesyltransferase

Protein farnesylation is a post-translational modification where a 15-carbon farnesyl isoprenoid is appended to the C-terminal end of a protein by farnesyltransferase (FTase). This process often causes proteins to associate with the membrane and participate in signal transduction pathways. The most common substrates of FTase are proteins that have C-terminal tetrapeptide CaaX box sequences where the cysteine is the site of modification. However, recent work has shown that five amino acid sequences can also be recognized, including the pentapeptides CMIIM and CSLMQ. In this work, peptide libraries were initially used to systematically vary the residues in those two parental sequences using an assay based on Matrix Assisted Laser Desorption Ionization–Mass Spectrometry (MALDI-MS). In addition, 192 pentapeptide sequences from the human proteome were screened using that assay to discover additional extended CaaaX-box motifs. Selected hits from that screening effort were rescreened using an in vivo yeast reporter protein assay. The X-ray crystal structure of CMIIM bound to FTase was also solved, showing that the C-terminal tripeptide of that sequence interacted with the enzyme in a similar manner as the C-terminal tripeptide of CVVM, suggesting that the tripeptide comprises a common structural element for substrate recognition in both tetrapeptide and pentapeptide sequences. Molecular dynamics simulation of CMIIM bound to FTase further shed light on the molecular interactions involved, showing that a putative catalytically competent Zn(II)-thiolate species was able to form. Bioinformatic predictions of tetrapeptide (CaaX-box) reactivity correlated well with the reactivity of pentapeptides obtained from in vivo analysis, reinforcing the importance of the C-terminal tripeptide motif. This analysis provides a structural framework for understanding the reactivity of extended CaaaX-box motifs and a method that may be useful for predicting the reactivity of additional FTase substrates bearing CaaaX-box sequences.

59 BASIC BIOLOGICAL SCIENCES↗

Repetitive proteins that undergo large conformational changes evade structural prediction algorithms

Protein structure prediction algorithms, such as AlphaFold, have accelerated protein design and advanced the understanding of the relationship between amino acid sequence and protein structure. However, these algorithms are limited in their ability to predict the structures of conformationally dynamic, intrinsically disordered, and stimuli-responsive proteins. To evaluate sequence-to-structure predictions of such challenging proteins, we explored a class of conformationally dynamic, repeats-in-toxin (RTX) proteins. RTX proteins adopt intrinsically disordered conformations in the absence of calcium and undergo reversible folding into β-roll structures upon binding to calcium. RTX proteins are characterized by tandem repeats of the sequence GGXGXDXUX, in which X can be any amino acid and U is an aliphatic amino acid. We designed RTX sequence variants with global substitutions of nonconserved amino acids, tandem repeats of consensus sequences GGAGXDTLY, and tandem repeats of scrambled sequences GGAGXDTYL. AlphaFold2 and AlphaFold3 predicted that all of these RTX variants adopt β-roll structures, characteristic of wild-type RTX bound to calcium. However, modeling the predicted structures with molecular dynamics simulations and characterizing the protein variants with circular dichroism spectroscopy, small-angle x-ray scattering, and x-ray crystallography revealed that variants adopt diverse, sequence-dependent structures in the absence and presence of calcium. To better design proteins for applications in biotechnology and sustainability, it is critical to build predictive tools that consider intrinsically disordered protein states and validate these tools with multi-mode, multi-scale experimental data.

Chang, Marina P. [Stanford Univ., CA (United State↗