Search NASA⌕ Search

SEARCH · Search NASA

Results for “Amino acid encoding”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Amino Acid Encoding for Deep Learning Applications

Background: The number of applications of deep learning algorithms in bioinformatics is increasing as they usually achieve superior performance over classical approaches, especially, when bigger training datasets are available. In deep learning applications, discrete data, e.g. words or n-grams in language, or amino acids or nucleotides in bioinformatics, are generally represented as a continuous vector through an embedding matrix. Recently, learning this embedding matrix directly from the data as part of the continuous iteration of the model to optimize the target prediction – a process called ‘end-to-end learning’ – has led to state-of-the-art results in many fields. Although usage of embeddings is well described in the bioinformatics literature, the potential of end-to-end learning for single amino acids, as compared to more classical manually-curated encoding strategies, has not been systematically addressed. To this end, we compared classical encoding matrices, namely one-hot, VHSE8 and BLOSUM62, to end-to-end learning of amino acid embeddings for two different prediction tasks using three widely used architectures, namely recurrent neural networks (RNN), convolutional neural networks (CNN), and the hybrid CNN-RNN. Results: By using different deep learning architectures, we show that end-to-end learning is on par with classical encodings for embeddings of the same dimension even when limited training data is available, and might allow for a reduction in the embedding dimension without performance loss, which is critical when deploying the models to devices with limited computational capacities. We found that the embedding dimension is a major factor in controlling the model performance. Surprisingly, we observed that deep learning models are capable of learning from random vectors of appropriate dimension. Conclusion: Our study shows that end-to-end learning is a flexible and powerful method for amino acid encoding. Further, due to the flexibility of deep learning systems, amino acid encoding schemes should be benchmarked against random vectors of the same dimension to disentangle the information content provided by the encoding scheme from the distinguishability effect provided by the scheme.

Deep-learning↗

Amino Acid Encoding for Deep Learning Applications

Background: The number of applications of deep learning algorithms in bioinformatics is increasing as they usually achieve superior performance over classical approaches, especially, when bigger training datasets are available. In deep learning applications, discrete data, e.g. words or n-grams in language, or amino acids or nucleotides in bioinformatics, are generally represented as a continuous vector through an embedding matrix. Recently, learning this embedding matrix directly from the data as part of the continuous iteration of the model to optimize the target prediction – a process called ‘end-to-end learning’ – has led to state-ofthe-art results in many fields. Although usage of embeddings is well described in the bioinformatics literature, the potential of end-to-end learning for single amino acids, as compared to more classical manually-curated encoding strategies, has not been systematically addressed. To this end, we compared classical encoding matrices, namely one-hot, VHSE8 and BLOSUM62, to end-to-end learning of amino acid embeddings for two different prediction tasks using three widely used architectures, namely recurrent neural networks (RNN), convolutional neural networks (CNN), and the hybrid CNN-RNN. Results: By using different deep learning architectures, we show that end-to-end learning is on par with classical encodings for embeddings of the same dimension even when limited training data is available, and might allow for a reduction in the embedding dimension without performance loss, which is critical when deploying the models to devices with limited computational capacities. We found that the embedding dimension is a major factor in controlling the model performance. Surprisingly, we observed that deep learning models are capable of learning from random vectors of appropriate dimension. Conclusion: Our study shows that end-to-end learning is a flexible and powerful method for amino acid encoding. Further, due to the flexibility of deep learning systems, amino acid encoding schemes should be benchmarked against random vectors of the same dimension to disentangle the information content provided by the encoding scheme from the distinguishability effect provided by the scheme.

Hesham ElAbd↗

Molecular cloning and characterization of a tumor-associated, growth-related, and time-keeping hydroquinone (NADH) oxidase (tNOX) of the HeLa cell surface

NOX proteins are growth-related cell surface proteins that catalyze both hydroquinone or NADH oxidation and protein disulfide interchange and exhibit prion-like properties. The two enzymatic activities alternate to generate a regular period length of about 24 min. Here we report the expression, cloning, and characterization of a tumor-associated NADH oxidase (tNOX). The cDNA sequence of 1830 bp is located on gene Xq25-26 with an open reading frame encoding 610 amino acids. The activities of the bacterially expressed tNOX oscillate with a period length of 22 min as is characteristic of tNOX activities in situ. The activities are inhibited completely by capsaicin, which represents a defining characteristic of tNOX activity. Functional motifs identified by site-directed mutagenesis within the C-terminal portion of the tNOX protein corresponding to the processed plasma membrane-associated form include quinone (capsaicin), copper and adenine nucleotide binding domains, and two cysteines essential for catalytic activity. Four of the six cysteine to alanine replacements retained enzymatic activity, but the period lengths of the oscillations were increased. A single protein with two alternating enzymatic activities indicative of a time-keeping function is unprecedented in the biochemical literature.

NASA Discipline Cell Biology↗

A Closer Look at Non-Random Patterns Within Chemistry Space for a Smaller, Earlier Amino Acid Alphabet

Recent findings, in vitro and in silico, are strengthening the idea of a simpler, earlier stage of genetically encoded proteins which used amino acids produced by prebiotic chemistry. These findings motivate a re-examination of prior work which has identified unusual properties of the set of twenty amino acids found within the full genetic code, while leaving it unclear whether similar patterns also characterize the subset of prebiotically plausible amino acids. We have suggested previously that this ambiguity may result from the low number of amino acids recognized by the definition of prebiotic plausibility used for the analysis. Here, we test this hypothesis using significantly updated data for organic material detected within meteorites, which contain several coded and non-coded amino acids absent from prior studies. In addition to confirming the well-established idea that “late” arriving amino acids expanded the chemistry space encoded by genetic material, we find that a prebiotically plausible subset of coded amino acids generally emulates the patterns found in the full set of 20, namely an exceptionally broad and even distribution of volumes and an exceptionally even distribution of hydrophobicities (quantified as logP) over a narrow range. However, the strength of this pattern varies depending on both the size and composition the library used to create a background (null model) for a random alphabet, and the precise definition of exactly which amino acids were present in a simpler, earlier code. Findings support the idea that a small sample size of amino acids caused previous ambiguous results, and further improvements in meteorite analysis, and/or prebiotic simulations will further clarify the nature and extent of unusual properties. We discuss the case of sulfur-containing amino acids as a specific and clear example and conclude by reviewing the potential impact of better understanding the chemical “logic” of a smaller forerunner to the standard amino acid alphabet.

Amino acids↗

A mutation in the Arabidopsis HYL1 gene encoding a dsRNA binding protein affects responses to abscisic acid, auxin, and cytokinin

Both physiological and genetic evidence indicate interconnections among plant responses to different hormones. We describe a pleiotropic recessive Arabidopsis transposon insertion mutation, designated hyponastic leaves (hyl1), that alters the plant's responses to several hormones. The mutant is characterized by shorter stature, delayed flowering, leaf hyponasty, reduced fertility, decreased rate of root growth, and an altered root gravitropic response. It also exhibits less sensitivity to auxin and cytokinin and hypersensitivity to abscisic acid (ABA). The auxin transport inhibitor 2,3,5-triiodobenzoic acid normalizes the mutant phenotype somewhat, whereas another auxin transport inhibitor, N-(1-naph-thyl)phthalamic acid, exacerbates the phenotype. The gene, designated HYL1, encodes a 419-amino acid protein that contains two double-stranded RNA (dsRNA) binding motifs, a nuclear localization motif, and a C-terminal repeat structure suggestive of a protein-protein interaction domain. We present evidence that the HYL1 gene is ABA-regulated and encodes a nuclear dsRNA binding protein. We hypothesize that the HYL1 protein is a regulatory protein functioning at the transcriptional or post-transcriptional level.

NASA Discipline Plant Biology↗

Characterization and distribution of a maize cDNA encoding a peptide similar to the catalytic region of second messenger dependent protein kinases

Maize (Zea mays) roots respond to a variety of environmental stimuli which are perceived by a specialized group of cells, the root cap. We are studying the transduction of extracellular signals by roots, particularly the role of protein kinases. Protein phosphorylation by kinases is an important step in many eukaryotic signal transduction pathways. As a first phase of this research we have isolated a cDNA encoding a maize protein similar to fungal and animal protein kinases known to be involved in the transduction of extracellular signals. The deduced sequence of this cDNA encodes a polypeptide containing amino acids corresponding to 33 out of 34 invariant or nearly invariant sequence features characteristic of protein kinase catalytic domains. The maize cDNA gene product is more closely related to the branch of serine/threonine protein kinase catalytic domains composed of the cyclic-nucleotide- and calcium-phospholipid-dependent subfamilies than to other protein kinases. Sequence identity is 35% or more between the deduced maize polypeptide and all members of this branch. The high structural similarity strongly suggests that catalytic activity of the encoded maize protein kinase may be regulated by second messengers, like that of all members of this branch whose regulation has been characterized. Northern hybridization with the maize cDNA clone shows a single 2400 base transcript at roughly similar levels in maize coleoptiles, root meristems, and the zone of root elongation, but the transcript is less abundant in mature leaves. In situ hybridization confirms the presence of the transcript in all regions of primary maize root tissue.

NASA Program Space Biology↗

Identification of a new EF-hand superfamily member from Trypanosoma brucei

We identified several open reading frames between the regions encoding calmodulin and ubiquitin-EP52/1 in the genome of Trypanosoma brucei. One of these, EFH5, encodes a protein 192 amino acids long. The EFH5 transcript is present in poly(A)+ mRNA and is present at similar levels in the mammalian bloodstream form and the insect procyclic form. EFH5 contains four EF-hand homolog domains, two of which are inferred to bind Ca2+ ions. We expressed EFH5 as a fusion protein in Escherichia coli and demonstrated calcium-binding activity of the fusion protein using the 45Ca-overlay technique. The function of EFH5 remains unknown; however, as the fourth EF-hand homolog identified in trypanosomes, it attests to the broad range of functions assumed by calcium functioning as a second messenger. EFH5, which is most closely related to LAV1-2 from Physarum, represents a distinct subfamily among the EF-hand-containing proteins.

Non-NASA Center↗

Isolation and characterization of a novel calmodulin-binding protein from potato

Tuberization in potato is controlled by hormonal and environmental signals. Ca(2+), an important intracellular messenger, and calmodulin (CaM), one of the primary Ca(2+) sensors, have been implicated in controlling diverse cellular processes in plants including tuberization. The regulation of cellular processes by CaM involves its interaction with other proteins. To understand the role of Ca(2+)/CaM in tuberization, we have screened an expression library prepared from developing tubers with biotinylated CaM. This screening resulted in isolation of a cDNA encoding a novel CaM-binding protein (potato calmodulin-binding protein (PCBP)). Ca(2+)-dependent binding of the cDNA-encoded protein to CaM is confirmed by (35)S-labeled CaM. The full-length cDNA is 5 kb long and encodes a protein of 1309 amino acids. The deduced amino acid sequence showed significant similarity with a hypothetical protein from another plant, Arabidopsis. However, no homologs of PCBP are found in nonplant systems, suggesting that it is likely to be specific to plants. Using truncated versions of the protein and a synthetic peptide in CaM binding assays we mapped the CaM-binding region to a 20-amino acid stretch (residues 1216-1237). The bacterially expressed protein containing the CaM-binding domain interacted with three CaM isoforms (CaM2, CaM4, and CaM6). PCBP is encoded by a single gene and is expressed differentially in the tissues tested. The expression of CaM, PCBP, and another CaM-binding protein is similar in different tissues and organs. The predicted protein contained seven putative nuclear localization signals and several strong PEST motifs. Fusion of the N-terminal region of the protein containing six of the seven nuclear localization signals to the reporter gene beta-glucuronidase targeted the reporter gene to the nucleus, suggesting a nuclear role for PCBP.

NASA Discipline Plant Biology↗

A Novel Kinesin-Like Protein with a Calmodulin-Binding Domain

Calcium regulates diverse developmental processes in plants through the action of calmodulin. A cDNA expression library from developing anthers of tobacco was screened with S-35-labeled calmodulin to isolate cDNAs encoding calmodulin-binding proteins. Among several clones isolated, a kinesin-like gene (TCK1) that encodes a calmodulin-binding kinesin-like protein was obtained. The TCK1 cDNA encodes a protein with 1265 amino acid residues. Its structural features are very similar to those of known kinesin heavy chains and kinesin-like proteins from plants and animals, with one distinct exception. Unlike other known kinesin-like proteins, TCK1 contains a calmodulin-binding domain which distinguishes it from all other known kinesin genes. Escherichia coli-expressed TCK1 binds calmodulin in a Ca(2+)-dependent manner. In addition to the presence of a calmodulin-binding domain at the carboxyl terminal, it also has a leucine zipper motif in the stalk region. The amino acid sequence at the carboxyl terminal of TCK1 has striking homology with the mechanochemical motor domain of kinesins. The motor domain has ATPase activity that is stimulated by microtubules. Southern blot analysis revealed that TCK1 is coded by a single gene. Expression studies indicated that TCKI is expressed in all of the tissues tested. Its expression is highest in the stigma and anther, especially during the early stages of anther development. Our results suggest that Ca(2+)/calmodulin may play an important role in the function of this microtubule-associated motor protein and may be involved in the regulation of microtubule-based intracellular transport.

Wang, W.↗

Cloning and Characterization of an alpha-amylase Gene from the Hyperthermophilic Archaeon Thermococcus Thioreducens

The gene encoding an extracellular alpha-amylase, TTA, from the hyperthermophilic archaeon Thermococcus thioreducens was cloned and expressed in Escherichia coli. Primary structural analysis revealed high similarity with other a-amylases from the Thermococcus and Pyrococcus genera, as well as the four highly conserved regions typical for a-amylases. The 1374 bp gene encodes a protein of 457 amino acids, of which 435 constitute the mature protein preceded by a 22 amino acid signal peptide. The molecular weight of the purified recombinant enzyme was estimated to be 43 kDa by denaturing gel electrophoresis. Maximal enzymatic activity of recombinant TTA was observed at 90 C and pH 5.5 in the absence of exogenous Ca(2+), and the enzyme was considerably stable even after incubation at 90 C for 2 hours. The thermostability at 90 and 102 C was enhanced in the presence of 5 mM Ca(2+). The extraordinarily high specific activity (about 7.4 x 10(exp 3) U/mg protein at 90 C, pH 5.5 with soluble starch as substrate) together with its low pH optimum makes this enzyme an interesting candidate for starch processing applications.

Bernhardsdotter, Eva C. M. J.↗

Cloning and Characterization of an Alpha-amylase Gene from the Hyperthermophilic Archaeon Thermococcus Thioreducens

The gene encoding an extracellular a-amylase, TTA, from the hyperthermophilic archaeon Thermococcus thioreducens was cloned and expressed in Escherichia coli. Primary structural analysis revealed high similarity with other a-amylases from the Thermococcus and Pyrococcus genera, as well as the four highly conserved regions typical for a-amylases. The 1374 bp gene encodes a protein of 457 amino acids, of which 435 constitute the mature protein preceded by a 22 amino acid signal peptide. The molecular weight of the purified recombinant enzyme was estimated to be 43 kDa by denaturing gel electrophoresis. Maximal enzymatic activity of recombinant TTA was observed at 90 C and pH 5.5 in the absence of exogenous Ca(2+), and the enzyme was considerably stable even after incubation at 90 C for 2 hours. The thermostability at 90 and 102 C was enhanced in the presence of 5 mM Ca(2+). The extraordinarily high specific activity (about 7.4 x 10(exp 3) U/mg protein at 90 C, pH 5.5 with soluble starch as substrate) together with its low pH optimum makes this enzyme an interesting candidate for starch processing applications.

Bernhardsdotter, Eva C. M. J.↗

Mutations in the putative calcium-binding domain of polyomavirus VP1 affect capsid assembly

Calcium ions appear to play a major role in maintaining the structural integrity of the polyomavirus and are likely involved in the processes of viral uncoating and assembly. Previous studies demonstrated that a VP1 fragment extending from Pro-232 to Asp-364 has calcium-binding capabilities. This fragment contains an amino acid stretch from Asp-266 to Glu-277 which is quite similar in sequence to the amino acids that make up the calcium-binding EF hand structures found in many proteins. To assess the contribution of this domain to polyomavirus structural integrity, the effects of mutations in this region were examined by transfecting mutated viral DNA into susceptible cells. Immunofluorescence studies indicated that although viral protein synthesis occurred normally, infective viral progeny were not produced in cells transfected with polyomavirus genomes encoding either a VP1 molecule lacking amino acids Thr-262 through Gly-276 or a VP1 molecule containing a mutation of Asp-266 to Ala. VP1 molecules containing the deletion mutation were unable to bind 45Ca in an in vitro assay. Upon expression in Escherichia coli and purification by immunoaffinity chromatography, wild-type VP1 was isolated as pentameric, capsomere-like structures which could be induced to form capsid-like structures upon addition of CaCl2, consistent with previous studies. However, although VP1 containing the point mutation was isolated as pentamers which were indistinguishable from wild-type VP1 pentamers, addition of CaCl2 did not result in their assembly into capsid-like structures. Immunogold labeling and electron microscopy studies of transfected mammalian cells provided in vivo evidence that a mutation in this region affects the process of viral assembly.

Non-NASA Center↗

PCR cloning and characterization of multiple ADP-glucose pyrophosphorylase cDNAs from tomato

Four ADP-glucose pyrophosphorylase (AGP) cDNAs were cloned from tomato fruit and leaves by the PCR techniques. Three of them (agp S1, agp S2, and agp S3) encode the large subunit of AGP, the fourth one (agp B) encodes the small subunit. The deduced amino acid sequences of the cDNAs show very high identities (96-98%) to the corresponding potato AGP isoforms, although there are major differences in tissue expression profiles. All four tomato AGP transcripts were detected in fruit and leaves; the predominant ones in fruit are agp B and agp S1, whereas in leaves they are agp B and agp S3. Genomic southern analysis suggests that the four AGP transcripts are encoded by distinct genes.

Non-NASA Center↗

The arabidopsis thaliana AGRAVITROPIC 1 gene encodes a component of the polar-auxin-transport efflux carrier

Auxins are plant hormones that mediate many aspects of plant growth and development. In higher plants, auxins are polarly transported from sites of synthesis in the shoot apex to their sites of action in the basal regions of shoots and in roots. Polar auxin transport is an important aspect of auxin functions and is mediated by cellular influx and efflux carriers. Little is known about the molecular identity of its regulatory component, the efflux carrier [Estelle, M. (1996) Current Biol. 6, 1589-1591]. Here we show that mutations in the Arabidopsis thaliana AGRAVITROPIC 1 (AGR1) gene involved in root gravitropism confer increased root-growth sensitivity to auxin and decreased sensitivity to ethylene and an auxin transport inhibitor, and cause retention of exogenously added auxin in root tip cells. We used positional cloning to show that AGR1 encodes a putative transmembrane protein whose amino acid sequence shares homologies with bacterial transporters. When expressed in Saccharomyces cerevisiae, AGR1 promotes an increased efflux of radiolabeled IAA from the cells and confers increased resistance to fluoro-IAA, a toxic IAA-derived compound. AGR1 transcripts were localized to the root distal elongation zone, a region undergoing a curvature response upon gravistimulation. We have identified several AGR1-related genes in Arabidopsis, suggesting a global role of this gene family in the control of auxin-regulated growth and developmental processes.

Non-NASA Center↗

The sequence, and its evolutionary implications, of a Thermococcus celer protein associated with transcription

Through random search, a gene from Thermococcus celer has been identified and sequenced that appears to encode a transcription-associated protein (110 amino acid residues). The sequence has clear homology to approximately the last half of an open reading frame reported previously for Sulfolobus acidocaldarius [Langer, D. & Zillig, W. (1993) Nucleic Acids Res. 21, 2251]. The protein translations of these two archaeal genes in turn are homologs of a small subunit found in eukaryotic RNA polymerase I (A12.2) and the counterpart of this from RNA polymerase II (B12.6). Homology is also seen with the eukaryotic transcription factor TFIIS, but it involves only the terminal 45 amino acids of the archaeal proteins. Evolutionary implications of these homologies are discussed.

Non-NASA Center↗

Overexpression of Human Bone Alkaline Phosphatase in Pichia Pastoris

The Pichiapastoris expression system was utilized to produce functionally active human bone alkaline phosphatase in gram quantities. Bone alkaline phosphatase is a key enzyme in bone formation and biomineralization, yet important questions about its structural chemistry and interactions with other cellular enzymes in mineralizing tissues remain unanswered. A soluble form of human bone alkaline phosphatase was constructed by deletion of the 25 amino acid hydrophobic C-terminal region of the encoding cDNA and inserted into the X-33 Pichiapastoris strain. An overexpression system was developed in shake flasks and converted to large-scale fermentation. Alkaline phosphatase was secreted into the medium to a level of 32mgAL when cultured in shake flasks. Enzyme activity was 12U/mg measured by a spectrophotometric assay. Fermentation yielded 880mgAL with enzymatic activity of 968U/mg. Gel electrophoresis analysis indicates that greater than 50% of the total protein in the fermentation is alkaline phosphatase. A purification scheme has been developed using ammonium sulfate precipitation followed by hydrophobic interaction chromatography. We are currently screening crystallization conditions of the purified recombinant protein for subsequent X-ray diffraction analyses. Structural data should provide additional information on the role of alkaline phosphatase in normal bone mineralization and in certain bone mineralization anomalies.

Karr, Laurel↗

Mutations in a new Arabidopsis cyclophilin disrupt its interaction with protein phosphatase 2A

The heterotrimeric protein phosphatase 2A (PP2A) is a component of multiple signaling pathways in eukaryotes. Disruption of PP2A activity in Arabidopsis is known to alter auxin transport and growth response pathways. We demonstrated that the regulatory subunit A of an Arabidopsis PP2A interacts with a novel cyclophilin, ROC7. The gene for this cyclophilin encodes a protein that contains a unique 30-amino acid extension at the N-terminus, which distinguishes the gene product from all previously identified Arabidopsis cyclophilins. Altered forms of ROC7 cyclophilin with mutations in the conserved DENFKL domain did not bind to PP2A. Unlike protein phosphatase 2B, PP2A activity in Arabidopsis extracts was not affected by the presence of the cyclophilin-binding molecule cyclosporin. The ROC7 transcript was expressed to high levels in all tissues tested. Expression of an ROC7 antisense transcript gave rise to increased root growth. These results indicate that cyclophilin may have a role in regulating PP2A activity, by a mechanism that differs from that employed for cyclophilin regulation of PP2B.

NASA Discipline Plant Biology↗

Chimeric calcium/calmodulin-dependent protein kinase in tobacco: differential regulation by calmodulin isoforms

cDNA clones of chimeric Ca2+/calmodulin-dependent protein kinase (CCaMK) from tobacco (TCCaMK-1 and TCCaMK-2) were isolated and characterized. The polypeptides encoded by TCCaMK-1 and TCCaMK-2 have 15 different amino acid substitutions, yet they both contain a total of 517 amino acids. Northern analysis revealed that CCaMK is expressed in a stage-specific manner during anther development. Messenger RNA was detected when tobacco bud sizes were between 0.5 cm and 1.0 cm. The appearance of mRNA coincided with meiosis and became undetectable at later stages of anther development. The reverse polymerase chain reaction (RT-PCR) amplification assay using isoform-specific primers showed that both of the CCaMK mRNAs were expressed in anther with similar expression patterns. The CCaMK protein expressed in Escherichia coli showed Ca2+-dependent autophosphorylation and Ca2+/calmodulin-dependent substrate phosphorylation. Calmodulin isoforms (PCM1 and PCM6) had differential effects on the regulation of autophosphorylation and substrate phosphorylation of tobacco CCaMK, but not lily CCaMK. The evolutionary tree of plant serine/threonine protein kinases revealed that calmodulin-dependent kinases form one subgroup that is distinctly different from Ca2+-dependent protein kinases (CDPKs) and other serine/threonine kinases in plants.

NASA Discipline Plant Biology↗