Search NASA⌕ Search

SEARCH · Search NASA

Results for “Protein Sequences”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Technical advance: stringent control of transgene expression in Arabidopsis thaliana using the Top10 promoter system

We show that the tightly regulated tetracycline-sensitive Top10 promoter system (Weinmann et al. Plant J. 1994, 5, 559-569) is functional in Arabidopsis thaliana. A pure breeding A. thaliana line (JL-tTA/8) was generated which expressed a chimeric fusion of the tetracycline repressor and the activation domain of Herpes simplex virus (tTA), from a single transgenic locus. Plants from this line were crossed with transgenics carrying the ER-targeted green fluorescent protein coding sequence (mGFP5) under control of the Top10 promoter sequence. Progeny from this cross displayed ER-targeted GFP fluorescence throughout the plant, indicating that the tTA-Top10 promoter interaction was functional in A. thaliana. GFP expression was repressed by 100 ng ml-1 tetracycline, an order of magnitude lower than the concentration used previously to repress expression in Nicotiana tabacum. Moreover, the level of GFP expression was controlled by varying the concentration of tetracycline in the medium, allowing a titred regulation of transgenic activity that was previously unavailable in A. thaliana. The kinetics of GFP activity were determined following de-repression of the Top10:mGFP5 transgene, with a visible ER-targeted GFP signal appearing from 24 to 48 h after de-repression.

NASA Discipline Plant Biology↗

The regulation and regulatory role of collagenase in bone

Interstitial collagenase plays an important role in both the normal and pathological remodeling of collagenous extracellular matrices, including skeletal tissues. The enzyme is a member of the family of matrix metalloproteinases. Only one rodent interstitial collagenase has been found but there are two human enzymes, human collagenase-1 and -3, the latter being the homologue of the rat enzyme. In developing rat and mouse bone, collagenase is expressed by hypertrophic chondrocytes, osteoblasts, and osteocytes, a situation that is replicated in a fracture callus. Cultured osteoblasts derived from neonatal rat calvariae show greater amounts of collagenase transcripts late in differentiation. These levels can be regulated by parathyroid hormone (PTH), retinoic acid, and insulin-like growth factors, as well as the degree of matrix mineralization. Much of the work on collagenase in bone has been derived from studies on the rat osteosarcoma cell line, UMR 106-01. All bone-resorbing agents stimulate these cells to produce collagenase mRNA and protein, with PTH being the most potent stimulator. Determination of secreted levels of collagenase has been difficult because UMR cells, normal rat osteoblasts, and rat fibroblasts possess a scavenger receptor that removes the enzyme from the extracellular space, internalizes and degrades it, thus imposing another level of control. PTH can also regulate the abundance of the receptor as well as the expression and synthesis of the enzyme. Regulation of the collagenase gene by PTH appears to involve the cAMP pathway as well as a primary response gene, possibly Fos, which then contributes to induction of the collagenase gene. The rat collagenase gene contains an activator protein-1 sequence that is necessary for basal expression, but other promoter regions may also participate in PTH regulation. Thus, there are many levels of regulation of collagenase in bone perhaps constraining what would otherwise be a rampant enzyme.

NASA Discipline Musculoskeletal↗

Expression of the Acyl-Coenzyme A: Cholesterol Acyltransferase GFP Fusion Protein in Sf21 Insect Cells

The enzyme acyl-coenzyme A:cholesterol acyltransferase (ACAT) is an important contributor to the pathological expression of plaque leading to artherosclerosis n a major health problem. Adequate knowledge of the structure of this protein will enable pharmaceutical companies to design drugs specific to the enzyme. ACAT is a membrane protein located in the endoplasmic reticulum.t The protein has never been purified to homogeneity.T.Y. Chang's laboratory at Dartmouth College provided a 4-kb cDNA clone (K1) coding for a structural gene of the protein. We have modified the gene sequence and inserted the cDNA into the BioGreen His Baculovirus transfer vector. This was successfully expressed in Sf2l insect cells as a GFP-labeled ACAT protein. The advantage to this ACAT-GFP fusion protein (abbreviated GCAT) is that one can easily monitor its expression as a function of GFP excitation at 395 nm and emission at 509 nm. Moreover, the fusion protein GCAT can be detected on Western blots with the use of commercially available GFP antibodies. Antibodies against ACAT are not readily available. The presence of the 6xHis tag in the transfer vector facilitates purification of the recombinant protein since 6xHis fusion proteins bind with high affinity to Ni-NTA agarose. Obtaining highly pure protein in large quantities is essential for subsequent crystallization. The purified GCAT fusion protein can readily be cleaved into distinct GFP and ACAT proteins in the presence of thrombin. Thrombin digests the 6xHis tag linking the two protein sequences. Preliminary experiments have indicated that both GCAT and ACAT are expressed as functional proteins. The ultimate aim is to obtain large quantities of the ACAT protein in pure and functional form appropriate for protein crystal growth. Determining protein structure is the key to the design and development of effective drugs. X-ray analysis requires large homogeneous crystals that are difficult to obtain in the gravity environment of earth. Protein crystals grown in microgravity are often larger and have fewer defects than those grown on earth. The analysis of higher quality space-grown crystals will assist in structure-based drug design. We have successfully grown GCAT-infected Sf21 cells in both adhesion and suspension cultures. Expression levels of GCAT in cell lines such as Sf9 and High Five appear to be reduced. We intend to replicate GCAT expression in all three cell lines using the NASA rotating wall bioreactor which effectively duplicates a microgravity environment. The bioreactor itself could be launched to study the expression of the GFP and GCAT proteins in the actual microgravity environment achieved in orbit.

Mahtani, H. K.↗

Characterization and distribution of a maize cDNA encoding a peptide similar to the catalytic region of second messenger dependent protein kinases

Maize (Zea mays) roots respond to a variety of environmental stimuli which are perceived by a specialized group of cells, the root cap. We are studying the transduction of extracellular signals by roots, particularly the role of protein kinases. Protein phosphorylation by kinases is an important step in many eukaryotic signal transduction pathways. As a first phase of this research we have isolated a cDNA encoding a maize protein similar to fungal and animal protein kinases known to be involved in the transduction of extracellular signals. The deduced sequence of this cDNA encodes a polypeptide containing amino acids corresponding to 33 out of 34 invariant or nearly invariant sequence features characteristic of protein kinase catalytic domains. The maize cDNA gene product is more closely related to the branch of serine/threonine protein kinase catalytic domains composed of the cyclic-nucleotide- and calcium-phospholipid-dependent subfamilies than to other protein kinases. Sequence identity is 35% or more between the deduced maize polypeptide and all members of this branch. The high structural similarity strongly suggests that catalytic activity of the encoded maize protein kinase may be regulated by second messengers, like that of all members of this branch whose regulation has been characterized. Northern hybridization with the maize cDNA clone shows a single 2400 base transcript at roughly similar levels in maize coleoptiles, root meristems, and the zone of root elongation, but the transcript is less abundant in mature leaves. In situ hybridization confirms the presence of the transcript in all regions of primary maize root tissue.

NASA Program Space Biology↗

Inferring the palaeoenvironment of ancient bacteria on the basis of resurrected proteins

Features of the physical environment surrounding an ancestral organism can be inferred by reconstructing sequences of ancient proteins made by those organisms, resurrecting these proteins in the laboratory, and measuring their properties. Here, we resurrect candidate sequences for elongation factors of the Tu family (EF-Tu) found at ancient nodes in the bacterial evolutionary tree, and measure their activities as a function of temperature. The ancient EF-Tu proteins have temperature optima of 55-65 degrees C. This value seems to be robust with respect to uncertainties in the ancestral reconstruction. This suggests that the ancient bacteria that hosted these particular genes were thermophiles, and neither hyperthermophiles nor mesophiles. This conclusion can be compared and contrasted with inferences drawn from an analysis of the lengths of branches in trees joining proteins from contemporary bacteria, the distribution of thermophily in derived bacterial lineages, the inferred G + C content of ancient ribosomal RNA, and the geological record combined with assumptions concerning molecular clocks. The study illustrates the use of experimental palaeobiochemistry and assumptions about deep phylogenetic relationships between bacteria to explore the character of ancient life.

Bacteria/classification/metabolism↗

Simultaneous enhancement of multiple functional properties using evolution-informed protein design

Abstract A major challenge in protein design is to augment existing functional proteins with multiple property enhancements. Altering several properties likely necessitates numerous primary sequence changes, and novel methods are needed to accurately predict combinations of mutations that maintain or enhance function. Models of sequence co-variation (e.g., EVcouplings), which leverage extensive information about various protein properties and activities from homologous protein sequences, have proven effective for many applications including structure determination and mutation effect prediction. We apply EVcouplings to computationally design variants of the model protein TEM-1 β -lactamase. Nearly all the 14 experimentally characterized designs were functional, including one with 84 mutations from the nearest natural homolog. The designs also had large increases in thermostability, increased activity on multiple substrates, and nearly identical structure to the wild type enzyme. This study highlights the efficacy of evolutionary models in guiding large sequence alterations to generate functional diversity for protein design applications.

59 BASIC BIOLOGICAL SCIENCES↗

The sequence, and its evolutionary implications, of a Thermococcus celer protein associated with transcription

Through random search, a gene from Thermococcus celer has been identified and sequenced that appears to encode a transcription-associated protein (110 amino acid residues). The sequence has clear homology to approximately the last half of an open reading frame reported previously for Sulfolobus acidocaldarius [Langer, D. & Zillig, W. (1993) Nucleic Acids Res. 21, 2251]. The protein translations of these two archaeal genes in turn are homologs of a small subunit found in eukaryotic RNA polymerase I (A12.2) and the counterpart of this from RNA polymerase II (B12.6). Homology is also seen with the eukaryotic transcription factor TFIIS, but it involves only the terminal 45 amino acids of the archaeal proteins. Evolutionary implications of these homologies are discussed.

Non-NASA Center↗

High throughput, accurate gene annotation through AI and HPC-enabled structural analysis

With the advances in next generation sequencing technologies, the number of sequenced genomes is growing exponentially, resulting in a technology bottleneck for the translation of sequence information into usable hypotheses about the function of each gene. We have proposed leveraging our leadership high-performance computing (HPC) resources to help break this annotation bottleneck. Here we design an HPC-based framework to infer gene function from gene sequence by incorporating information about protein structure and interactions predicted by deep learning approaches. Accurate functional prediction and gene annotation using computational methods will facilitate breakthroughs in the genomic sciences essential to understanding and harnessing life processes in bacteria, fungi and plants. The development and applications of the state-of-the-art deep neural networks to protein structural modeling, interaction prediction, sequence comparison, and quality assessment of protein structural models will be made possible by leadership computational resources. These HPC-enabled bioinformatics and molecular modeling tools will provide powerful insights into molecular functions of genes.

59 BASIC BIOLOGICAL SCIENCES↗

Archaebacterial rhodopsin sequences: Implications for evolution

It was proposed over 10 years ago that the archaebacteria represent a separate kingdom which diverged very early from the eubacteria and eukaryotes. It follows that investigations of archaebacterial characteristics might reveal features of early evolution. So far, two genes, one for bacteriorhodopsin and another for halorhodopsin, both from Halobacterium halobium, have been sequenced. We cloned and sequenced the gene coding for the polypeptide of another one of these rhodopsins, a halorhodopsin in Natronobacterium pharaonis. Peptide sequencing of cyanogen bromide fragments, and immuno-reactions of the protein and synthetic peptides derived from the C-terminal gene sequence, confirmed that the open reading frame was the structural gene for the pharaonis halorhodopsin polypeptide. The flanking DNA sequences of this gene, as well as those of other bacterial rhodopsins, were compared to previously proposed archaebacterial consensus sequences. In pairwise comparisons of the open reading frame with DNA sequences for bacterio-opsin and halo-opsin from Halobacterium halobium, silent divergences were calculated. These indicate very considerable evolutionary distance between each pair of genes, even in the dame organism. In spite of this, three protein sequences show extensive similarities, indicating strong selective pressures.

Lanyi, J. K.↗

Machine learning approaches for integrating multi-omics data to expand microbiome annotation (Final Technical Report)

We fulfilled all original three aims of the proposal. Following the earlier release (during the first phase of the project at Montana) of software that identifies and fills gaps in the annotation of metabolic proteins within bacterial genomes, we have nearly completed a second gap-filling tool that improves accuracy and explainability. We completed software for alignment-based annotation of protein coding DNA, allowing for coding frameshifts caused by sequencing error. Finally, we completed a neural embedding model for identifying similarities between protein sequences based on amino-wise latent vectors.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning approaches for integrating multi-omics data to expand microbiome annotation

Preliminary: This final report corresponds to a grant (DE-SC0021216) that was awarded to the University of Montana. Mid-way through the grant period, I relocated from the University of Montana to the University of Arizona. The grant was ended at University of Montana in late 2022, with all efforts concluding on 08/26/22; the remaining funds supporting the project were relinquished by University of Montana, and were later awarded to University of Arizona under a new grant, with start date 04/01/23. This report focuses on results of research efforts at UMontana through 08/26/22. Results: We made progress in each of the three aims of the proposal. We released software that identifies and fills gaps in the annotation of metabolic proteins within bacterial genomes. We made substantial progress in developing software for alignment-based annotation of protein coding DNA, allowing for coding frameshifts caused by sequencing error. Finally, we made notable progress in developing AI methods (specifically: a neural embedding model) for identifying similarities between protein sequences based on amino-wise latent vectors. These efforts were supplemented by development of methods for protein modeling in support of predicting protein-drug binding activity, and by my leadership of a team in the NIH/DOE 2021 Petabyte-Scale Sequence Search hack-a-thon.

59 BASIC BIOLOGICAL SCIENCES↗

Frameshifting Stimulatory Sequence Induces Large Structural Change of Ribosomal Proteins When Bound to E. coli Ribosomes

Biological macromolecular machines occupy a continuum of structural conformations to perform cellular tasks. Mapping this conformational space provides an insight into its functionality. While the cryo-electron microscopy resolution revolution has expanded our ability to characterize the conformational continuums, there are obstacles in structurally characterizing regions of high flexibility. These technical barriers have impeded characterization of flexible ribosomal proteins when the ribosome is interacting with mRNA stem-loop structures such as a frameshifting stimulatory sequence (FSS). Small-angle neutron/X-ray scattering and electron microscopy were used to study ribosomal samples and compared structural differences between a ribosome that is bound to an FSS stem-loop compared to a ribosome bound to linear mRNA. This comparison shows that a large protein stalk elongates by 22% when the 70S interacts with an mRNA stem-loop. Finally, our results suggest that ribosomal proteins have extensive flexibility and may influence important ribosomal mechanisms, such as those that involve FSS.

36 MATERIALS SCIENCE↗

Origins of the protein synthesis cycle

Largely derived from experiments in molecular evolution, a theory of protein synthesis cycles has been constructed. The sequence begins with ordered thermal proteins resulting from the self-sequencing of mixed amino acids. Ordered thermal proteins then aggregate to cell-like structures. When they contained proteinoids sufficiently rich in lysine, the structures were able to synthesize offspring peptides. Since lysine-rich proteinoid (LRP) also catalyzes the polymerization of nucleoside triphosphate to polynucleotides, the same microspheres containing LRP could have synthesized both original cellular proteins and cellular nucleic acids. The LRP within protocells would have provided proximity advantageous for the origin and evolution of the genetic code.

Fox, S. W.↗

A pollen-specific novel calmodulin-binding protein with tetratricopeptide repeats

Calcium is essential for pollen germination and pollen tube growth. A large body of information has established a link between elevation of cytosolic Ca(2+) at the pollen tube tip and its growth. Since the action of Ca(2+) is primarily mediated by Ca(2+)-binding proteins such as calmodulin (CaM), identification of CaM-binding proteins in pollen should provide insights into the mechanisms by which Ca(2+) regulates pollen germination and tube growth. In this study, a CaM-binding protein from maize pollen (maize pollen calmodulin-binding protein, MPCBP) was isolated in a protein-protein interaction-based screening using (35)S-labeled CaM as a probe. MPCBP has a molecular mass of about 72 kDa and contains three tetratricopeptide repeats (TPR) suggesting that it is a member of the TPR family of proteins. MPCBP protein shares a high sequence identity with two hypothetical TPR-containing proteins from Arabidopsis. Using gel overlay assays and CaM-Sepharose binding, we show that the bacterially expressed MPCBP binds to bovine CaM and three CaM isoforms from Arabidopsis in a Ca(2+)-dependent manner. To map the CaM-binding domain several truncated versions of the MPCBP were expressed in bacteria and tested for their ability to bind CaM. Based on these studies, the CaM-binding domain was mapped to an 18-amino acid stretch between the first and second TPR regions. Gel and fluorescence shift assays performed with CaM and a CaM-binding synthetic peptide further confirmed MPCBP binding to CaM. Western, Northern, and reverse transcriptase-polymerase chain reaction analysis have shown that MPCBP expression is specific to pollen. MPCBP was detected in both soluble and microsomal proteins. Immunoblots showed the presence of MPCBP in mature and germinating pollen. Pollen-specific expression of MPCBP, its CaM-binding properties, and the presence of TPR motifs suggest a role for this protein in Ca(2+)-regulated events during pollen germination and growth.

NASA Discipline Plant Biology↗

Blueprinting extendable nanomaterials with standardized protein blocks

A wooden house frame consists of many different lumber pieces, but because of the regularity of these building blocks, the structure can be designed using straightforward geometrical principles. The design of multicomponent protein assemblies, in comparison, has been much more complex, largely owing to the irregular shapes of protein structures. Here we describe extendable linear, curved and angled protein building blocks, as well as inter-block interactions, that conform to specified geometric standards; assemblies designed using these blocks inherit their extendability and regular interaction surfaces, enabling them to be expanded or contracted by varying the number of modules, and reinforced with secondary struts. Using X-ray crystallography and electron microscopy, we validate nanomaterial designs ranging from simple polygonal and circular oligomers that can be concentrically nested, up to large polyhedral nanocages and unbounded straight ‘train track’ assemblies with reconfigurable sizes and geometries that can be readily blueprinted. Because of the complexity of protein structures and sequence–structure relationships, it has not previously been possible to build up large protein assemblies by deliberate placement of protein backbones onto a blank three-dimensional canvas; the simplicity and geometric regularity of our design platform now enables construction of protein nanomaterials according to ‘back of an envelope’ architectural blueprints.

36 MATERIALS SCIENCE↗

Quantifying Structural Relationships of Metal-Binding Sites Suggests Origins of Biological Electron Transfer

Biological redox reactions drive planetary biogeochemical cycles. Using a novel, structure-guided sequence analysis of proteins, we explored the patterns of evolution of enzymes responsible for these reactions. Our analysis reveals that the folds that bind transition metal–containing ligands have similar structural geometry and amino acid sequences across the full diversity of proteins. Similarity across folds reflects the availability of key transition metals over geological time and strongly suggests that transition metal–ligand binding had a small number of common peptide origins. We observe that structures central to our similarity network come primarily from oxidoreductases, suggesting that ancestral peptides may have also facilitated electron transfer reactions. Last, our results reveal that the earliest biologically functional peptides were likely available before the assembly of fully functional protein domains over 3.8 billion years ago. Thus, life is a special, very complex form of motion of matter, but this form did not always exist, and it is not separated from inorganic nature by an impassable abyss; rather, it arose from inorganic nature as a new property in the process of evolution of the world. We must study the history of this evolution if we want to solve the problem of the origin of life.

Yana Bromberg↗

Water, Solute, and Ion Transport in De Novo-Designed Membrane Protein Channels

Biological organisms engineer peptide sequences to fold into membrane pore proteins capable of performing a wide variety of transport functions. Synthetic de novo-designed membrane pores can mimic this approach to achieve a potentially even larger set of functions. Here, in this work, we explore water, solute, and ion transport in three de novo designed β-barrel membrane channels in the 5–10 Å pore size range. We show that these proteins form passive membrane pores with high water transport efficiencies and size rejection characteristics consistent with the pore size encoded in the protein structure. Ion conductance and ion selectivity measurements also show trends consistent with the pore size, with the two larger pores showing weak cation selectivity. MD simulations of water and ion transport and solute size exclusion are consistent with the experimental trends and provide further insights into structure–function correlations in these membrane pores.

59 BASIC BIOLOGICAL SCIENCES↗