Search NASASearch

SEARCH · Search NASA

Results for “Codon”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Genomic factors shaping codon usage across the Saccharomycotina subphylum

Codon usage bias, or the unequal use of synonymous codons, is observed across genes, genomes, and between species. It has been implicated in many cellular functions, such as translation dynamics and transcript stability, but can also be shaped by neutral forces. We characterized codon usage across 1,154 strains from 1,051 species from the fungal subphylum Saccharomycotina to gain insight into the biases, molecular mechanisms, evolution, and genomic features contributing to codon usage patterns. We found a general preference for A/T-ending codons and correlations between codon usage bias, GC content, and tRNA-ome size. Codon usage bias is distinct between the 12 orders to such a degree that yeasts can be classified with an accuracy >90% using a machine learning algorithm. We also characterized the degree to which codon usage bias is impacted by translational selection. We found it was influenced by a combination of features, including the number of coding sequences, BUSCO count, and genome length. Our analysis also revealed an extreme bias in codon usage in the Saccharomycodales associated with a lack of predicted arginine tRNAs that decode CGN codons, leaving only the AGN codons to encode arginine. Analysis of Saccharomycodales gene expression, tRNA sequences, and codon evolution suggests that avoidance of the CGN codons is associated with a decline in arginine tRNA function. Consistent with previous findings, codon usage bias within the Saccharomycotina is shaped by genomic features and GC bias. However, we find cases of extreme codon usage preference and avoidance along yeast lineages, suggesting additional forces may be shaping the evolution of specific codons.

59 BASIC BIOLOGICAL SCIENCES

Codon bias, nucleotide selection, and genome size predict in situ bacterial growth rate and transcription in rewetted soil

In soils, the first rain after a prolonged dry period represents a major pulse event impacting soil microbial community function, yet we lack a full understanding of the genomic traits associated with the microbial response to rewetting. Genomic traits such as codon usage bias and genome size have been linked to bacterial growth in soils—however, often through measurements in culture. Here, we used metagenome-assembled genomes (MAGs) with 18 O-water stable isotope probing and metatranscriptomics to track genomic traits associated with growth and transcription of soil microorganisms over one week following rewetting of a grassland soil. We found that codon bias in ribosomal protein genes was the strongest predictor of growth rate. We also found higher growth rates in bacteria with smaller genomes, suggesting that reduced genome size enables a faster response to pulses in soil bacteria. Faster transcriptional upregulation of ribosomal protein genes was associated with high codon bias and increased nucleotide skew. We found that several of these relationships existed within phyla, indicating that these associations between genomic traits and activity could be generalized characteristics of soil bacteria. Finally, we used publicly available metagenomes to assess the distribution of codon bias across a pH gradient and found that microbial communities in higher pH soils—which are often more water limited and pulse driven—have higher codon usage bias in their ribosomal protein genes. Together, these results provide evidence that genomic characteristics affect soil microbial activity during rewetting and pose a potential fitness advantage for soil bacteria where water and nutrient availability are episodic.

59 BASIC BIOLOGICAL SCIENCES

Role of Ribosomal Protein bS1 in Orthogonal mRNA Start Codon Selection

In many bacteria, the location of the mRNA start codon is determined by a short ribosome binding site sequence that base pairs with the 3'-end of 16S rRNA (rRNA) in the 30S subunit. Many groups have changed these short sequences, termed the Shine-Dalgarno (SD) sequence in the mRNA and the anti-Shine-Dalgarno (ASD) sequence in 16S rRNA, to create "orthogonal" ribosomes to enable the synthesis of orthogonal polymers in the presence of the endogenous translation machinery. However, orthogonal ribosomes are prone to SD-independent translation. Ribosomal protein bS1, which binds to the 30S ribosomal subunit, is thought to promote translation initiation by shuttling the mRNA to the ribosome. Thus, a better understanding of how the SD and bS1 contribute to start codon selection could help efforts to improve the orthogonality of ribosomes. Here, we engineered the Escherichia coli ribosome to prevent binding of bS1 to the 30S subunit and separate the activity of bS1 binding to the ribosome from the role of the mRNA SD sequence in start codon selection. We find that ribosomes lacking bS1 are slightly less active than wild-type ribosomes in vitro. Furthermore, orthogonal 30S subunits lacking bS1 do not have an improved orthogonality. Our findings suggest that mRNA features outside the SD sequence and independent of binding of bS1 to the ribosome likely contribute to start codon selection and the lack of orthogonality of present orthogonal ribosomes.

59 BASIC BIOLOGICAL SCIENCES

An archaeal genetic code with all TAG codons as pyrrolysine

Multiple genetic codes developed during the evolution of eukaryotes and bacteria, yet no alternative genetic code is known for archaea. We used proteomics to confirm our prediction that certain archaea consistently incorporate pyrrolysine (Pyl) at TAG codons, supporting an alternative archaeal genetic code that we designate the Pyl code. This genetic code has 62 sense codons encoding 21 amino acids. In contrast to monophyletic genetic code distributions in bacteria, the archaeal Pyl code occurs sporadically, indicating that it arose independently in multiple lineages. We discovered that more than 1800 archaeal proteins contain Pyl, increasing the number of such proteins by two orders of magnitude. Additionally, five Pyl transfer RNA (tRNA) pyrrolysyl–tRNA synthetase pairs from Pyl-code archaea were used to introduce Pyl analogs into proteins in Escherichia coli.

Kivenson, Veronika [University of California, Berk

Plastome evolution in annual Brachypodium species reveals widespread heteroplasmy and chloroplast capture, lineage-specific codon usage bias, and low positive selection

Comparative genomics and plastome phylogenomics have advanced significantly in recent years, highlighting the diversity, possible admixture, and non-neutral evolution of the predominantly considered non-recombinant chloroplast genomes in angiosperms. The grass genus Brachypodium serves as a powerful model for studying evolutionary processes in monocots. We analyzed 287 plastomes across the native circum-Mediterranean range of the three annual Brachypodium species ( B. distachyon, B. stacei, B. hybridum ), focusing on their structural variation, selection patterns and phylogenomic relationships. Our analyses confirmed the differentiation of the S and D plastomes, inherited respectively from the diploid progenitor species B. stacei and B. distachyon . We identified novel structural rearrangements and indels, and unique repeat motifs, along with widespread heteroplasmy, particularly in ancestral B. hybridum -D plastotypes. SNP diversity varied among plastotypes, reflecting population dynamics and evolutionary histories, with B. hybridum -D plastotypes showing the highest normalized diversity and B. hybridum -S the lowest. Positive selection was detected in 29 plastid genes by Tajima’s neutrality test, and in nine genes by site and branch-site evolutionary models, including matK, ndhF, rbcL, and rpoC2. Phylogenomic analyses revealed well-supported clades corresponding to the S and D plastome lineages, with frequent chloroplast capture events and long-distance dispersals shaping their evolutionary trajectories.

allopolyploidy

The relationship between gene traits and transcription in soil microbial communities varies by environmental stimulus

Codon and nucleotide frequencies are known to relate to the rate of gene transcription, yet how these traits shape transcriptional profiles of soil microbial communities remains unclear. Here we test the prediction that functional genes with high codon optimization and energetically lower cost nucleotides (i.e., nucleotides requiring less adenosine triphosphate (ATP) for synthesis) have higher transcriptional expression in a soil microbial community. In laboratory incubations, we subjected an agricultural soil to two separate short-term environmental changes: labile carbon (glucose) addition or a sudden 30-min increase in temperature from 20 °C to 60 °C. Using the total genomic codon frequencies to predict preferred codon usage for each taxon, we then estimated codon optimization for each transcript. On the community level, we found a higher average level of codon optimization after the addition of glucose. Synonymous nucleotide composition in the transcript pool also shifted towards energetically cheaper nucleotides, favoring uracil (U) over adenine (A) and cytosine (C) over guanine (G). Similarly, we found that encoded amino acid usage shifted towards energetically cheaper amino acids in response to labile carbon. In contrast, in communities responding to heat shock, there were no significant differences in the averaged gene traits of expressed transcripts. We used metagenome-assembled-genomes to further examine the ability of gene traits to predict transcriptional responses within and between taxa. We found that traits of individual genes could not reliably predict the level of transcription of a gene within or between taxa—highlighting the limits of this approach. However, we did find that when traits were averaged across several related genes, codon optimization was able to predict levels of transcription in metabolic pathways associated with growth and nutrient uptake in response to glucose. Similar relationships were not observed in response to heat, or for functions associated with stress—such as genes associated with sporulation or heat shock. These results demonstrate that gene traits, such as codon usage, nucleotide selection, and amino acid selection, relate to the transcriptional expression of genes in soil microbial communities and suggests that these relationships may be dependent on both gene function and the specific type of environmental stimuli.

Biological and medical sciences

Three stages during the evolution of the genetic code

A diversification of the genetic code based on the number of codons available for the proteinous amino acids is established. Three groups of amino acids during evolution of the code are distinguished. On the basis of their chemical complexity and a small codon number those amino acids emerging later in a translation process are derived. Both criteria indicate that His, Phe, Tyr, Cys and either Lys or Asn were introduced in the second stage, whereas the number of codons alone gives evidence that Trp and Met were introduced in the third stage. The amino acids of stage one use purines rich codons, thus purines have been retained in their third codon position. All the amino acids introduced in the second stage, in contrast, use pyrimidines in this codon position. A low abundance of pyrimidines during early translation is derived. This assumption is supported by experiments on non enzymatic replication and interactions of DNA hairpin loops with a complementary strand. A back extrapolation concludes a high purine content of the first nucleic acids which gradually decreased during their evolution. Amino acids independently available form prebiotic synthesis were thus correlated to purine rich codons. Conclusions on prebiotic replication are discussed also in the light of recent codon usage data.

Baumann, U.

Recent evidence for evolution of the genetic code

The genetic code, formerly thought to be frozen, is now known to be in a state of evolution. This was first shown in 1979 by Barrell et al. (G. Barrell, A. T. Bankier, and J. Drouin, Nature [London] 282:189-194, 1979), who found that the universal codons AUA (isoleucine) and UGA (stop) coded for methionine and tryptophan, respectively, in human mitochondria. Subsequent studies have shown that UGA codes for tryptophan in Mycoplasma spp. and in all nonplant mitochondria that have been examined. Universal stop codons UAA and UAG code for glutamine in ciliated protozoa (except Euplotes octacarinatus) and in a green alga, Acetabularia. E. octacarinatus uses UAA for stop and UGA for cysteine. Candida species, which are yeasts, use CUG (leucine) for serine. Other departures from the universal code, all in nonplant mitochondria, are CUN (leucine) for threonine (in yeasts), AAA (lysine) for asparagine (in platyhelminths and echinoderms), UAA (stop) for tyrosine (in planaria), and AGR (arginine) for serine (in several animal orders) and for stop (in vertebrates). We propose that the changes are typically preceded by loss of a codon from all coding sequences in an organism or organelle, often as a result of directional mutation pressure, accompanied by loss of the tRNA that translates the codon. The codon reappears later by conversion of another codon and emergence of a tRNA that translates the reappeared codon with a different assignment. Changes in release factors also contribute to these revised assignments. We also discuss the use of UGA (stop) as a selenocysteine codon and the early history of the code.

Review

Three stages in the evolution of the genetic code

A diversification of the genetic code based on the number of codons available for the proteinous amino acids is established. Three groups of amino acids during evolution of the code are distinguished. On the basis of their chemical complexity those amino acids emerging later in a translation process are derived. Codon number and chemical complexity indicate that His, Phe, Tyr, Cys and either Lys or Asn were introduced in the second stage, whereas the number of codons alone gives evidence that Trp and Met were introduced in the third stage. The amino acids of stage 1 use purine-rich codons, while all the amino acids introduced in the second stage, in contrast, use pyrimidines in the third position of their codons. A low abundance of pyrimidines during early translation is derived. This assumption is supported by experiments on non-enzymatic replication and interactions of hairpin loops with a complementary strand. A back extrapolation concludes a high purine content of the first nucleic acids, which gradually decreased during their evolution. Amino acids independently available from prebiotic synthesis were thus correlated to purine-rich codons. Implications on the prebiotic replication are discussed also in the light of recent codon usage data.

Review

On the possible origin and evolution of the genetic code

The genetic code is examined for indications of possible preceding codes that existed during early evolution. Eight of the 20 amino acids are coded by 'quartets' of codons with fourfold degeneracy, and 16 such quartets can exist, so that an earlier code could have provided for 15 or 16 amino acids, rather than 20. If twofold degeneracy is postulated for the first position of the codon, there could have been ten amino acids in the code. It is speculated that these may have been phenylalanine, valine, proline, alanine, histidine, glutamine, glutanic acid, aspartic acid, cysteine and glycine. There is a notable deficiency of arginine in proteins, despite the fact that it has six codons. Simultaneously, there is more lysine in proteins than would be expected from its two codons, if the four bases in mRNA are equiprobable and are arranged randomly. It is speculated that arginine is an 'intruder' into the genetic code, and that it may have displayed another amino acid such as ornithine, or may even have displayed lysine from some of its previous codon assignments. As a result, natural selection has favored lysine against the fact that it has only two codons.

Jukes, T. H.

Theoretical foundations for quantitative paleogenetics. III - The molecular divergence of nucleic acids and proteins for the case of genetic events of unequal probability

Theoretical equations are derived for molecular divergence with respect to gene and protein structure in the presence of genetic events with unequal probabilities: amino acid and base compositions, the frequencies of nucleotide replacements, the usage of degenerate codons, the distribution of fixed base replacements within codons and the distribution of fixed base replacements among codons. Results are presented in the form of tables relating the probabilities of given numbers of codon base changes with respect to the original codon for the alpha hemoglobin, beta hemoglobin, myoglobin, cytochrome c and parvalbumin group gene families. Application of the calculations to the rabbit alpha and beta hemoglobin mRNAs and proteins indicates that the genes are separated by about 425 fixed based replacements distributed over 114 codon sites, which is a factor of two greater than previous estimates. The theoretical results also suggest that many more base replacements are required to effect a given gene or protein structural change than previously believed.

Holmquist, R.

An analysis of the metabolic theory of the origin of the genetic code

A computer program was used to test Wong's coevolution theory of the genetic code. The codon correlations between the codons of biosynthetically related amino acids in the universal genetic code and in randomly generated genetic codes were compared. It was determined that many codon correlations are also present within random genetic codes and that among the random codes there are always several which have many more correlations than that found in the universal code. Although the number of correlations depends on the choice of biosynthetically related amino acids, the probability of choosing a random genetic code with the same or greater number of codon correlations as the universal genetic code was found to vary from 0.1% to 34% (with respect to a fairly complete listing of related amino acids). Thus, Wong's theory that the genetic code arose by coevolution with the biosynthetic pathways of amino acids, based on codon correlations between biosynthetically related amino acids, is statistical in nature.

NASA Discipline Exobiology

Probing the limits of genetic recoding using multi-omics-guided evolution

Engineering the genetic code—by reassigning multiple of the 64 natural codons—enables making organisms resistant to all viruses, preventing genetic information exchange, and allowing the biosynthesis of genetically encoded unnatural polymers. However, synonymous codon replacement—recoding—is frequently lethal, and how recoding impacts fitness remains poorly explored. Here, we explore these effects using genome synthesis, directed evolution, and genome-transcriptome-translatome-proteome co-profiling on multiple synthetic Escherichia coli genomes. We construct six partially recoded E. coli strains bearing up to 45.8% of a synthetic genome with a deleterious 57-codon genetic code. As our analyses revealed widespread defects—including unassigned codons in Syn61 and Syn57—we apply multi-omics to revise our genome design and mitigate defects. Using multi-omics, we show that recoding induces transcriptional and translational changes leading to fitness defects under hundreds of conditions. Finally, we develop a multi-omics-guided evolution strategy that rapidly restores fitness, enabling genome synthesis with radical changes.

Nyerges, Akos [Harvard Medical School, Boston, MA

Profiling expression strategies for a type III polyketide synthase in a lysate-based, cell-free system

Abstract Some of the most metabolically diverse species of bacteria (e.g., Actinobacteria) have higher GC content in their DNA, differ substantially in codon usage, and have distinct protein folding environments compared to tractable expression hosts like Escherichia coli . Consequentially, expressing biosynthetic gene clusters (BGCs) from these bacteria in E. coli often results in a myriad of unpredictable issues with regard to protein expression and folding, delaying the biochemical characterization of new natural products. Current strategies to achieve soluble, active expression of these enzymes in tractable hosts can be a lengthy trial-and-error process. Cell-free expression (CFE) has emerged as a valuable expression platform as a testbed for rapid prototyping expression parameters. Here, we use a type III polyketide synthase from Streptomyces griseus , RppA, which catalyzes the formation of the red pigment flaviolin, as a reporter to investigate BGC refactoring techniques. We applied a library of constructs with different combinations of promoters and rppA coding sequences to investigate the synergies between promoter and codon usage. Subsequently, we assess the utility of cell-free systems for prototyping these refactoring tactics prior to their implementation in cells. Overall, codon harmonization improves natural product synthesis more than traditional codon optimization across cell-free and cellular environments. More importantly, the choice of coding sequences and promoters impact protein expression synergistically, which should be considered for future efforts to use CFE for high-yield protein expression. The promoter strategy when applied to RppA was not completely correlated with that observed with GFP, indicating that different promoter strategies should be applied for different proteins. In vivo experiments suggest that there is correlation, but not complete alignment between expressing in cell free and in vivo. Refactoring promoters and/or coding sequences via CFE can be a valuable strategy to rapidly screen for catalytically functional production of enzymes from BCGs, which advances CFE as a tool for natural product research.

59 BASIC BIOLOGICAL SCIENCES

A change in the genetic code in Mycoplasma capricolum

Mycoplasma capricolum was previously found to use UGA instead of UGG as its codon for tryptophan and to contain 75 percent A + T in its DNA. The codon change could have been due to mutational pressure to replace C + G by A + T, resulting in the replacement of UGA stop codons by UAA, change of the anticodon in tryptophan tRNA from CCA to UCA, and replacement of UGG tryptophan codons by UGA. None of these changes should have been deleterious.

Jukes, T. H.