Search NASASearch

SEARCH · Search NASA

Results for “GENETIC CODE”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

An archaeal genetic code with all TAG codons as pyrrolysine

Multiple genetic codes developed during the evolution of eukaryotes and bacteria, yet no alternative genetic code is known for archaea. We used proteomics to confirm our prediction that certain archaea consistently incorporate pyrrolysine (Pyl) at TAG codons, supporting an alternative archaeal genetic code that we designate the Pyl code. This genetic code has 62 sense codons encoding 21 amino acids. In contrast to monophyletic genetic code distributions in bacteria, the archaeal Pyl code occurs sporadically, indicating that it arose independently in multiple lineages. We discovered that more than 1800 archaeal proteins contain Pyl, increasing the number of such proteins by two orders of magnitude. Additionally, five Pyl transfer RNA (tRNA) pyrrolysyl–tRNA synthetase pairs from Pyl-code archaea were used to introduce Pyl analogs into proteins in Escherichia coli.

Kivenson, Veronika [University of California, Berk

Developing a pipeline to expand the genetic code of diverse bacteria for microbial engineering

Microbial biotechnologies are key to addressing grand challenges to promote human health, reverse carbon emissions, recycle mixed plastic waste, remediate contaminated soils, and achieve sustainable economies. Synthetic biology has enabled design of diverse microbes and their proteins for useful purposes, but the narrowness of the natural genetic code limits functional diversity (e.g., biosynthesis) of engineered microbes. The natural genetic code defines the fundamental rules of translating genetic information into proteins comprised of 22 ‘canonical’ amino acids. However, using a technique called genetic code expansion (GCE), the chemical properties and therefore functions of proteins can be transformed by incorporation of one or more of ~200 chemically diverse ‘non-canonical’ amino acids. The effective application of genetic code expansion in diverse microbes has the potential to revolutionize biotechnology. However, despite over 50 years of research and its transformative potential, the application of genetic code expansion has been limited to a handful of bacterial species. In this project, we will perform three tasks to both overcome the barriers that prevent wide spread adoption of GCE as molecular tool and demonstrate its potential for biotechnological applications. Specifically, we will (1) develop a genetic engineering methodology that will enable use of GCE in a broad range of bacterial hosts, (2) use high-throughput functional genomics methods to identify physiological responses to both genetic code expansion and exposure to non-canonical amino acids in three different bacteria, and (3) demonstrate an application of GCE by selectively incorporate non-canonical amino acids into surface displayed peptides such as those used for biomining.

59 BASIC BIOLOGICAL SCIENCES

Efficient genetic code expansion tools enable in vivo study of lysine acetylation in non-model bacteria

Recent proteomic advancements have revealed widespread Nε-lysine acetylation in pathways governing pathogenicity, metabolism, and antibiotic resistance in bacteria. The spontaneous, non-specific nature of this modification in prokaryotes obscures its biological role, necessitating prokaryotic specific in vivo interrogation systems. Genetic Code Expansion (GCE) offers a powerful method to investigate the roles and regulation dynamics of acetyl-lysine in vivo with the precise incorporation of a suite of non-canonical amino acids, including acetyl-lysine analogs. However, its use has been largely restricted to E. coli strains due to challenges associated with implementation and optimization of the technology in more diverse bacterial strains. Here, we present a bacterial host-agnostic, readily optimizable GCE platform designed to site-specifically incorporate non-canonical amino acids into target proteins within living bacteria. We further demonstrate the versatility of this technology by showcasing, for the first time, the successful incorporation of acetyl-lysine in a non- E. coli bacterium.

59 BASIC BIOLOGICAL SCIENCES

tonalli : an asexual genetic code to characterize APOGEE-2 stellar spectra. I. Validation with synthetic and solar spectra

ABSTRACT We present tonalli, a spectroscopic analysis python code that efficiently predicts effective temperature, stellar surface gravity, metallicity, $\alpha$-element abundance, and rotational and radial velocities for stars with effective temperatures between 3200 and 6250 K, observed with the Apache Point Observatory Galactic Evolution Experiment 2 (APOGEE-2). tonalli implements an asexual genetic algorithm to optimize the finding of the best comparison between a target spectrum and the continuum-normalized synthetic spectra library from the Model Atmospheres with a Radiative and Convective Scheme (MARCS), which is interpolated in each generation. Using simulated observed spectra and the APOGEE-2 solar spectrum of Vesta, we study the performance, limitations, accuracy, and precision of our tool. Finally, a Monte Carlo realization was implemented to estimate the uncertainties of each derived stellar parameter.

Adame, Lucía (ORCID:0000000263286099)

A Route to Design Novel Functional Peptides by Applying a Denoising Diffusional Model to mRNA Display Libraries

In vitro directed evolution techniques, such as mRNA display, enable peptide ligand discovery and optimization. However, physical libraries that rely on a genetic code can only search a small fraction of sequence space due to inherent biases in the genetic code and experimental limitations. To address this challenge, denoising diffusion implicit models (DDIMs) are applied to generate novel peptide ligands against B‐cell lymphoma extra‐large (Bcl‐x L ), a key cancer target. Starting with high‐throughput sequencing data from previous selections, a DDIM is trained to produce novel sequences with high affinity binding. Experimental validation confirms that most generated sequences are functionally equivalent to the original library members for Bcl‐x L binding and demonstrated comparable binding kinetics and affinity relative to the wildtype and nearest original neighbors. Importantly, this approach generated rare sequences not easily accessible via mutation and directed evolution. These results indicate that DDIMs can complement and expand directed evolution data, efficiently exploring underrepresented regions of sequence space. This approach provides a broadly applicable framework for accelerating ligand discovery and optimizing molecular properties across diverse targets.

Qi, Pearl [Mork Family Department of Chemical Engi

Probing the limits of genetic recoding using multi-omics-guided evolution

Engineering the genetic code—by reassigning multiple of the 64 natural codons—enables making organisms resistant to all viruses, preventing genetic information exchange, and allowing the biosynthesis of genetically encoded unnatural polymers. However, synonymous codon replacement—recoding—is frequently lethal, and how recoding impacts fitness remains poorly explored. Here, we explore these effects using genome synthesis, directed evolution, and genome-transcriptome-translatome-proteome co-profiling on multiple synthetic Escherichia coli genomes. We construct six partially recoded E. coli strains bearing up to 45.8% of a synthetic genome with a deleterious 57-codon genetic code. As our analyses revealed widespread defects—including unassigned codons in Syn61 and Syn57—we apply multi-omics to revise our genome design and mitigate defects. Using multi-omics, we show that recoding induces transcriptional and translational changes leading to fitness defects under hundreds of conditions. Finally, we develop a multi-omics-guided evolution strategy that rapidly restores fitness, enabling genome synthesis with radical changes.

Nyerges, Akos [Harvard Medical School, Boston, MA

A One‐Pot Biocatalytic Cascade to Access Diverse l ‐Phenylalanine Derivatives from Aldehydes or Carboxylic Acids

Abstract Nonstandard amino acids (nsAAs) that are l ‐phenylalanine derivatives with aryl ring functionalization have long been harnessed in natural product synthesis, therapeutic peptide synthesis, and diverse applications of genetic code expansion. Yet, to date, these chiral molecules have often been the products of poorly enantioselective and environmentally harsh organic synthesis routes. Here, we reveal the broad specificity of multiple natural pyridoxal 5′‐phosphate (PLP)‐dependent enzymes, specifically an l ‐threonine transaldolase, a phenylserine dehydratase, and an aminotransferase, toward substrates that contain aryl side chains with diverse substitutions. We exploit this tolerance to construct a one‐pot biocatalytic cascade that achieves high‐yield synthesis of 18 diverse l ‐phenylalanine derivatives from aldehydes under mild aqueous reaction conditions. We demonstrate the addition of a carboxylic acid reductase module to this cascade to enable the biosynthesis of l ‐phenylalanine derivatives from carboxylic acids that may be less expensive or less reactive than the corresponding aldehydes. Finally, we investigate the scalability of the cascade by developing a lysate‐based route for preparative‐scale synthesis of 4‐formyl‐ l ‐phenylalanine, a nsAA with a bio‐orthogonal handle that is not readily market‐accessible. Overall, this work offers an efficient, versatile, and scalable route with the potential to lower manufacturing costs and democratize synthesis for many valuable nsAAs.

Anderson, Shelby R. [Department of Chemical and Bi

NCAP: Noncanonical Amino Acid Parameterization Software for CHARMM Potentials

Noncanonical Amino Acids (NCAAs) provide numerous avenues for introduction of novel functionality to peptides and proteins. NCAAs can be incorporated through solid phase synthesis or genetic code expansion in conjugation with heterologous expression of the encoded protein modification. Due to the difficulty of synthesis, wide chemical space and lack of empirically resolved structures modeling the effects of NCAA mutation is critical for rational protein design. To evaluate the structural and functional perturbations NCAAs introduce we utilize molecular potentials that describe the forces in protein structure. Most potentials such as CHARMM are designed to model canonical residues but can be parameterized in include novel NCAAs. Here, in this work, we introduce NCAP a software package to generate CHARMM compatible parameters from quantum chemical calculation. Unlike currently available tools NCAP is designed to recognize NCAA structure and automatically bridge the gap between DFT calculations and potential parameters. For our software we discuss workflow, validation against canonical parameter sets and comparison to published NCAA-protein structures.

59 BASIC BIOLOGICAL SCIENCES

Rapid discovery and evolution of nanosensors containing fluorogenic amino acids

Binding-activated optical sensors are powerful tools for imaging, diagnostics, and biomolecular sensing. However, biosensor discovery is slow and requires tedious steps in rational design, screening, and characterization. Here we report on a platform that streamlines biosensor discovery and unlocks directed nanosensor evolution through genetically encodable fluorogenic amino acids (FgAAs). Building on the classical knowledge-based semisynthetic approach, we engineer ~15 kDa nanosensors that recognize specific proteins, peptides, and small molecules with up to 100-fold fluorescence increases and subsecond kinetics, allowing real-time and wash-free target sensing and live-cell bioimaging. An optimized genetic code expansion chemistry with FgAAs further enables rapid (~3 h) ribosomal nanosensor discovery via the cell-free translation of hundreds of candidates in parallel and directed nanosensor evolution with improved variant-specific sensitivities (up to ~250-fold) for SARS-CoV-2 antigens. Altogether, this platform could accelerate the discovery of fluorogenic nanosensors and pave the way to modify proteins with other non-standard functionalities for diverse applications.

Biosensors

Enhancing chemical bioproduction with rational control of bacterial post-translational modifications

Efficient conversion of inexpensive feedstocks to valuable chemicals by microbes is critical for a robust bioeconomy, but the ability to rationally design bacteria is hampered by insufficient knowledge of how post translational modifications (PTMs) control bacterial protein function and thus bioproduction phenotypes. Our study will focus on the lysine acetylation, a ubiquitous bacterial PTM that can affect the function of enzymes in central metabolism that are often critical for bioproduction processes, disrupt transcriptional regulation, and reduce translation. However, most lysine acetylation data is observational, which means that we do not know when, how, and what specific acetylated residues affect protein function and bacterial physiology. For our model host, we will use a Pseudomonas putida strain that we previously engineered to convert lignocellulosic feedstocks into chemicals such as itaconic acid (ITA). With this strain, we use a dynamic two-stage bioproduction process in which ITA is produced during a non-growth associated production phase. Production is highest during growth stages when lysine acetylation is low in other organisms (early stationary phase) and stalls in conditions where acetylation is highest (late stationary phase). The switch from high to stalled ITA production is also correlated with an unexpected increase in acetate levels – the precursor to non-enzymatic lysine acetylation. As such, we predict that lysine acetylation plays a substantial role in regulating the metabolic pathways required for ITA production. We will develop a generalizable approach that combines high-throughput genetic screens and cutting-edge genome engineering with state-of-the-art proteomics, metabolomics, and genetic code expansion methods to identify and modulate lysine acetylation patterns in bacteria. Ultimately, these strategies aim to manipulate protein expression and acetylation patterns to enhance bioproduction phenotypes (e.g., sustained ITA production in late stationary phase).

60 APPLIED LIFE SCIENCES

Multimodal Approaches for Leveraging Domain Knowledge with State-of-the-Art Machine Learning to Engineer Biocatalysts

This grant aimed to accelerate the development of specialized enzymes—biological catalysts essential for sustainable manufacturing and medicine—by integrating traditional laboratory evolution with cutting-edge artificial intelligence. To achieve this, we developed a suite of high-throughput sequencing tools and a centralized database to bridge the gap between a protein’s genetic "code" and its physical function. By training machine learning models on large datasets, we also demonstrated the ability to move beyond slow, trial-and-error testing to a "generative" approach, where AI can independently design new, versatile enzymes like tryptophan synthases. Ultimately, these findings demonstrate that combining laboratory data with computer-guided design enables the engineering of highly efficient biological tools with unprecedented speed and precision.

59 BASIC BIOLOGICAL SCIENCES

XES_Neo_Public

XES Neo is a fitting software that was based on the already approved EXAFS Neo genetic algorithm fitting software code. Using the principles of genetics, data fitting is done for x-ray emission spectroscopy data. Using the EXAFS Neo open source code, as well as the open source xes_neo code written by other collaborators, a new final repository for XES Neo was created with several necessary changes for general user use.

Humiston, Alaina [Los Alamos National Laboratory]

Data for "Which plant traits increase soil carbon sequestration? Empirical evidence from a long-term poplar genetic diversity trial"

This archive contains all data and code used by the following publication: Field, J. L., Sloan, B. P., Craig, M. E., Calloway, P., Ottinger, S. L., Mead, T., Abramoff, R. Z., Venegas, M. P., Chhetri, H. B., Haiby, K., Kalluri, U. C., Muchero, W., Schadt, C. W., & Mayes, M. A. (2025). Which plant traits increase soil carbon sequestration? Empirical evidence from a long-term poplar genetic diversity trial (p. 2025.02.17.638464). bioRxiv. https://doi.org/10.1101/2025.02.17.638464 Our analysis combined several soil and root data sets collected by Oak Ridge National Laboratory (ORNL) researchers/collaborators from the Clatskanie Poplar Common Garden in Clatskanie, OR by from 2009-2024. The raw data data files are located */02-data/01-raw/* which we harmonized using the codes in */01-codes/01-harmonize-clatskanie-data-pub.qmd*. The final processed data set used in the paper is found at */02-data/02-processed/clatskanie-c-fit-data.csv* and its columns are described in the table below.

Sloan, Brandon [ORNL] (ORCID:0000000316304271)

ldrd_virus_work

This is a Python code base that takes openly-available genetic information on known viruses and performs supervised machine learning and feature importance analysis on the relationship of the viral genomes to the competence to infect humans or bind to a specific host cell receptor.

Reddy, Tyler [LANL]

Inferring demographic and selective histories from population genomic data using a 2-step approach in species with coding-sparse genomes: an application to human data

Abstract The demographic history of a population, and the distribution of fitness effects (DFE) of newly arising mutations in functional genomic regions, are fundamental factors dictating both genetic variation and evolutionary trajectories. Although both demographic and DFE inference has been performed extensively in humans, these approaches have generally either been limited to simple demographic models involving a single population, or, where a complex population history has been inferred, without accounting for the potentially confounding effects of selection at linked sites. Taking advantage of the coding-sparse nature of the genome, we propose a 2-step approach in which coalescent simulations are first used to infer a complex multi-population demographic model, utilizing large non-functional regions that are likely free from the effects of background selection. We then use forward-in-time simulations to perform DFE inference in functional regions, conditional on the complex demography inferred and utilizing expected background selection effects in the estimation procedure. Throughout, recombination and mutation rate maps were used to account for the underlying empirical rate heterogeneity across the human genome. Importantly, within this framework it is possible to utilize and fit multiple aspects of the data, and this inference scheme represents a generalized approach for such large-scale inference in species with coding-sparse genomes.

Soni, Vivak (ORCID:0000000294969562)

A modular GUI-based program for genetic algorithm-based feedback-assisted wavefront shaping

Abstract We have developed a modular graphical user interface (GUI)-based program for use in genetic algorithm-based feedback-assisted wavefront shaping. The program uses a class-based structure to separate out the universal modules (e.g. GUI, multithreading, optimization algorithms) and hardware-specific modules (e.g. code for different SLMs and cameras). This modular design makes the program easily adaptable to a wide range of lab equipment, while providing easy access to a GUI, multithreading, and three optimization algorithms (phase-stepping, simple genetic, and microgenetic).

97 MATHEMATICS AND COMPUTING