Search NASA⌕ Search

SEARCH · Search NASA

Results for “sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

CABO-16S—a Combined Archaea, Bacteria, Organelle 16S rRNA database framework for amplicon analysis of prokaryotes and eukaryotes in environmental samples

Abstract Identification of both prokaryotic and eukaryotic microorganisms in environmental samples is currently challenged by the need for additional sequencing to obtain separate 16S and 18S ribosomal RNA (rRNA) amplicons or the constraints imposed by “universal” primers. Organellar 16S rRNA sequences are amplified and sequenced along with prokaryote 16S rRNA and provide an alternative method to identify eukaryotic microorganisms. CABO-16S combines bacterial and archaeal sequences from the SILVA database with 16S rRNA sequences of plastids and other organelles from the PR2 database to enable identification of all 16S rRNA sequences. Comparison of CABO-16S with SILVA 138.2 results in equivalent taxonomic classification of mock communities and increased classification of diverse environmental samples. In particular, identification of phototrophic eukaryotes in shallow seagrass environments, marine waters, and lake waters was increased. The CABO-16S framework allows users to add custom sequences for further classification of underrepresented clades and can be easily updated with future releases of reference databases. Addition of sequences obtained from Sanger sequencing of methane seep sediments and curated sequences of the polyphyletic SEEP-SRB1 clade resulted in differentiation of syntrophic and non-syntrophic SEEP-SRB1 in hydrothermal vent sediments. CABO-16S highlights the benefit of combining and amending existing training sets when studying microorganisms in diverse environments.

Eitel, Eryn M. (ORCID:0009000723919297)↗

Hypermut 3: identifying specific mutational patterns in a defined nucleotide context that allows multistate characters

Abstract Motivation The detection of APOBEC3F- and APOBEC3G-induced mutations in virus sequences is useful for identifying hypermutated sequences. These sequences are not representative of viral evolution and can therefore alter the results of downstream sequence analyses if included. We previously published the software Hypermut, which detects hypermutation events in sequences relative to a reference. Two versions of this method are available as a webtool. Neither of these methods consider multistate characters or gaps in the sequence alignment. Results Here, we present an updated, user-friendly web and command-line version of Hypermut with functionality to handle multistate characters and gaps in the sequence alignment. This tool allows for straightforward integration of hypermutation detection into sequence analysis pipelines. As with the previous tool, while the main purpose is to identify G to A hypermutation events, any mutational pattern and context can be specified. Availability and implementation Hypermut 3 is written in Python 3. It is available as a command-line tool at https://github.com/MolEvolEpid/hypermut3 and as a webtool at https://www.hiv.lanl.gov/content/sequence/HYPERMUT/hypermutv3.html.

59 BASIC BIOLOGICAL SCIENCES↗

Decomposing a San Francisco estuary microbiome using long-read metagenomics reveals species- and strain-level dominance from picoeukaryotes to viruses

ABSTRACT Although long-read sequencing has enabled obtaining high-quality and complete genomes from metagenomes, many challenges still remain to completely decompose a metagenome into its constituent prokaryotic and viral genomes. This study focuses on decomposing an estuarine metagenome to obtain a more accurate estimate of microbial diversity. To achieve this, we developed a new bead-based DNA extraction method, a novel bin refinement method, and obtained 150 Gbp of Nanopore sequencing. We estimate that there are ~500 bacterial and archaeal species in our sample and obtained 68 high-quality bins (>90% complete, <5% contamination, ≤5 contigs, contig length of >100 kbp, and all ribosomal and tRNA genes). We also obtained many contigs of picoeukaryotes, environmental DNA of larger eukaryotes such as mammals, and complete mitochondrial and chloroplast genomes and detected ~40,000 viral populations. Our analysis indicates that there are only a few strains that comprise most of the species abundances. IMPORTANCE Ocean and estuarine microbiomes play critical roles in global element cycling and ecosystem function. Despite the importance of these microbial communities, many species still have not been cultured in the lab. Environmental sequencing is the primary way the function and population dynamics of these communities can be studied. Long-read sequencing provides an avenue to overcome limitations of short-read technologies to obtain complete microbial genomes but comes with its own technical challenges, such as needed sequencing depth and obtaining high-quality DNA. We present here new sampling and bioinformatics methods to attempt decomposing an estuarine microbiome into its constituent genomes. Our results suggest there are only a few strains that comprise most of the species abundances from viruses to picoeukaryotes, and to fully decompose a metagenome of this diversity requires 1 Tbp of long-read sequencing. We anticipate that as long-read sequencing technologies continue to improve, less sequencing will be needed.

Lui, Lauren M.↗

Data for FUN-PROSE: A Deep Learning Approach to Predict Condition-Specific Gene Expression in Fungi

mRNA levels of all genes in a genome is a critical piece of information defining the overall state of the cell in a given environmental condition. Being able to reconstruct such condition-specific expression in fungal genomes is particularly important to metabolically engineer these organisms to produce desired chemicals in industrially scalable conditions. Most previous deep learning approaches focused on predicting the average expression levels of a gene based on its promoter sequence, ignoring its variation across different conditions. Here we present FUN-PROSE—a deep learning model trained to predict differential expression of individual genes across various conditions using their promoter sequences and expression levels of all transcription factors. We train and test our model on three fungal species and get the correlation between predicted and observed condition-specific gene expression as high as 0.85. We then interpret our model to extract promoter sequence motifs responsible for variable expression of individual genes. We also carried out input feature importance analysis to connect individual transcription factors to their gene targets. A sizeable fraction of both sequence motifs and TF-gene interactions learned by our model agree with previously known biological information, while the rest corresponds to either novel biological facts or indirect correlations.

Genomics↗

Repetitive proteins that undergo large conformational changes evade structural prediction algorithms

Protein structure prediction algorithms, such as AlphaFold, have accelerated protein design and advanced the understanding of the relationship between amino acid sequence and protein structure. However, these algorithms are limited in their ability to predict the structures of conformationally dynamic, intrinsically disordered, and stimuli-responsive proteins. To evaluate sequence-to-structure predictions of such challenging proteins, we explored a class of conformationally dynamic, repeats-in-toxin (RTX) proteins. RTX proteins adopt intrinsically disordered conformations in the absence of calcium and undergo reversible folding into β-roll structures upon binding to calcium. RTX proteins are characterized by tandem repeats of the sequence GGXGXDXUX, in which X can be any amino acid and U is an aliphatic amino acid. We designed RTX sequence variants with global substitutions of nonconserved amino acids, tandem repeats of consensus sequences GGAGXDTLY, and tandem repeats of scrambled sequences GGAGXDTYL. AlphaFold2 and AlphaFold3 predicted that all of these RTX variants adopt β-roll structures, characteristic of wild-type RTX bound to calcium. However, modeling the predicted structures with molecular dynamics simulations and characterizing the protein variants with circular dichroism spectroscopy, small-angle x-ray scattering, and x-ray crystallography revealed that variants adopt diverse, sequence-dependent structures in the absence and presence of calcium. To better design proteins for applications in biotechnology and sustainability, it is critical to build predictive tools that consider intrinsically disordered protein states and validate these tools with multi-mode, multi-scale experimental data.

Chang, Marina P. [Stanford Univ., CA (United State↗

Repainting the colour–mass diagrams by unearthing the green mountain: dust-rich S0 galaxies in the colour–(galaxy stellar mass) diagram, and the colour–(black hole mass) relations for dust-poor versus dust-rich galaxies

ABSTRACT Lenticular galaxies are notoriously misclassified as elliptical galaxies and, as such, a (disc inclination)-dependent correction for dust is often not applied to the magnitudes of dusty lenticular galaxies. This results in overly red galaxy colours, impacting their distribution in the colour–magnitude diagram. It is revealed how this has led to an underpopulation of the ‘green valley’ by hiding a ‘green mountain’ of massive dust-rich lenticular galaxies – known to be built from gas-rich major mergers – within the ‘red sequence’ of colour–(stellar mass) diagrams. Correcting for dust, a ‘green mountain’ appears at M*,gal ∼ 1011 M⊙, along with signs of an extension to lower masses producing a ‘green range’ or ‘green ridge’ on the green side of the ‘red sequence’ and ‘blue cloud.’ The ‘red sequence’ is shown to be comprised of two components: a red plateau defined by elliptical galaxies with a near-constant colour and by lower-mass dust-poor lenticular galaxies, which are mostly a primordial population but may include faded/transformed spiral galaxies. The presence of the quasi-triangular-shaped galaxy evolution sequence, previously called the ‘Triangal’, is revealed in the galaxy colour–(stellar mass) diagram. It tracks the speciation of galaxies and their associated migration through the diagram. The connection of the ‘Triangal’ to previous galaxy morphology sequences (Fork, Trident, and Comb) is also shown herein. Finally, the colour–(black hole mass) diagram is revisited, revealing how the dust correction generates a blue–green sequence for the spiral and dust-rich lenticular galaxies that is offset from a green–red sequence defined by the dust-poor lenticular and elliptical galaxies.

Graham, Alister W. (ORCID:0000000264969414)↗

VirJenDB: a FAIR (meta)data and bioinformatics platform for all viruses

High-throughput sequencing has generated an unprecedented volume of data. However, researcher-submitted data in repositories requires extensive curation and quality control for reuse. These tasks are hindered by the multiplicity of repositories, the sheer volume of the data, and the complexity of virus (meta)data curation. To address these challenges, VirJenDB offers a user-friendly platform to facilitate versioned, community-driven curation, and ontology development. Virus sequences were ingested from 16 sources, including ~200 fields of metadata or standards, covering taxonomy, sample, and host information. Up to 85 metadata fields have undergone at least one round of curation, and are linked to 15.4 million virus sequences, with 88 % from those infecting eukaryotes and the remaining infecting prokaryotes. Subsets were created, including a novel collection of 0.91 million viral operational taxonomic unit (vOTU) sequences across all viruses, while keeping the original sequences from each vOTU to facilitate downstream analyses, e.g. sequence variation. The VirJenDB web portal (https://www.virjendb.org) provides HTTPS and Application Programming Interface (API) access to the sequence datasets and metadata, offering a search engine, filtering, download, visualizations, and documentation. VirJenDB aims to connect the phage and eukaryotic virus research communities by supporting webtool integration, meta-analyses, and metadata schema extensions.

Saghaei, Shahram↗

Signatures of Mollicutes-related endobacteria in publicly available Mucoromycota genomes

ABSTRACT Mucoromycota fungi and their Mollicutes-related endobacteria (MRE) are an ideal system for studying bacterial–fungal interactions and evolution due to the long-term and intimate nature of their interactions. However, methods for detecting MRE face specific challenges due to the poor representation of MRE in sequencing databases coupled with the high sequence divergence of their genomes, making traditional similarity searches unreliable. This has precluded estimations on the diversity of MRE associated with Mucoromycota. To determine the prevalence of previously undetected MRE in fungal genome sequences, we scanned 389 Mucoromycota genome assemblies available from the National Center for Biotechnology Information for the presence of MRE sequences using publicly available tools to map contigs from fungal assemblies to publicly available MRE genomes. We demonstrate a higher diversity of MRE genomes than previously described in Mucoromycota and a lack of cophylogeny between MRE and the majority of their fungal hosts. This supports the late invasion hypothesis regarding MRE acquisition across most of the examined fungal families. In contrast with other Mucoromycota lineages, MRE from the Gigasporaceae displayed some degree of cophylogeny with their hosts, which may indicate that horizontal transmission is restricted between members of this family or that transmission is strictly vertical. These results underscore the need for a refined process to capture sequencing data from potential fungal endosymbionts to discern their evolution and transmission. Screens of fungal genomes for MRE can help improve the quality of fungal genome assemblies while identifying new MRE lineages to further test hypotheses on their origin and evolution. IMPORTANCE Mollicutes-related endobacteria (MRE) are obligate intracellular bacteria found within Mucoromycota fungi. Despite their frequent detection, MRE roles in host functioning are still unknown. Comparative genomic investigations can improve our understanding of the impact of MRE on their fungal hosts by identifying similarities and differences in MRE genome evolution. However, MRE genomes have only been assembled from a small fraction of Mucoromycota hosts. Here, we demonstrate that MRE can be present yet undetected in publicly available Mucoromycota genome assemblies. We use these newfound sequences to assess the broader diversity of MRE and their phylogenetic relationships with respect to their hosts. We demonstrate that publicly available tools can be used to extract novel MRE sequences from assembled fungal genomes leading to insights on MRE evolution. This work contributes to a greater understanding of the fungal microbiome, which is crucial to improving knowledge on the dynamics and impacts of fungi in microbial ecosystems.

59 BASIC BIOLOGICAL SCIENCES↗

SARS-CoV-2 wastewater variant surveillance: pandemic response leveraging FDA’s GenomeTrakr network

ABSTRACT Wastewater surveillance has emerged as a crucial public health tool for population-level pathogen surveillance. Supported by funding from the American Rescue Plan Act of 2021, the FDA‘s genomic epidemiology program, GenomeTrakr, was leveraged to sequence SARS-CoV-2 from wastewater sites across the United States. This initiative required the evaluation, optimization, development, and publication of new methods and analytical tools spanning sample collection through variant analyses. Version-controlled protocols for each step of the process were developed and published on protocols.io. A custom data analysis tool and a publicly accessible dashboard were built to facilitate real-time visualization of the collected data, focusing on the relative abundance of SARS-CoV-2 variants and sub-lineages across different samples and sites throughout the project. From September 2021 through June 2023, a total of 3,389 wastewater samples were collected, with 2,517 undergoing sequencing and submission to NCBI under the umbrella BioProject,PRJNA757291. Sequence data were released with explicit quality control (QC) tags on all sequence records, communicating our confidence in the quality of data. Variant analysis revealed wide circulation of Delta in the fall of 2021 and captured the sweep of Omicron and subsequent diversification of this lineage through the end of the sampling period. This project successfully achieved two important goals for the FDA’s GenomeTrakr program: first, contributing timely genomic data for the SARS-CoV-2 pandemic response, and second, establishing both capacity and best practices for culture-independent, population-level environmental surveillance for other pathogens of interest to the FDA. IMPORTANCE This paper serves two primary objectives. First, it summarizes the genomic and contextual data collected during a Covid-19 pandemic response project, which utilized the FDA’s laboratory network, traditionally employed for sequencing foodborne pathogens, for sequencing SARS-CoV-2 from wastewater samples. Second, it outlines best practices for gathering and organizing population-level next generation sequencing (NGS) data collected for culture-free, surveillance of pathogens sourced from environmental samples.

Microbiology↗

SynBio QC Dual Barcode QC (DBC) v1.0

This software was designed as a sequence validation tool for the assembly of synthetic constructs, where the constructs have a high degree of similarity and thus are barcoded prior to the sequencing library prep. It demultiplexes each FASTQ file for each barcode, then analyzes the resulting FASTQ files against a list of reference sequences for that barcode/library, combining the results from eight sequencing libraries to generate a summary, and the files needed to view the results in the Integrative Genomics Viewer (IGV) application for manual verification. This was developed for FASTQ files generated by PacBio sequencing, but could be used on any FASTQ files that do not have paired end reads. It can be used to analyze one - eight libraries at a time. Each construct is independently analyzed with only the sequences with the same barcode, in the same pooled library. Then the results are combined into a user friendly summary. This is used to identify which libraries of pooled sequences contains a perfect match, or fixable match to the reference file. This pipeline uses many freely available open source libraries, the value added is that in our application the steps of the pipeline are defined in Workflow Description Language (WDL) and run through the Cromwell workflow engine in Docker containers, for easy distribution and set up, as well as the user friendly html summary that is generated.

Simirenko, Lisa↗

Multiplexing and Demultiplexing Signals for Radiography Application Using the Discrete Fourier Transform

Our goal is to develop an X-ray phase-contrast imaging system that can provide excellent soft tissue contrast of phase, attenuation, and small-angle scatter. We propose to replace the common system of G0, G1, and G2 gradings with a biprism array to replace the G1 grading and introduce a novel X-ray tube designed to replace the motion of the phase stepping grading G2. The proposed X-ray tube uses temporal multiplexing to provide simultaneous virtual “electronic phase stepping.” In this work the discrete Fourier transform is used to separate from the composite measurement individual X-ray phase contrast measurements sampled at different frequencies. The method performs a discrete Fourier transform of a composite refence sequence to obtain using the frequency amplitudes calibration factors needed to extract the X-ray phase contrast measurement amplitudes from the composite image. The composite reference sequence is the sum of the individual sequences, at different frequencies, with amplitudes of one. The method takes the discrete Fourier transform of this composite reference sequence; whereby, the amplitude of each frequency component is compared with the total sum of its stand-alone sequence amplitude. A calibration factor is determined so that the amplitude of this composite reference frequency times the calibration factor must equal the total sum of the sequence amplitude—the zero-frequency amplitude of the discrete Fourier transform of its stand-alone sequence. To demultiplex the composite measured signal these calibration factors are multiplied by the amplitudes of the frequency components of the discrete Fourier transform of the composite X-phase-contrast measurement to obtain the amplitude of each frequency encoded measurement. Using these calibration factors, we demonstrate with the discrete Fourier transform in Mathematica the extraction of individual images from a composite image that one would expect obtaining from our proposed new X-ray phase contrast imaging system. We then demonstrate as an example how using images from X-ray phase contrast data one can calculate phase, attenuation and the dark field images using grading phase step data supplied to use from Microworks, GmbH in Karlsruhe, Germany.

42 ENGINEERING↗

Populus VariantDB v3.2 facilitates CRISPR and functional genomics research

The success of CRISPR genome editing studies depends critically on the precision of guide RNA (gRNA) design. Sequence polymorphisms in outcrossing tree species pose design hazards that can render CRISPR genome editing ineffective. Despite recent advances in tree genome sequencing with haplotype resolution, sequence polymorphism information remains largely inaccessible to various functional genomics research efforts. The Populus VariantDB v3.2 addresses these challenges by providing a user-friendly search engine to query sequence polymorphisms of heterozygous genomes. The database accepts short sequences, such as gRNAs and primers, as input for searching against multiple poplar genomes, including hybrids, with customizable parameters. We provide examples to showcase the utilities of VariantDB in improving the precision of gRNA or primer design. The platform-agnostic nature of the probe search design makes Populus VariantDB v3.2 a versatile tool for the rapidly evolving CRISPR field and other sequence-sensitive functional genomics applications. The database schema is expandable and can accommodate additional tree genomes to broaden its user base.

59 BASIC BIOLOGICAL SCIENCES↗

Rapid Detection and Quick Characterization of African Swine Fever Virus Using the VolTRAX Automated Library Preparation Platform

African swine fever virus (ASFV) is the causative agent of a severe and highly contagious viral disease affecting domestic and wild swine. The current ASFV pandemic strain has a high mortality rate, severely impacting pig production and, for countries suffering outbreaks, preventing the export of their pig products for international trade. Early detection and diagnosis of ASFV is necessary to control new outbreaks before the disease spreads rapidly. One of the rate-limiting steps to identify ASFV by next-generation sequencing platforms is library preparation. Here, we investigated the capability of the Oxford Nanopore Technologies’ VolTRAX platform for automated DNA library preparation with downstream sequencing on Nanopore sequencing platforms as a proof-of-concept study to rapidly identify the strain of ASFV. Within minutes, DNA libraries prepared using VolTRAX generated near-full genome sequences of ASFV. Thus, our data highlight the use of the VolTRAX as a platform for automated library preparation, coupled with sequencing on the MinION Mk1C for field sequencing or GridION within a laboratory setting. These results suggest a proof-of-concept study that VolTRAX is an effective tool for library preparation that can be used for the rapid and real-time detection of ASFV.

60 APPLIED LIFE SCIENCES↗

Methods for safely sharing dual-use genetic data

Background: Some genetic data has dual-use potential. Sharing pathogen data has shown tremendous value. For example therapeutic development and lineage tracking during the COVID pandemic. This data sharing is complicated by the fact that these data have the potential to be used for harm. The genome sequence of a pathogen can be used to enable malicious genetic engineering approaches or to recreate the pathogen from synthetic DNA. Standard data security methods can be applied to genetic data, but when data is shared between institutions, ensuring appropriate security can be difficult. Sensitive data that is shared internationally among a wide array of institutions can be especially difficult to control. Methods for securely storing and sharing genetic data with potential for dual-use are needed to mitigate this potential harm.Results: Here we propose new methods that allow genetic data to be shared in a data format that prevents a nefarious actor from accessing sensitive aspects of the data. Our methods obfuscate raw sequence data by pooling reads from different samples. This approach can ensure that data is secure while stored and during electronic transfer. We demonstrate that by pooling raw sequence data from multiple samples of the same organism, the ability to fully reconstruct any individual sample is prevented. In the pooled data, most genomic information remains, but reads or mutations cannot be directly attributed to any individual sample. To further restrict access to information, regions of a genome can be removed from the reads.Conclusion: Our methods obscure genomic information within raw sequence reads. This method can allow genetic data to be stored and shared while preventing a nefarious actor from being able to perfectly reconstruct an organism. Broad-scale sequence information remains, while fine scale details about specific samples are difficult or impossible to reconstruct. Our software is available at https://github.com/Geneinfosec-Inc/ReadMixer.

59 BASIC BIOLOGICAL SCIENCES↗

A Route to Design Novel Functional Peptides by Applying a Denoising Diffusional Model to mRNA Display Libraries

In vitro directed evolution techniques, such as mRNA display, enable peptide ligand discovery and optimization. However, physical libraries that rely on a genetic code can only search a small fraction of sequence space due to inherent biases in the genetic code and experimental limitations. To address this challenge, denoising diffusion implicit models (DDIMs) are applied to generate novel peptide ligands against B‐cell lymphoma extra‐large (Bcl‐x L ), a key cancer target. Starting with high‐throughput sequencing data from previous selections, a DDIM is trained to produce novel sequences with high affinity binding. Experimental validation confirms that most generated sequences are functionally equivalent to the original library members for Bcl‐x L binding and demonstrated comparable binding kinetics and affinity relative to the wildtype and nearest original neighbors. Importantly, this approach generated rare sequences not easily accessible via mutation and directed evolution. These results indicate that DDIMs can complement and expand directed evolution data, efficiently exploring underrepresented regions of sequence space. This approach provides a broadly applicable framework for accelerating ligand discovery and optimizing molecular properties across diverse targets.

Qi, Pearl [Mork Family Department of Chemical Engi↗

Apatite geochemistry as a tool for understanding the petrogenesis of layered mafic-ultramafic rocks in the Bushveld Complex, South Africa

The sources of the magmas that formed the Rustenburg Layered Suite of the Bushveld Complex in South Africa remain debated, despite decades of research. Vertical and lateral variation in bulk rock and mineral separate Sr-Nd isotopic compositions, which generally indicate enriched sources, demonstrate that the layered sequence was formed by the emplacement of multiple batches of magma, crucially resulting in episodes of PGE-Cr-V mineralisation. The Lu-Hf isotope compositions of zircon are, however, at odds with the bulk rock Sr-Nd isotopic heterogeneity as they show near homogeneous compositions throughout the layered sequence (εHf (2.06 Ga) =−8). This lack of variation in Hf isotope composition has been attributed to deep, continental lithospheric mantle-related and/or crustal contamination of plume-derived Bushveld magmas. In this study, we analysed the major, trace element and Sr-Nd isotope geochemistry of apatite in the Rustenburg Layered Suite. Apatite occurs as an intercumulus mineral in the lowermost regions and a cumulus mineral in the uppermost regions of the layered sequence and can therefore be used to test existing models for the isotopic disequilibrium between bulk rock Sr-Nd and zircon Hf isotopic compositions. Apatite is largely chlorapatite in the lowermost regions and fluorapatite in the uppermost regions of the layered sequence. The Merensky Reef is unusual in that it comprises both chlorapatite and fluorapatite. Apatite throughout the layered sequence is generally unzoned and shows no evidence of late-stage alteration. Trace element data show that apatite is enriched in L/HREE, with common negative Eu-Sr anomalies. These trace element signatures are consistent with a magmatic origin for the apatite grains, with prior, or concurrent, plagioclase crystallization from the same melt. Variability in in situ Sr and Nd isotope compositions of apatite is recorded throughout the layered sequence with εNd (2.06 Ga) compositions varying between −2.5 and − 10.2 and initial 87 Sr/ 86 Sr compositions varying between 0.7079 and 0.7103 (for the Marikana dikes only). The variability in Sr-Nd isotope compositions of apatite is consistent with the bulk rock (and mineral separate) variation in Sr-Nd isotope compositions, suggesting apatite preserves primary magmatic compositions in the Rustenburg Layered Suite.

Apatite↗

Role of Ribosomal Protein bS1 in Orthogonal mRNA Start Codon Selection

In many bacteria, the location of the mRNA start codon is determined by a short ribosome binding site sequence that base pairs with the 3'-end of 16S rRNA (rRNA) in the 30S subunit. Many groups have changed these short sequences, termed the Shine-Dalgarno (SD) sequence in the mRNA and the anti-Shine-Dalgarno (ASD) sequence in 16S rRNA, to create "orthogonal" ribosomes to enable the synthesis of orthogonal polymers in the presence of the endogenous translation machinery. However, orthogonal ribosomes are prone to SD-independent translation. Ribosomal protein bS1, which binds to the 30S ribosomal subunit, is thought to promote translation initiation by shuttling the mRNA to the ribosome. Thus, a better understanding of how the SD and bS1 contribute to start codon selection could help efforts to improve the orthogonality of ribosomes. Here, we engineered the Escherichia coli ribosome to prevent binding of bS1 to the 30S subunit and separate the activity of bS1 binding to the ribosome from the role of the mRNA SD sequence in start codon selection. We find that ribosomes lacking bS1 are slightly less active than wild-type ribosomes in vitro. Furthermore, orthogonal 30S subunits lacking bS1 do not have an improved orthogonality. Our findings suggest that mRNA features outside the SD sequence and independent of binding of bS1 to the ribosome likely contribute to start codon selection and the lack of orthogonality of present orthogonal ribosomes.

59 BASIC BIOLOGICAL SCIENCES↗

Native Chemical Ligation of Peptoid Oligomers

Bioorganic chemists are inspired by natural biopolymers to design peptidomimetic oligomers that can exhibit sequence-structure-function relationships. Biomimetic polymers can be synthesized to incorporate a specific sequence of nonbiological monomer units using a variety of iterative solution-phase or solid-phase reaction schemes. These protocols generally provide access to a vast diversity of oligomeric compounds but are limited with respect to their ability to attain protein-like chain lengths. This constraint can preclude access to sequence-defined synthetic macromolecules with sufficient sizes required to exhibit tertiary structure and other protein-mimetic attributes. In contrast, peptide chemists have overcome this limitation by developing convergent synthetic methods, such as native chemical ligation, to join individual, smaller peptide chains together to make larger peptides or full proteins. A similar convergent approach is needed to establish efficient synthetic routes to non-natural sequence-defined macromolecules. Herein, we adapt the peptide native chemical ligation method to peptoid oligomers, demonstrating how short chains can be conjoined to create sequence-defined peptoid macromolecules. Nanosheet-forming peptoid polymers with distinct surface loop display domains were generated by sequential ligation of several discrete fragments. This method provides a reliable convergent ligation route for sequence-defined polypeptoids that results in a native amide bond joining the fragments. We envision that this strategy will be useful in synthesizing peptoid-based proteomimetics that incorporate diverse chemical features.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗