Search NASA⌕ Search

Engineering topics

Boob, Aashutosh Girish

Publications and source records attributed to Boob, Aashutosh Girish.

CRISPR-COPIES: an in silico platform for discovery of neutral integration sites for CRISPR/Cas-facilitated gene integration

Abstract The CRISPR/Cas system has emerged as a powerful tool for genome editing in metabolic engineering and human gene therapy. However, locating the optimal site on the chromosome to integrate heterologous genes using the CRISPR/Cas system remains an open question. Selecting a suitable site for gene integration involves considering multiple complex criteria, including factors related to CRISPR/Cas-mediated integration, genetic stability, and gene expression. Consequently, identifying such sites on specific or different chromosomal locations typically requires extensive characterization efforts. To address these challenges, we have developed CRISPR-COPIES, a COmputational Pipeline for the Identification of CRISPR/Cas-facilitated intEgration Sites. This tool leverages ScaNN, a state-of-the-art model on the embedding-based nearest neighbor search for fast and accurate off-target search, and can identify genome-wide intergenic sites for most bacterial and fungal genomes within minutes. As a proof of concept, we utilized CRISPR-COPIES to characterize neutral integration sites in three diverse species: Saccharomyces cerevisiae, Cupriavidus necator, and HEK293T cells. In addition, we developed a user-friendly web interface for CRISPR-COPIES (https://biofoundry.web.illinois.edu/copies/). We anticipate that CRISPR-COPIES will serve as a valuable tool for targeted DNA integration and aid in the characterization of synthetic biology toolkits, enable rapid strain construction to produce valuable biochemicals, and support human gene and cell therapy applications.

59 BASIC BIOLOGICAL SCIENCES↗

CRISPR-COPIES: Web Tool

CRISPR/Cas system has emerged as a powerful genome-editing tool for metabolic engineering and human gene therapy. However, the conundrum of where to integrate heterologous genes on the chromosome using the CRISPR/Cas system remains an open question. Selecting a site for gene integration requires incorporation of complex criteria such as factors involved in CRISPR/Cas-mediated integration, genetic stability, and gene expression and therefore, usually requires strenuous characterization of sites on particular or different chromosomal locations. To address these issues, we developed CRISPR-COPIES, a COmputational Pipeline for the Identification of CRISPR/Cas-facilitated intEgration Sites. The tool applies ScaNN, a state-of-the-art model on the embedding-based nearest neighbor search for fast and accurate off-target search and can identify genome-wide intergenic sites for most bacterial and fungal genomes within minutes. This submission contains the code we developed to create a user-friendly web interface for CRISPR-COPIES (https://biofoundry.web.illinois.edu/copies/). We anticipate CRISPR-COPIES will serve as a useful tool for targeted DNA integration and aid in the characterization of synthetic biology toolkits, rapid strain construction to produce valuable biochemicals, and human gene and cell therapy.

Bioinformatics↗

A landing pad system for multicopy gene integration in Issatchenkia orientalis

The robust nature of the non-conventional yeast Issatchenkia orientalis allows it to grow under highly acidic conditions and therefore, has gained increasing interest in producing organic acids using a variety of carbon sources. Recently, the development of a genetic toolbox for I. orientalis, including an episomal plasmid, characterization of multiple promoters and terminators, and CRISPR-Cas9 tools, has eased the metabolic engineering efforts in I. orientalis. However, multiplex engineering is still hampered by the lack of efficient multicopy integration tools. To facilitate the construction of large, complex metabolic pathways by multiplex CRISPR-Cas9-mediated genome editing, we developed a bioinformatics pipeline to identify and prioritize genome-wide intergenic loci and characterized 47 gRNAs located in 21 intergenic regions. These loci are screened for guide RNA cutting efficiency, integration efficiency of a gene cassette, the resulting cellular fitness, and GFP expression level. We further developed a landing pad system using components from these well-characterized loci, which can aid in the integration of multiple genes using single guide RNA and multiple repair templates of the user’s choice. We have demonstrated the use of the landing pad for simultaneous integrations of 2, 3, 4, or 5 genes to the target loci with efficiencies greater than 80%. As a proof of concept, we showed how the production of 5-aminolevulinic acid can be improved by integrating five copies of genes at multiple sites in one step. We have further demonstrated the efficiency of this tool by constructing a metabolic pathway for succinic acid production by integrating five gene expression cassettes using a single guide RNA along with five different repair templates, leading to the production of 9 g/L of succinic acid in batch fermentations. Furthermore, this study demonstrates the effectiveness of a single gRNA-mediated CRISPR platform to build complex metabolic pathways in a non-conventional yeast. This landing pad system will be a valuable tool for the metabolic engineering of I. orientalis.

59 BASIC BIOLOGICAL SCIENCES↗

In vitro continuous protein evolution empowered by machine learning and automation

Directed evolution has become one of the most successful and powerful tools for protein engineering. However, the efforts required for designing, constructing, and screening a large library of variants can be laborious, time-consuming, and costly. With the recent advent of machine learning (ML) in the directed evolution of proteins, researchers can now evaluate variants in silico and guide a more efficient directed evolution campaign. Furthermore, recent advancements in laboratory automation have enabled the rapid execution of long, complex experiments for high-throughput data acquisition in both industrial and academic settings, thus providing the means to collect a large quantity of data required to develop ML models for protein engineering. In this perspective, here we propose a closed-loop in vitro continuous protein evolution framework that leverages the best of both worlds, ML and automation, and provide a brief overview of the recent developments in the field.

59 BASIC BIOLOGICAL SCIENCES↗