Search NASASearch

SEARCH · Search NASA

Results for “Protein function predictions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Assessing the potential of deep learning for protein–ligand docking

The effects of ligand binding on protein structures and their in vivo functions carry numerous implications for modern biomedical research and biotechnology development efforts such as drug discovery. Although several deep learning (DL) methods and benchmarks designed for protein–ligand docking have recently been introduced, so far no previous works have systematically studied the behaviour of the latest docking and structure prediction methods within the broadly applicable context of: (1) using predicted (apo) protein structures for docking (for example, for applicability to new proteins); (2) binding multiple (cofactor) ligands concurrently to a given target protein (for example, for enzyme design); and (3) having no previous knowledge of binding pockets (for example, for generalization to unknown pockets). To enable a deeper understanding of the real-world utility of docking methods, we introduce PoseBench, a comprehensive benchmark for broadly applicable protein–ligand docking. PoseBench enables researchers to rigorously and systematically evaluate DL methods for apo-to-holo protein–ligand docking and protein–ligand structure prediction using both primary ligand and multiligand benchmark datasets, the latter of which we introduce to the DL community. Empirically, using PoseBench, we find that: (1) DL cofolding methods generally outperform comparable conventional and DL docking baseline algorithms, but popular methods such as AlphaFold 3 are still challenged by prediction targets with new protein–ligand binding poses; (2) certain DL cofolding methods are highly sensitive to their input multiple sequence alignments, whereas others are not; and (3) DL methods struggle to strike a balance between structural accuracy and chemical specificity when predicting new or multiligand protein targets.

Morehead, Alex [Lawrence Berkeley National Laborat

A goldilocks computational protocol for inhibitor discovery targeting DNA damage responses including replication-repair functions

While many researchers can design knockdown and knockout methodologies to remove a gene product, this is mainly untrue for new chemical inhibitor designs that empower multifunctional DNA Damage Response (DDR) networks. Here, we present a robust Goldilocks (GL) computational discovery protocol to efficiently innovate inhibitor tools and preclinical drug candidates for cellular and structural biologists without requiring extensive virtual screen (VS) and chemical synthesis expertise. By computationally targeting DDR replication and repair proteins, we exemplify the identification of DDR target sites and compounds to probe cancer biology. Our GL pipeline integrates experimental and predicted structures to efficiently discover leads, allowing early-structure and early-testing (ESET) experiments by many laboratories. By employing an efficient VS protocol to examine protein-protein interfaces (PPIs) and allosteric interactions, we identify ligand binding sites beyond active sites, leveraging in silico advances for molecular docking and modeling to screen PPIs and multiple targets. A diverse 3,174 compound ESET library combines Diamond Light Source DSI-poised, Protein Data Bank fragments, and FDA-approved drugs to span relevant chemotypes and facilitate downstream hit evaluation efficiency for academic laboratories. Two VS per library and multiple ranked ligand binding poses enable target testing for several DDR targets. This GL library and protocol can thus strategically probe multiple DDR network targets and identify readily available compounds for early structural and activity testing to overcome bottlenecks that can limit timely breakthrough drug discoveries. By testing accessible compounds to dissect multi-functional DDRs and suggesting inhibitor mechanisms from initial docking, the GL approach may enable more groups to help accelerate discovery, suggest new sites and compounds for challenging targets including emerging biothreats and advance cancer biology for future precision medicine clinical trials.

59 BASIC BIOLOGICAL SCIENCES

Advancing Protein Display on Bacterial Spores through an Extensive Survey of Coat Components

The profound stability of bacterial spores makes them a promising platform for biotechnological applications like biocatalysis, bioremediation, drug delivery, etc. However, though the Bacillus subtilis spore is composed of >40 types of proteins, only ∼12 have been explored as fusion carriers for protein display. Here, we assessed the suitability of 33 spore proteins (SPs) as enzyme display carriers by direct allele tagging at native genomic loci. Of the 33 SPs investigated, 26 formed functional fusions with β-glucuronidase (GUS)─a ∼272 kDa homotetramer. This almost triples the number of SPs assessed for enzyme display and doubles the number of functional fusions documented in the literature. We quantitatively assessed 1) SP promoter activation dynamics, 2) GUS activity on spores, 3) surface availability, and 4) protection from thermal and proteolytic degradation. Multicopy expression and pairwise coexpression of the most promising SP-GUS fusions highlighted the complexity of spore structure/assembly and the difficulty in predicting compatibility between different SP fusions. We also assessed the suitability of engineered spores to degrade PET (polyethylene terephthalate) films and found that surface-exposed SPs were most effective. Beyond the broad survey, a key outcome of our work was the identification of SscA (small spore coat assembly protein A) as an effective spore display carrier. SscA supported enzyme activity at least 4-fold higher than any other SP, including the well-established anchor, CotY. We attribute this to its promoter, which demonstrated early and sustained activation relative to other SPs and its small size (∼3 kDa), which likely minimally interferes with enzyme folding, oligomerization, and activity. Labeling and genetic studies, its hydrophobic nature, and low surface availability suggest that SscA assembles within the inner spore coat, which makes it stabilizing and suitable for many biocatalytic applications. Overall, this work serves as a knowledge base to advance the biotechnological utility of B. subtilis spores.

Bacillus subtilis

A minimal complex of KHNYN and zinc-finger antiviral protein binds and degrades single-stranded RNA

Detecting viral infection is a key role of the innate immune system. The genomes of some RNA viruses have a high CpG dinucleotide content relative to most vertebrate cell RNAs, making CpGs a molecular marker of infection. The human zinc-finger antiviral protein (ZAP) recognizes CpG, mediates clearance of the foreign CpG-rich RNA, and causes attenuation of CpG-rich RNA viruses. While ZAP binds RNA, it lacks enzymatic activity that might be responsible for RNA degradation and thus requires interacting cofactors for its function. One of these cofactors, KHNYN, has a predicted nuclease domain. Using biochemical approaches, we found that the KHNYN NYN domain is a single-stranded RNA ribonuclease that does not have sequence specificity and digests RNA with or without CpG dinucleotides equivalently in vitro. We show that unlike most KH domains, the KHNYN KH domain does not bind RNA. Indeed, a crystal structure of the KH region revealed a double-KH domain with a negatively charged surface that accounts for the lack of RNA binding. Rather, the KHNYN C-terminal domain (CTD) interacts with the ZAP RNA-binding domain (RBD) to provide target RNA specificity. We define a minimal complex composed of the ZAP RBD and the KHNYN NYN-CTD and use a fluorescence polarization assay to propose a model for how this complex interacts with a CpG dinucleotide-containing RNA. In the context of the cell, this module would represent the minimum ZAP and KHNYN domains required for CpG-recognition and ribonuclease activity essential for attenuation of viruses with clusters of CpG dinucleotides.

Yeoh, Zoe C. (ORCID:0000000226949068)

Structural and biochemical basis for regiospecificity of the flavonoid glycosyltransferase UGT95A1

Glycosylation is a predominant strategy plants use to fine-tune the properties of small molecule metabolites to affect their bioactivity, transport, and storage. It is also important in biotechnology and medicine as many glycosides are utilized in human health. Small molecule glycosylation is largely carried out by family 1 glycosyltransferases. Here, we report a structural and biochemical investigation of UGT95A1, a family 1 GT enzyme from Pilosella officinarum that exhibits a strong, unusual regiospecificity for the 3'-O position of flavonoid acceptor substrate luteolin. We obtained an apo crystal structure to help drive the analyses of a series of binding site mutants, revealing that while most residues are tolerant to mutations, key residues M145 and D464 are important for overall glycosylation activity. Interestingly, E347 is crucial for maintaining the strong preference for 3'-O glycosylation, while R462 can be mutated to increase regioselectivity. The structural determinants of regioselectivity were further confirmed in homologous enzymes. Our study also suggests that the enzyme contains large, highly dynamic, disordered regions. We showed that while most disordered regions of the protein have little to no implication in catalysis, the disordered regions conserved among investigated homologs are important to both the overall efficiency and regiospecificity of the enzyme. This report represents a comprehensive in-depth analysis of a family 1 GT enzyme with a unique substrate regiospecificity and may provide a basis for enzyme functional prediction and engineering.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

A Trichomonas vaginalis C2-XYPPX-repeat protein with a structured C2 domain displaying dampened flexibility upon binding calcium

C2 domains are ubiquitous membrane-binding modules of ∼130 residues in eukaryotes that are often associated with proteins involved in membrane trafficking and lipid modification. The genome of Trichomonas vaginalis, the most common, non-viral, sexually transmitted human pathogen, encodes eight genes that contain a N-terminal C2 module linked to a XYPPX-repeat domain of more than four XYPPX repeats (C2-XYPPX). While the function of the XYPPX-repeat domain remains unknown, its multiple association with C2 domains in T. vaginalis suggests it is important. Here, the C2 domain from one of these C2-XYPPX-repeat proteins, Tv-C2-1, was structurally and physically characterized using X-ray crystallography and NMR spectroscopy. The crystal structure for Tv-C2-1 shows that this domain shares a fold common to all C2 domains, a compact Greek-key motif composed of eight anti-parallel β-strands in the type-2 topology. An NMR chemical shift perturbation study with Ca 2+ showed that Tv-C2-1 bound two Ca 2+ atoms primarily via two loops (loop-1 and loop-3) on the predicted calcium binding face of the protein with K d s of 58.0 ± 0.1 μM and 232 ± 6 μM. Estimations of the overall rotational correlation time, τ c , in the apo (11.1 ns) and Ca 2+ -bound (9.2 ns) state suggests the protein becomes more compact upon Ca 2+ binding, consistent with a decrease in dynamics in loop-3 and marginally in loop-1 suggested by amide 15 N heteronuclear steady-state { 1 H}- 15 N NOEs. Showing Tv-C2-1 binds calcium and adopts a compact Greek-key motif structure, two primary features of C2 domains, suggests understanding the function of the XYPPX-repeat domain may be warranted.

NMR spectroscopy

Leveraging High-resolution Molecular Composition of Soil Organic Matter to Enhance Carbon Cycling Modeling

Soils store more carbon than the atmosphere and vegetation combined, yet Earth system models still struggle to predict how this vast reservoir will respond to environmental change. A central limitation is that most soil biogeochemical models represent organic matter using bulk conceptual pools or chemically homogeneous fractions, preventing direct use of rapidly expanding molecular-scale datasets. Here we develop and test a new soil decomposition framework that explicitly integrates high-resolution information on organic matter composition. First, we construct a molecularly informed litter decomposition module in which plant inputs are partitioned into five functional compound classes—carbohydrates, proteins, lignin-like aromatics, lipids, and carbonyls—using a molecular mixing model calibrated to solid-state 13 C Nuclear Magnetic Resonance (NMR) spectra. Class-specific kinetics, lignin-dependent physical protection, and substrate-driven microbial carbon use efficiency allow the module to capture metabolic tradeoffs associated with enzyme production and nutrient limitation. We then embed this litter module within a microbially explicit whole-soil model that tracks the transformation of these compound classes through particulate organic matter, dissolved organic matter, mineral-associated organic matter, and microbial biomass. High-resolution Fourier Transform Ion Cyclotron Resonance mass spectrometry (FTICR-MS) data are used to link internal pools to measurable soil organic matter fractions and to constrain key process parameters. Applications at soil-core and ecosystem scales demonstrate that the new model reproduces observed soil respiration dynamics while providing mechanistic attribution of CO 2 fluxes to specific chemical classes and pools. Compared to existing frameworks such as the Community Land Model soil biogeochemistry module and the Millennial model, our approach maintains competitive predictive skill while substantially improving interpretability and opportunities for data–model integration. This work illustrates a viable pathway for leveraging molecular-scale observations to reduce structural uncertainty in soil carbon–climate feedback projections.

54 ENVIRONMENTAL SCIENCES

The Factors Governing Metal Dependence of an Emergent Superfamily of Bimetallic Oxygenases

Metalloenzyme superfamilies are typically defined by their protein scaffolds and active sites. Owing to the high tunability of protein structures, members of a single superfamily can catalyze diverse reactions with the same metallocofactor. Some superfamilies, such as amidohydrolase-related dinuclear oxygenases (AROs), display further versatility by utilizing multiple metallocofactors. We have shown that certain AROs catalyze monooxygenation reactions with diiron, dimanganese, and/or mixed manganese−iron cofactors, but the molecular factors governing the selection of a particular cofactor remain unknown, and the extent of this superfamily in biology is unclear. Here, we report bioinformatic analyses that expand the ARO superfamily to approximately 17,000 unique UniProt sequences, far exceeding the number of previously characterized enzymes. Through the integration of structural, spectroscopic, and thermodynamic analyses of representative proteins with a bioinformatic pipeline that identifies key secondary- and tertiary-sphere residues, we can predict in silico the metal preference for the majority of reported ARO sequences. These annotations were validated via the characterization of multiple new AROs, including ones implicated in key oxidative steps of natural product biosyntheses. This study establishes the key structure−function relationships governing metal preferences in AROs and highlights their vastly underappreciated role in myriad biological processes.

Liu, Chang [University of California, Berkeley, CA

Fast myosin binding protein C knockout in skeletal muscle alters length-dependent activation and myofilament structure

In striated muscle, the sarcomeric protein myosin-binding protein-C (MyBP-C) is bound to the myosin thick filament and is predicted to stabilize myosin heads in a docked position against the thick filament, which limits crossbridge formation. Here, we use the homozygous Mybpc2 knockout (C2 -/- ) mouse line to remove the fast-isoform MyBP-C from fast skeletal muscle and then conduct mechanical functional studies in parallel with small-angle X-ray diffraction to evaluate the myofilament structure. We report that C2 -/- fibers present deficits in force production and calcium sensitivity. Structurally, passive C2 -/- fibers present altered sarcomere length-independent and -dependent regulation of myosin head conformations, with a shift of myosin heads towards actin. At shorter sarcomere lengths, the thin filament is axially extended in C2 -/- , which we hypothesize is due to increased numbers of low-level crossbridges. These findings provide testable mechanisms to explain the etiology of debilitating diseases associated with MyBP-C.

59 BASIC BIOLOGICAL SCIENCES

Editorial: Structure and mechanism of microbial membrane active transporters

Membrane active transporters play essential roles in microbial physiology. They couple energy transduction to conformational changes that drive translocation of nutrients, substrates and ions, as well as molecular communication. The structure and function of microbial membrane active transporters are highly diverse. Typical examples include the primary active transporters in the ATP-binding cassette (ABC) superfamily (Thomas and Tampé, 2020; Davidson et al., 2008; Locher et al., 2002), the secondary active transporters in the Major Facilitator Superfamily (MFS) (Drew et al., 2021; Kaback and Guan, 2019), and the ligand-gated porins in the TonB-dependent transporter (TBDT) family (Klebba et al., 2021). As structural, proteogenomic, and computational methods advance, active transporters are increasingly recognized as dynamic molecular machines whose mechanisms can now be visualized and modeled with remarkable precision, building on decades of biochemical and biophysical discovery that established the foundations of this field. The transporter studies recruited in this Research Topic provide us with new insights into the field including structure-function of sugar transporters in yeast, structural prediction and classification of ABC complexes in Bacillus subtilis, Type VI Secretion System (T6SS) in Bacteroides fragilis, amino acids uptake in Escherichia coli and bacterial spore germination.

mechanism

A characterization of recombinant Arabidopsis FRIABLE1 (FRB1) reveals robust rhamnogalacturonan-I rhamnosyltransferase activity and critical catalytic residues

Plant cell walls are glycan-rich extracellular matrices that fundamentally impact essential cellular processes, such as growth, adhesion, and cell shape acquisition. Understanding plant cell wall glycans requires the identification and characterization of the biosynthetic enzymes that produce these polymers. Most successful in vitro protein expression studies of plant cell wall glycosyltransferases have relied on insect, fungal/yeast, or human cell expression systems, whereas prokaryotic expression systems have been generally unsuccessful. Here, we show that Arabidopsis FRIABLE1 (FRB1)/rhamnogalacturonan-I rhamnosyltransferase 8 (RRT8) can be produced in Escherichia coli RosettaGami2 cells as N-terminal maltose-binding protein fusion proteins containing C-terminal 6X-His-tags. We also report the catalytic constants of FRB1/RRT8 with apparent K M and K cat values of 226 μM and 33 min -1 for UDP-Rhamnose and 117 μM and 28.7 min -1 for rhamnogalacturonan-I (RG-I), respectively. We examine the catalytic activities of mutated FRB1/RRT8 proteins based on an AlphaFold 3-generated FRB1/RRT8 protein structural model with a virtually docked UDP-Rha donor. Enzymatic characterization of the mutated and wildtype FRB1/RRT8 protein confirmed that mutation of predicted catalytic site amino acid residues resulted in a 20-fold reduction in RRT activity. FRB1 also robustly polymerizes RG-I in combination with RG-I galacturonosyltransferase 1. These results show how a robust E. coli expression system combined with artificial intelligence tools can be used to increase understanding of plant cell wall glycosyltransferase structure and function.

glycosyltransferase

Cohort-based pan-cancer analysis and experimental studies reveal ISG15 gene as a novel biomarker for prognosis and immunotherapy efficacy prediction

Abstract ISG15, an interferon-stimulated ubiquitin-like protein, plays a multifaceted role in tumorigenesis and immune regulation. This study comprehensively evaluates ISG15 as a prognostic biomarker and predictor of immunotherapy response through pan-cancer bioinformatics analysis and experimental validation. By integrating multiomics data from TCGA, GEO, and clinical cohorts, we found that ISG15 is significantly overexpressed in multiple cancers and generally correlates with poor prognosis. Elevated ISG15 expression is associated with increased immune checkpoint gene expression, particularly PD-L1, and immune infiltration, notably M2-like tumor-associated macrophages. Immunohistochemistry and multiplexed immunofluorescence confirmed a strong positive correlation between ISG15, PD-L1, and M2-TAM infiltration in lung and gastric cancer samples. Functional analysis at the single-cell level revealed significant associations between ISG15 and tumor proliferation, angiogenesis, and immune suppression. Immunotherapy cohort analysis demonstrated that tumors with high ISG15 expression responded favorably to PD-L1 inhibitors but exhibited resistance to CTLA-4 blockade, findings further validated in lung cancer patients receiving anti-PD-1 therapy. These results suggest that ISG15 is a promising biomarker for prognosis and immunotherapy response prediction across cancers. Its integration into clinical decision-making may enhance personalized treatment strategies, improve immunotherapy outcomes, and provide new insights into the tumor immune microenvironment, cancer progression, and potential therapeutic targets for future drug development.

Immunology

Novosphingobium aromaticivorans LigR coordinates transcription of genes involved in metabolism of multiple types of aromatics

Aromatic compounds are a ubiquitous and diverse family of chemicals with functions as biomolecules, natural products, industrial chemicals, and pollutants. Novosphingobium aromaticivorans DSM 12444 uses multiple inducible pathways to catabolize H-, G-, and S-type aromatics that contain zero, one, or two methoxy groups, respectively. Here, we obtain a systems-level view of the transcriptional control of its aromatic metabolic pathways. Several in vitro analyses found that a N. aromaticivorans homolog of the Sphingobium lignivorans SYK-6 transcription factor LigR bound genomic DNA upstream of genes involved in metabolism of multiple aromatic types. We found that a ΔLigR mutant had growth defects on all three types of aromatics as sole carbon sources. Transcriptomic analysis revealed that LigR was required to increase expression of gene products that function in metabolism of all three aromatic types. We also found that, in media containing both glucose and an aromatic carbon source, the ΔLigR mutant directed intermediates through alternative aromatic metabolic pathways. Protein-DNA binding assays showed that N. aromaticivorans LigR binds immediately upstream of promoters of genes involved in aromatic metabolism. We found that N. aromaticivorans LigR coordinates the expression of enzymes that function in the catabolism of H-, G-, and S-type aromatics, and that there are differences in the role of LigR in N. aromaticivorans and S. lignivorans. A comparative genomic analysis predicted that LigR homologs and the aromatic-metabolizing genes that it directly regulates are often co-localized in the genomes of Sphingomonadales, but often not found in this arrangement in many other known aromatic metabolizing bacteria.

Aromatic Compound Degradation

Protein Data Bank (PDB): Fifty-three years young and having a transformative impact on science and society

This review article describes the co-evolution of structural biology as a discipline and the Protein Data Bank (PDB), established in 1971 as the first open-access data resource in biology by like-minded structural scientists. As the PDB archive grew in size and scope to encompass macromolecular crystallography, NMR spectroscopy, and cryo-electron microscopy, new technologies were developed to ingest, validate, curate, store, and distribute the information. Community engagement ensured that the needs of structural biologists (data depositors) and data consumers were met. Today, the archive houses more than 230,000 experimentally determined structures of proteins, nucleic acids, and macromolecular machines and their complexes with one another and small-molecule ligands. Aggregate costs of PDB data preservation are ~1% of the cost of structure determination. The enormous impact of PDB data on basic and applied research and education across the natural and medical sciences is presented and highlighted with illustrative examples. Enablement of de novo protein structure prediction (AlphaFold2, RoseTTAfold, OpenFold, etc.) is the most widely appreciated benefit of having a corpus of rigorously validated, expertly curated 3D biostructure data.

bioinformatics

The protein structurome of Orthornavirae and its dark matter

Metatranscriptomics is uncovering more and more diverse families of viruses with RNA genomes comprising the viral kingdom Orthornavirae in the realm Riboviria. Thorough protein annotation and comparison are essential to get insights into the functions of viral proteins and virus evolution. In addition to sequence- and hmm profile-based methods, protein structure comparison adds a powerful tool to uncover protein functions and relationships. We constructed an Orthornavirae “structurome” consisting of already annotated as well as unannotated (“dark matter”) proteins and domains encoded in viral genomes. We used protein structure modeling and similarity searches to illuminate the remaining dark matter in hundreds of thousands of orthornavirus genomes. The vast majority of the dark matter domains showed either “generic” folds, such as single α-helices, or no high confidence structure predictions. Nevertheless, a variety of lineage-specific globular domains that were new either to orthornaviruses in general or to particular virus families were identified within the proteomic dark matter of orthornaviruses, including several predicted nucleic acid-binding domains and nucleases. In addition, we identified a case of exaptation of a cellular nucleoside monophosphate kinase as an RNA-binding protein in several virus families. Notwithstanding the continuing discovery of numerous orthornaviruses, it appears that all the protein domains conserved in large groups of viruses have already been identified. The rest of the viral proteome seems to be dominated by poorly structured domains including intrinsically disordered ones that likely mediate specific virus-host interactions.

59 BASIC BIOLOGICAL SCIENCES

Regulation of sarcomere formation and function in the healthy heart requires a titin intronic enhancer

Heterozygous truncating variants in the sarcomere protein titin (TTN) are the most common genetic cause of heart failure. To understand mechanisms that regulate abundant cardiomyocyte (CM) TTN expression, we characterized highly conserved intron 1 sequences that exhibited dynamic changes in chromatin accessibility during differentiation of human CMs from induced pluripotent stem cells (hiPSC-CMs). Homozygous deletion of these sequences in mice caused embryonic lethality, whereas heterozygous mice showed an allele-specific reduction in Ttn expression. A 296 bp fragment of this element, denoted E1, was sufficient to drive expression of a reporter gene in hiPSC-CMs. Deletion of E1 downregulated TTN expression, impaired sarcomerogenesis, and decreased contractility in hiPSC-CMs. Site-directed mutagenesis of predicted binding sites of NK2 homeobox 5 (NKX2-5) and myocyte enhancer factor 2 (MEF2) within E1 abolished its transcriptional activity. In embryonic mice expressing E1 reporter gene constructs, we validated in vivo cardiac-specific activity of E1 and the requirement for NKX2-5- and MEF2-binding sequences. Moreover, isogenic hiPSC-CMs containing a rare E1 variant in the predicted MEF2-binding motif that was identified in a patient with unexplained dilated cardiomyopathy (DCM) showed reduced TTN expression. Together, these discoveries define an essential, functional enhancer that regulates TTN expression. Manipulation of this element may advance therapeutic strategies to treat DCM caused by TTN haploinsufficiency.

Kim, Yuri

Enhancing chemical bioproduction with rational control of bacterial post-translational modifications

Efficient conversion of inexpensive feedstocks to valuable chemicals by microbes is critical for a robust bioeconomy, but the ability to rationally design bacteria is hampered by insufficient knowledge of how post translational modifications (PTMs) control bacterial protein function and thus bioproduction phenotypes. Our study will focus on the lysine acetylation, a ubiquitous bacterial PTM that can affect the function of enzymes in central metabolism that are often critical for bioproduction processes, disrupt transcriptional regulation, and reduce translation. However, most lysine acetylation data is observational, which means that we do not know when, how, and what specific acetylated residues affect protein function and bacterial physiology. For our model host, we will use a Pseudomonas putida strain that we previously engineered to convert lignocellulosic feedstocks into chemicals such as itaconic acid (ITA). With this strain, we use a dynamic two-stage bioproduction process in which ITA is produced during a non-growth associated production phase. Production is highest during growth stages when lysine acetylation is low in other organisms (early stationary phase) and stalls in conditions where acetylation is highest (late stationary phase). The switch from high to stalled ITA production is also correlated with an unexpected increase in acetate levels – the precursor to non-enzymatic lysine acetylation. As such, we predict that lysine acetylation plays a substantial role in regulating the metabolic pathways required for ITA production. We will develop a generalizable approach that combines high-throughput genetic screens and cutting-edge genome engineering with state-of-the-art proteomics, metabolomics, and genetic code expansion methods to identify and modulate lysine acetylation patterns in bacteria. Ultimately, these strategies aim to manipulate protein expression and acetylation patterns to enhance bioproduction phenotypes (e.g., sustained ITA production in late stationary phase).

60 APPLIED LIFE SCIENCES

The interplay of DNA repair context with target sequence predictably biases Cas9-generated mutations

Abstract Repair of double-stranded breaks generated by CRISPR/Cas9 is highly dependent on the flanking DNA sequence. To learn about interactions between DNA repair and target sequence, we measure frequencies of over 236,000 distinct Cas9-generated mutational outcomes at over 2800 synthetic target sequences in 18 DNA repair deficient mouse embryonic stem cells lines. We classify the outcomes in an unbiased way, finding a specialised role forPrkdc(DNA-PKcs protein) andPolmin creating 1 bp insertions matching the nucleotide on the protospacer-adjacent motif side of the break, a variable involvement ofNbnandPolqin the creation of different deletion outcomes, and uni-directional deletions dependent on both end-protection and end-resection. Using our dataset, we build predictive models of the mutagenic outcomes of Cas9 scission that outperform the current standards. This work improves our understanding of DNA repair gene function, and provides avenues for more precise modulation of Cas9-generated mutations.

Science & Technology - Other Topics