Conserved sequence features in intracellular domains of viral spike proteins
Not Available
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not Available
Rhodotorula toruloides has been increasingly explored as a host for bioproduction of lipids, fatty acid derivatives and terpenoids. Various genetic tools have been developed, but neither a centromere nor an autonomously replicating sequence (ARS), both necessary elements for stable episomal plasmid maintenance, has yet been reported. In this study, cleavage under targets and release using nuclease (CUT&RUN), a method used for genome-wide mapping of DNA–protein interactions, was used to identify R. toruloides IFO0880 genomic regions associated with the centromeric histone H3 protein Cse4, a marker of centromeric DNA. Fifteen putative centromeres ranging from 8 to 19 kb in length were identified and analyzed, and four were tested for, but did not show, ARS activity. These centromeric sequences contained below average GC content, corresponded to transcriptional cold spots, were primarily nonrepetitive and shared some vestigial transposon-related sequences but otherwise did not show significant sequence conservation. Future efforts to identify an ARS in this yeast can utilize these centromeric DNA sequences to improve the stability of episomal plasmids derived from putative ARS elements.
Pyrone-2,4-dicarboxylic acid (PDC) is a valuable polymer precursor that can be derived from the microbial degradation of lignin. The key enzyme in the microbial production of PDC is 4-carboxy-2-hydroxymuconate-6-semialdehyde (CHMS) dehydrogenase, which acts on the substrate CHMS. We present the crystal structure of CHMS dehydrogenase (PmdC from Comamonas testosteroni) bound to the cofactor NADP, shedding light on its three-dimensional architecture, and revealing residues responsible for binding NADP. Using a combination of structural homology, molecular docking, and quantum chemistry calculations, we have predicted the binding site of CHMS. Key histidine residues in a conserved sequence are identified as crucial for binding the hydroxyl group of CHMS and facilitating dehydrogenation with NADP. Mutating these histidine residues results in a loss of enzyme activity, leading to a proposed model for the enzyme's mechanism. These findings are expected to help guide efforts in protein and metabolic engineering to enhance PDC yields in biological routes to polymer feedstock synthesis.
The N-heptad repeat (NHR) of the HIV-1 gp41 prehairpin intermediate (PHI) is an attractive potential vaccine target with high sequence conservation across diverse strains. However, despite the potency of NHR-targeting peptides and clinical efficacy of the NHR-targeting entry inhibitor enfuvirtide, no potently neutralizing NHR-directed monoclonal antibodies (mAbs) nor antisera have been identified or elicited to date. The lack of potent NHR-binding mAbs both dampens enthusiasm for vaccine development efforts at this target and presents a barrier to performing passive immunization experiments with NHR-targeting antibodies. To address this challenge, we previously developed an improved variant of the NHR-directed mAb D5, called D5_AR, which is capable of neutralizing diverse tier-2 viruses. Building on that work, here we present the 2.7Å-crystal structure of D5_AR bound to NHR mimetic peptide IQN17. We then utilize protein language models and supervised machine learning to generate small (n < 100) libraries of D5_AR variants that are subsequently screened for improved neutralization potency. We identify a variant with 5-fold improved neutralization potency, D5_FI, which is the most potent NHR-directed monoclonal antibody characterized to date and exhibits broad neutralization of tier-2 and −3 pseudoviruses as well as replicating R5 and X4 challenge strains. Additionally, our work highlights the ability of protein language models to efficiently identify improved mAb variants from relatively small libraries.
Bifurcating enzymes employ energy from a favorable electron transfer to drive unfavorable transfer of a second electron, thereby generating a more reactive product. They are therefore highly desirable in catalytic systems, for example, to drive challenging reactions such as nitrogen fixation. While most bifurcating enzymes contain air-sensitive metal centers, bifurcating electron transfer flavoproteins (bETFs) employ flavins. However, they have not been successfully deployed on electrodes. Herein, we demonstrate immobilization and expected thermodynamic reactivity of a bETF from a hyperthermophilic archaeon, Sulfolobus acidocaldarius (SaETF). SaETF differs from previously biochemically characterized bETFs in being a single protein, representing a concatenation of the two subunits of known ETFs. However, SaETF retains the chemical properties of heterodimeric bETFs, including possession of two FADs: one that undergoes sequential 1-electron (1e) reductions at high E° and forms an anionic semiquinone, and another that is amenable to lower-E° 2e reduction, including by NADH. We found homologous monomeric ETF genes in archaeal and bacterial genomes, accompanied by genes that also commonly flank heterodimeric ETFs, and SaETF’s sequence conservation is 50% higher with bETFs than with canonical ETFs. Thus, SaETF is best described as a bETF. Our direct electrochemical trials capture reversible redox couples for all three thermodynamically expected redox events. We document electrochemical activity over a range of pH values and reveal a conformational change coupled to proton acquisition that affects the electrochemical activity of the higher-E° FAD. Thus, this well-behaved monomeric bETF opens the door to bioinspired bifurcating devices or bifurcation on a chip.
Regulation of Ras GTPases by GTPase-activating proteins (GAPs) is essential for their normal signaling. Nine of the ten GAPs for Ras contain a C2 domain immediately proximal to their canonical GAP domain, and in RasGAP (p120GAP, p120RasGAP;RASA1) mutation of this domain is associated with vascular malformations in humans. Here, we show that the C2 domain of RasGAP is required for full catalytic activity toward Ras. Analyses of the RasGAP C2-GAP crystal structure, AlphaFold models, and sequence conservation reveal direct C2 domain interaction with the Ras allosteric lobe. This is achieved by an evolutionarily conserved surface centered around RasGAP residue R707, point mutation of which impairs the catalytic advantage conferred by the C2 domain in vitro. In mice,R707Cmutation phenocopies the vascular and signaling defects resulting from constitutive disruption of theRASA1gene. In SynGAP, mutation of the equivalent conserved C2 domain surface impairs catalytic activity. Our results indicate that the C2 domain is required to achieve full catalytic activity of GAPs for Ras.
Fyn is a Src-family tyrosine kinase implicated in synaptic dysfunction and neuroinflammation across multiple neurodegenerative disorders, including Alzheimer’s disease (AD) and Parkinson’s disease (PD). Saracatinib (AZD0530) is a potent Src-family inhibitor that has been explored as a repurposed therapeutic; however, its clinical utility is limited by poor kinase selectivity caused by high sequence conservation within Src-family ATP-binding sites. Here, we combine surface plasmon resonance (SPR) and X-ray crystallography to define saracatinib recognition by the Fyn kinase domain (KD). SPR single-cycle kinetics shows that saracatinib binds the isolated Fyn KD and full-length Fyn with low-nanomolar affinity, whereas dasatinib binds with subnanomolar affinity and markedly slower dissociation. We determined the crystal structure of the Fyn KD-saracatinib complex at 2.22 Å resolution. The kinase adopts an active-like conformation with the DFG motif and αC-helix in the ‘in’ state and a conserved β3 αC Lys-Glu salt bridge. Saracatinib occupies the adenine and ribose pockets, and engages the hinge through direct and water-mediated hydrogen bonding while complementing a hydrophobic back pocket by van der Waals contacts. Comparison with reported saracatinib-bound structures of other kinases suggests that the active-state geometry observed for Fyn creates a pocket not observed in inactive-like complexes, providing a structural handle for designing Fyn-selective inhibitors. Comparison with all saracatinib-bound kinase co-structures currently available in the PDB (ALK2 and PKMYT1) indicates a conserved monodentate hinge binding mode but kinase-dependent αC-helix conformations, providing a structural rationale for designing Fyn-selective analogues.
Conserved non-coding sequences (CNSs) are integral elements of transcriptional regulation. Transcriptional tuning of PLETHORA (PLT) genes that encode master regulators of plant development is vital for embryogenesis and meristematic function. However, how the expression of PLT genes is modulated through CNSs remains unclear. Through motif-based mining of upstream sequences in 120 angiosperm genomes, we identified 21 conserved and lineage-specific CNSs, two of which are unusually long, similar, and colinear within eudicots. Using Arabidopsis thaliana, we demonstrate that these two deeply conserved elements, which we named BOX1 and BOX2, control PLT1 and PLT2 expression. CRISPR mutants within these elements specifically reduced PLT expression levels, and reporter lines revealed that deletion of either or both BOXes altered and/or abrogated the PLT2 expression pattern in the root tip, affecting the ability to rescue the plt1 plt2 double mutant. We further show that the influence of these elements on expression patterns is already exerted during embryogenesis and functional in the context of the early embryo. Finally, we reveal the existence of a BOX-mediated autoregulatory feedback loop that, in large part, explains CNS influence on expression patterns. We thus uncover a transcriptional mechanism by which genes encoding master regulators of embryo and root meristem development are regulated.
Gene regulatory evolution is a well-known source of phenotypic diversity and adaptive evolution. Although cis-regulatory elements (CREs) play a vital role in gene expression evolution, the molecular evolution of CREs remains mostly unknown due to the difficulty in identifying and characterizing these functional elements. Comparative genomic analyses of noncoding DNA can be leveraged to identify conserved noncoding sequences (CNS), many of which may harbor functional CREs conserved by purifying selection. However, purely computational inference of CREs from putative CNS can be erroneous due to the complex genomic architecture in plants. One promising experimental approach to identify CREs is by profiling accessible chromatin regions (ACRs) that are often associated with the location of CREs. In this study, we use comparative genomics along with the profiling of ACRs to study the molecular evolution of putative functional noncoding regulatory regions in Panicoid grasses. We identified sets of CNS that varied in relationship to the degree of evolutionary divergence among the studied taxa, including identifying core-Panicoid-CNS. We augmented this analysis by profiling ACRs in Panicum hallii ecotypes using ATAC-seq. ACRs had low SNP density at the summit, harbored a high frequency of core-Panicoid-CNS, and were enriched with expression QTL. These data help to annotate the P. hallii genome for putative functional elements and suggest that a large proportion of these ACRs are evolving under purifying selection. Turnover in CNS and ACR between ecotypes of P. hallii identifies a small set of putatively divergent CREs that may underlie differences in gene regulation between genotypes from inland and coastal habitats. In summary, we profiled ACRs in Panicoid grasses and integrated this data with our putative CNS prediction framework, which provides unique insight into patterns of polymorphism and divergence in CREs in C4 perennial grasses.
Schizophyllum commune is a mushroom-forming fungus notable for its distinctive fruiting bodies with split gills. It is used as a model organism to study mushroom development, lignocellulose degradation and mating type loci. It is a hypervariable species with considerable genetic and phenotypic diversity between the strains. In this study, we systematically phenotyped 16 dikaryotic strains for aspects of mushroom development and 18 monokaryotic strains for lignocellulose degradation. There was considerable heterogeneity among the strains regarding these phenotypes. The majority of the strains developed mushrooms with varying morphologies, although some strains only grew vegetatively under the tested conditions. Growth on various carbon sources showed strain-specific profiles. The genomes of seven monokaryotic strains were sequenced and analyzed together with six previously published genome sequences. Moreover, the related species Schizophyllum fasciatum was sequenced. Although there was considerable genetic variation between the genome assemblies, the genes related to mushroom formation and lignocellulose degradation were well conserved. These sequenced genomes, in combination with the high phenotypic diversity, will provide a solid basis for functional genomics analyses of the strains of S. commune.
Ubiquilins are molecular chaperones that play multifaceted roles in proteostasis, with point mutations in UBQLN2 leading to altered phase-separation properties and amyotrophic lateral sclerosis (ALS). Our mechanistic understanding of this essential process has been hindered by a lack of structural information on the STI1 domain, which is essential for ubiquilin chaperone activity and phase separation. Here, we present the first crystal structure of a ubiquilin-family STI1 domain bound to a transmembrane domain (TMD), and show that ALS mutations disrupt the STI1-TMD interaction. We further demonstrate that ubiquilins contain multiple conserved internal sequences that bind to the STI1 domain, including the PXX-repeat region that is a hotspot for ALS mutations. We propose that these placeholder sequences prevent solvent exposure of the STI1 hydrophobic groove and contribute to the multivalency that drives ubiquilin phase-separation. Together, this work provides a new paradigm for understanding how STI1 domains modulate ubiquilin chaperone activity and phase separation, and offers insights into the molecular basis of ALS pathogenesis.
Abstract Apidaecin 1b (Api), the first characterized Type II Proline-rich antimicrobial peptide (PrAMP), is encoded in the honey bee genome. It inhibits bacterial growth by binding in the nascent peptide exit tunnel of the ribosome after the release of the completed protein and trapping the release factors. By genome mining, we have identified 71 PrAMPs encoded in insect genomes as pre-pro-polyproteins. Having chemically synthesized and tested the activity of 26 peptides, we demonstrate that despite significant sequence variation in the N-terminal sequence, the majority of the PrAMPs that retain the conserved C-terminal sequence of Api are able to trap the ribosome at the stop codons and induce stop codon readthrough—all hallmarks of Type II PrAMP mode of action. Some of the characterized PrAMPs exhibit superior antibacterial activity in comparison with Api. The newly solved crystallographic structures of the ribosome complexed with Api and with the more active peptide Fva1 from the stingless bee demonstrate the universal placement of the PrAMPs’ C-terminal pharmacophore in the post-release ribosome despite variations in their N-terminal sequence.
Abstract Premise Target sequence capture (Hyb‐Seq) is a cost‐effective sequencing strategy that employs RNA probes to enrich for specific genomic sequences. By targeting conserved low‐copy orthologs, Hyb‐Seq enables efficient phylogenomic investigations. Here, we present Asparagaceae1726—a Hyb‐Seq probe set targeting 1726 low‐copy nuclear genes for phylogenomics in the angiosperm family Asparagaceae—which will aid the often‐challenging delineation and resolution of evolutionary relationships within Asparagaceae. Methods Here we describe and validate the Asparagaceae1726 probe set (https://github.com/bentzpc/Asparagaceae1726) in six of the seven subfamilies of Asparagaceae. We perform phylogenomic analyses with these 1726 loci and evaluate how inclusion of paralogs and bycatch plastome sequences can enhance phylogenomic inference with target‐enriched data sets. Results We recovered at least 82% of target orthologs from all sampled taxa, and phylogenomic analyses resulted in strong support for all subfamilial relationships. Additionally, topology and branch support were congruent between analyses with and without inclusion of target paralogs, suggesting that paralogs had limited effect on phylogenomic inference. Discussion Asparagaceae1726 is effective across the family and enables the generation of robust data sets for phylogenomics of any Asparagaceae taxon. Asparagaceae1726 establishes a standardized set of loci for phylogenomic analysis in Asparagaceae, which we hope will be widely used for extensible and reproducible investigations of diversification in the family.
ABSTRACT Global transcription factors (TFs) control metabolic processes in bacteria to efficiently utilize available carbon. The orderCaldicellulosiruptoraleshas drawn interest due to the ability of its members to degrade components of lignocellulosic biomass. Regulatory reconstruction ofAnaerocellum (f. Caldicellulosiruptor) besciiidentified two major global transcription factors for xylan utilization, XynR and XylR, and the corresponding putative transcription factor binding sites. Recombinant versions of XynR (LacI family) and XylR (ROK family) were subjected to fluorescence polarization (FP) and biolayer interferometry (BLI) analysis to confirm the predicted binding sites. Four XynR sites and two XylR sites were validated, accounting for 20 of 26 genes regulated by XynR and six of seven genes regulated by XylR. Bioinformatic analysis of the individual genes controlled by the two regulators showed an inter-dependent scheme for xylan conversion; the transport of xylooligosaccharides (XOS) is dependent on XylR, while enzymes responsible for hydrolysis are controlled by both regulators. For xylose catabolism by the xylose isomerase-xylulose kinase pathway, regulation is also split, with XylR controlling xylose isomerase and XynR controlling xylokinase. The XynR/XylR regulator pair withinA. besciiis conserved in all sequenced species ofCaldicellulosiruptorales, suggesting similarities in regulating linear xylan conversion. In other xylanolytic thermophiles, XylR homologs control xylan degradation, compared to just 6 out of 26 genes forA. bescii. These results show that two separate regulatory schemes (dual repression) are coordinated byA. besciito effectively regulate the hemicellulose inventory and xylan catabolism. IMPORTANCE To take full advantage of extreme thermophiles as platform metabolic engineering microorganisms, the tools for genetic manipulation must be further developed, and strategies that exploit a better understanding of metabolic regulation need to be discerned.Anaerocellum bescii, the most studied of the extremely thermophilic fermentative anaerobic bacteria that can utilize microcrystalline cellulose, can degrade microcrystalline cellulose and hemicellulose and has been metabolically engineered to convert the resulting sugars to products such as ethanol and acetone. For xylan, in particular, two major global transcription factors (TFs), XynR and XylR, play a role in sugar metabolism, although their predicted regulatory interdependence from bioinformatics analysis has not been elucidated experimentally. Here, fluorescence polarization (FP) and biolayer interferometry (BLI) were used to explore this issue to support metabolic engineering efforts aimed at improving carbohydrate processing to industrial chemicals.
Heterozygous truncating variants in the sarcomere protein titin (TTN) are the most common genetic cause of heart failure. To understand mechanisms that regulate abundant cardiomyocyte (CM) TTN expression, we characterized highly conserved intron 1 sequences that exhibited dynamic changes in chromatin accessibility during differentiation of human CMs from induced pluripotent stem cells (hiPSC-CMs). Homozygous deletion of these sequences in mice caused embryonic lethality, whereas heterozygous mice showed an allele-specific reduction in Ttn expression. A 296 bp fragment of this element, denoted E1, was sufficient to drive expression of a reporter gene in hiPSC-CMs. Deletion of E1 downregulated TTN expression, impaired sarcomerogenesis, and decreased contractility in hiPSC-CMs. Site-directed mutagenesis of predicted binding sites of NK2 homeobox 5 (NKX2-5) and myocyte enhancer factor 2 (MEF2) within E1 abolished its transcriptional activity. In embryonic mice expressing E1 reporter gene constructs, we validated in vivo cardiac-specific activity of E1 and the requirement for NKX2-5- and MEF2-binding sequences. Moreover, isogenic hiPSC-CMs containing a rare E1 variant in the predicted MEF2-binding motif that was identified in a patient with unexplained dilated cardiomyopathy (DCM) showed reduced TTN expression. Together, these discoveries define an essential, functional enhancer that regulates TTN expression. Manipulation of this element may advance therapeutic strategies to treat DCM caused by TTN haploinsufficiency.
The growth of metagenomics-derived amino acid sequence data has transformed our understanding of protein function, microbial diversity, and evolutionary relationships. However, the vast majority of these proteins remain functionally uncharacterized. Grouping the millions of such uncharacterized sequences with the few experimentally characterized ones allows the transfer of annotations, while the inspection of conserved residues with multiple sequence alignments can provide clues to function, even in the absence of existing functional information. To address the challenges associated with this data surge and the need to group sequences, we present a scalable, open-source, parametrizable Nextflow pipeline (nf-core/proteinfamilies) that generates nascent protein families or assigns new proteins to existing families. The computational benchmarks demonstrated that resource usage scales approximately linearly with input size, and the biological benchmarks showed that the generated protein families closely resemble manually curated families in widely used databases.
Biopolymer sequences dictate their functions, and protein-based polymers are a promising platform to establish sequence–function relationships for novel biopolymers. To efficiently explore vast sequence spaces of natural proteins, sequence repetition is a common strategy to tune and amplify specific functions. This strategy is applied to repeats-in-toxin (RTX) proteins with calcium-responsive folding behavior, which stems from tandem repeats of the nonapeptide GGXGXDXUX in which X can be any amino acid and U is a hydrophobic amino acid. To determine the functional range of this nonapeptide, we modified a naturally occurring RTX protein that forms β-roll structures in the presence of calcium. Sequence modifications focused on calcium-binding turns within the repetitive region, including either global substitution of nonconserved residues or complete replacement with tandem repeats of a consensus nonapeptide GGAGXDTLY. Some sequence modifications disrupted the typical transition from intrinsically disordered random coils to folded β rolls, despite conservation of the underlying nonapeptide sequence. Proteins enriched with smaller, hydrophobic amino acids adopted secondary structures in the absence of calcium and underwent structural rearrangements in calcium-rich environments. In contrast, proteins with bulkier, hydrophilic amino acids maintained intrinsic disorder in the absence of calcium. In conclusion, these results indicate a significant role of nonconserved amino acids in calcium-responsive folding, thereby revealing a strategy to leverage sequences in the design of tunable, calcium-responsive biopolymers.
Transcription factors (TFs) bind combinatorially to cis-regulatory elements, orchestrating transcriptional programs. Although studies of chromatin state and chromosomal interactions have demonstrated dynamic neurodevelopmental cis-regulatory landscapes, parallel understanding of TF interactions lags. To elucidate combinatorial TF binding driving mouse basal ganglia development, we integrated chromatin immunoprecipitation sequencing (ChIP-seq) for twelve TFs, H3K4me3-associated enhancer-promoter interactions, chromatin and gene expression data, and functional enhancer assays. We identified sets of putative regulatory elements with shared TF binding (TF-pRE modules) that orchestrate distinct processes of GABAergic neurogenesis and suppress other cell fates. The majority of pREs were bound by one or two TFs; however, a small proportion were extensively bound. These sequences had exceptional evolutionary conservation and motif density, complex chromosomal interactions, and activity as in vivo enhancers. Our results provide insights into the combinatorial TF-pRE interactions that activate and repress expression programs during telencephalon neurogenesis and demonstrate the value of TF binding toward modeling developmental transcriptional wiring.