Alternative Approach to Sequence-Specific Recognition of DNA: Cooperative Stacking of Dication Dimers─Sensitivity to Compound Curvature, Aromatic Structure, and DNA Sequence
Not provided.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not provided.
Understanding the inherent timescales of large bubbles in DNA is critical to a thorough comprehension of its physicochemical characteristics, as well as their potential role on helix opening and biological function. In this work, we employ the coarse-grained Peyrard–Bishop–Dauxois model of DNA to study relaxation dynamics of large bubbles in homopolymer DNA, using simulations up to the microsecond time scale. By studying energy autocorrelation functions of relatively large bubbles inserted into thermalised DNA molecules, we extract characteristic relaxation times from the equilibration process for both adenine–thymine (AT) and guanine–cytosine (GC) homopolymers. Bubbles of different amplitudes and widths are investigated through extensive statistics and appropriate fittings of their relaxation. Characteristic relaxation times increase with bubble amplitude and width. We show that, within the model, relaxation times are two orders of magnitude longer in GC sequences than in AT sequences. Overall, our results confirm that large bubbles leave a lasting impact on the molecule’s dynamics, for times between 0.5–500 ns depending on the homopolymer type and bubble shape, thus clearly affecting long-time evolutions of the molecule.
Nanopores in thin membranes play important roles in science and industry.
In-context promoter bashing via genome editing is a route to identify and characterize critical regulatory regions that govern expression of genes of interest. The outcomes of in-context promoter bashing can be used to inform editing strategies to modulate the expression of selected gene models in a desired fashion. Here we employed in-context promoter bashing to characterize the proximal upstream regulatory regions of sorghum genes encoding phosphoenolpyruvate carboxykinase (Sb.PEPCK.BS, SbiTx430.01G455400) and alanine aminotransferase (SbiTx430.02G006600, SbAlaAT.BS), two proteins involved in the PCK C4 pathway. Characterized germinal edits within the targeted regions upstream of these two genes ranged in size from 138 bp up to 1790 bp. A 138 bp within the Sb.PEPCK.BS upstream region and a 1643 bp element within the Sb.AlaAT.BS upstream region were determined to be important for maintenance of transcription levels. No change in development or various physiological parameters was observed in characterized lineages carrying promoter edits. However, significant changes in seed reserves and a reduction in 100 seed weight were consistently observed, under both greenhouse and field environments, in plants carrying an edit in the promoter of Sb.PEPCK.BS gene were significantly reduced in transcript accumulation for this gene.
SUMMARY: We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e. DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 28 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. AVAILABILITY AND IMPLEMENTATION: Model, codes, and data are publically available at https://github.com/MAGICS-LAB/DNABERT_S.
The present method of detection involves increasing an amount of analyte molecules by an isothermal molecular amplification approach. In the present approach a starting molecule of interest may be amplified through a reaction it induces with specifically engineered and functionalized particles, namely protected particles A and storage particles B. This reaction may result in a set of output DNA molecules that is larger in number than the input DNA molecules. Thus the reaction between nanoparticles for amplification of a certain DNA sequence (input DNA molecules) may occur when there is a match with a targeted molecule (stored molecules on storage particles B) and if the DNA sequence of the input DNA molecules does not match (partially or completely) the targeted molecule the reaction may not occur. Without a certain molecular input of the input DNA molecule the reaction may not occur.
The present invention provides a method of a method of designing an implementation of a DNA assembly. In an exemplary embodiment, the method includes (1) receiving a list of DNA sequence fragments to be assembled together and an order in which to assemble the DNA sequence fragments, (2) designing DNA oligonucleotides (oligos) for each of the DNA sequence fragments, and (3) creating a plan for adding flanking homology sequences to each of the DNA oligos. In an exemplary embodiment, the method includes (1) receiving a list of DNA sequence fragments to be assembled together and an order in which to assemble the DNA sequence fragments, (2) designing DNA oligonucleotides (oligos) for each of the DNA sequence fragments, and (3) creating a plan for adding optimized overhang sequences to each of the DNA oligos.
Polymeric gels crosslinked by DNA sequences can exploit DNA strand-displacement reactions to promote swelling through dynamic polymerization. The degree of swelling and the rate of swelling must be directly tunable to achieve the promise of programmable soft matter. Though the kinetics of the strand-displacement reaction provide insertion rates up to 10 4 /Molar/second as measured in bulk solution, DNA hydrogel swelling can take upwards of 30 h to complete. Computational modeling of the reaction-induced swelling of these gels with our recently-developed reactive electrochemomechanical theory (Zimmerman et al., 2024) suggests that their extraordinarily slow swelling is partly due to a scaling mismatch between the addition of charge and the addition of fluid volume, leading to a large transient increase in the fixed charge density. The significant increase in the gel’s fixed charge density, due to the binding of negatively charged DNA, sharply restricts the concentration of mobile hairpins through the phenomenon of Donnan charge exclusion, an effect commonly exploited in nanofiltration applications using polymeric membranes. The scaling problem is overcome when the mean additional swelling provided to the hydrogel by addition of a crosslink is above a critical value, thus the swelling outpaces the charge accumulation, leading the fixed charge density to drop and significantly accelerating the swelling process. This study shows that Donnan exclusion can explain the kinetics of DNA hydrogel swelling, and studies ways to modulate the reaction speed by either modifying the salt concentration or increasing or decreasing the number of base pairs in each DNA sequence.
The heterocylic diamidine DB2447 recognizes a single GC base pair in the minor groove of a mixed DNA sequence. The DNA is stabilized by a PU.1 protein bound to the major groove.
DNA-functionalized hydrogels are capable of sensing oligonucleotides, proteins, and small molecules, and specific DNA sequences sensed in the hydrogels’ environment can induce changes in these hydrogels’ shape and fluorescence. Fabricating DNA-functionalized hydrogel architectures with multiple domains could make it possible to sense multiple molecules and undergo more complicated macroscopic changes, such as changing fluorescence or changing the shapes of regions of the hydrogel architecture. However, automatically fabricating multi-domain DNA-functionalized hydrogel architectures, capable of enabling the construction of hydrogel architectures with tens to hundreds of different domains, presents a significant challenge. We describe a platform for fabricating multi-domain DNA-functionalized hydrogels automatically at the micron scale, where reaction and diffusion processes can be coupled to program material behavior. Using this platform, the hydrogels’ material properties, such as shape and fluorescence, can be programmed, and the fabricated hydrogels can sense their environment. DNA-functionalized hydrogel architectures with domain sizes as small as 10 microns and with up to 4 different types of domains can be automatically fabricated using ink volumes as low as 50 μL. We also demonstrate that hydrogels fabricated using this platform exhibit responses similar to those of DNA-functionalized hydrogels fabricated using other methods by demonstrating that DNA sequences can hybridize within them and that they can undergo DNA sequence-induced shape change.
Build Optimization Software Tools (BOOST) accelerate the design of DNA, RNA and protein sequences for their synthesis and assembly in an automated and scalable fashion. The BOOST Juggler automates the following two design tasks: - Reverse-Translation of protein sequences into DNA sequences - Codon Juggling of DNA sequences The BOOST Polisher provides the following design tasks: - The Verification of DNA sequences against DNA synthesis constraints - The Modification of DNA sequences in case they violate any DNA synthesis constraint In addition, the BOOST Polisher can be instructed to verify and modify DNA sequences against sequence patterns, such as restriction sites. The BOOST Partitioner supports both Gibson/chewback and Yeast assembly (also known as Transformation-Associated Recombination, or TAR) methods when performing the decomposition of large sequences that exceed the maximum length of DNA synthesis into synthesizable building blocks. The Partitioner can be instructed to find building blocks with appropriate overlap sequences, making their assembly into the complete construct more efficient. BOOST Workflows enable, in a customized fashion, the execution of the Juggler, Polisher, and Partitioner tools on a batch of sequences.
EndofungalMycetohabitans(formerlyBurkholderia) spp. rely on a type III secretion system to deliver mostly unidentified effector proteins when colonizing their host fungus,Rhizopus microsporus. The one known secreted effector family fromMycetohabitansconsists of homologues of transcription activator-like (TAL) effectors, which are used by plant pathogenicXanthomonasandRalstoniaspp. to activate host genes that promote disease. These ‘BurkholderiaTAL-like (Btl)’ proteins bind corresponding specific DNA sequences in a predictable manner, but their genomic target(s) and impact on transcription in the fungus are unknown. Recent phenotyping of Btl mutants of twoMycetohabitansstrains revealed that the single Btl in oneMycetohabitans endofungorumstrain enhances fungal membrane stress tolerance, while others in aMycetohabitans rhizoxinicastrain promote bacterial colonization of the fungus. The phenotypic diversity underscores the need to assess the sequence diversity and, given that sequence diversity translates to DNA targeting specificity, the functional diversity of Btl proteins. Using a dual approach to maximize capture of Btl protein sequences for our analysis, we sequenced and assembled nineMycetohabitansspp. genomes using long-read PacBio technology and also mined available short-read Illumina fungal–bacterial metagenomes. We show thatbtlgenes are present across diverseMycetohabitansstrains from Mucoromycota fungal hosts yet vary in sequences and predicted DNA binding specificity. Phylogenetic analysis revealed distinct clades of Btl proteins and suggested thatMycetohabitansmight contain more species than previously recognized. Within our data set, Btl proteins were more conserved acrossM. rhizoxinicastrains than acrossM. endofungorum, but there was also evidence of greater overall strain diversity within the latter clade. Overall, the results suggest that Btl proteins contribute to bacterial–fungal symbioses in myriad ways.
The advance of high-throughput molecular biology tools allows in-depth profiling of microbial communities in soils, which possess a high diversity of prokaryotic microorganisms. Amplicon-based sequencing of 16S rRNA genes is the most common approach to studying the richness and composition of soil prokaryotes. To reliably detect different taxonomic lineages of microorganisms in a single soil sample, an adequate pipeline including DNA isolation, primer selection, PCR amplification, library preparation, DNA sequencing, and bioinformatic post-processing is required. Besides DNA sequencing quality and depth, the selection of PCR primers and PCR amplification reactions arguably have the largest influence on the results. This study tested the performance and potential bias of two primer pairs, i.e., 515F (Parada)-806R (Apprill) and 515F (Parada)-926R (Quince) in the standard pipelines of 16S rRNA gene Illumina amplicon sequencing protocol developed by the Earth Microbiome Project (EMP), against shotgun metagenome-based 16S rRNA gene reads. The evaluation was conducted using five differently managed soils. We observed a higher richness of soil total prokaryotes by using reverse primer 806R compared to 926R, contradicting to in silico evaluation results. Both primer pairs revealed various degrees of taxon-specific bias compared to metagenome-derived 16S rRNA gene reads. Nonetheless, we found consistent patterns of microbial community variation associated with different land uses, irrespective of primers used. Total microbial communities, as well as ammonia oxidizing archaea (AOA), the predominant ammonia oxidizers in these soils, shifted along with increased soil pH due to agricultural management. In the unmanaged low pH plot abundance of AOA was dominated by the acid-tolerant NS-Gamma clade, whereas limed agricultural plots were dominated by neutral-alkaliphilic NS-Delta/NS-Alpha clades. This study stresses how primer selection influences community composition and highlights the importance of primer selection for comparative and integrative studies, and that conclusions must be drawn with caution if data from different sequencing pipelines are to be compared.
Genomad aims to identify mobile genetic elements (namely, viruses and plasmids) from DNA sequence data. It uses a combination of marker gene identification and machine learning models to find likely virus/plasmid candidates in environmental DNA sequencing data. It has an improved classification performance over similar tools.
The accurate identification of SARS-CoV-2 (SC2) variants and estimation of their abundance in mixed population samples (e.g., air or wastewater) is imperative for successful surveillance of community level trends. Assessing the performance of SC2 variant composition estimators (VCEs) should improve our confidence in public health decision making. Here, we introduce a linear regression based VCE and compare its performance to four other VCEs: two re-purposed DNA sequence read classifiers (Kallisto and Kraken2), a maximum-likelihood based method (Lineage deComposition for Sars-Cov-2 pooled samples (LCS)), and a regression based method (Freyja). We simulated DNA sequence datasets of known variant composition from both Illumina and Oxford Nanopore Technologies (ONT) platforms and assessed the performance of each VCE. We also evaluated VCEs performance using publicly available empirical wastewater samples collected for SC2 surveillance efforts. Bioinformatic analyses were performed with a custom NextFlow workflow (C-WAP, CFSAN Wastewater Analysis Pipeline). Relative root mean squared error (RRMSE) was used as a measure of performance with respect to the known abundance and concordance correlation coefficient (CCC) was used to measure agreement between pairs of estimators. Based on our results from simulated data, Kallisto was the most accurate estimator as it had the lowest RRMSE, followed by Freyja. Kallisto and Freyja had the most similar predictions, reflected by the highest CCC metrics. We also found that accuracy was platform and amplicon panel dependent. For example, the accuracy of Freyja was significantly higher with Illumina data compared to ONT data; performance of Kallisto was best with ARTICv4. However, when analyzing empirical data there was poor agreement among methods and variations in the number of variants detected (e.g., Freyja ARTICv4 had a mean of 2.2 variants while Kallisto ARTICv4 had a mean of 10.1 variants). This work provides an understanding of the differences in performance of a number of VCEs and how accurate they are in capturing the relative abundance of SC2 variants within a mixed sample (e.g., wastewater). Such information should help officials gauge the confidence they can have in such data for informing public health decisions.
The work performed in this project has demonstrated the ability to construct proteolytic enzyme substrates that are PCR and sequencing-readable reporter molecules. Specifically, the goal was to detect those reporter molecules via PCR and Oxford Nanopore Technologies MinION sequencing methods following exposure to the biomarker protease thrombin. The assay development focused on binding the constructed peptide-oligonucleotide chimera to immobilized streptavidin. The action of thrombin on the peptide portion of the molecule released the oligonucleotide for detection. Detection of protease activity was demonstrated in a concentration-dependent manner using MALDI-MS, RT-PCR and DNA sequencing. Additional steps to remove background release of reporter molecules during the assay was used to improve the difference in detected oligonucleotide reporter following protease activity. Additional steps in assay development will be to (1) test the assay in an appropriate matrix, (2) investigate detection using additional DNA sequencing platforms and (3) demonstrate multiplexed detection of multiple protease markers in a single reaction.
Not provided.