Search NASASearch

SEARCH · Search NASA

Results for “motifs”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Discovering methylated DNA motifs in bacterial nanopore sequencing data with MIJAMP

Abstract Bacterial DNA methylation is involved in diverse cellular functions, including modulation of gene expression, DNA repair, and restriction–modification systems for defense against viruses and other foreign DNA. Restriction systems hinder efforts to engineer organisms to produce fuels and chemicals from waste and renewable feedstocks by degrading DNA during transformation. Methylome analysis allows identification of motifs within a bacterial chromosome that may be targeted by native restriction enzymes. Further expression of the corresponding methyltransferases in Escherichia coli allows plasmid DNA to be protected from restriction in the target organism, thereby drastically enhancing transformation efficiency. Nanopore sequencing can detect methylated bases, but software is needed to transform modified base coordinates into methylated motifs. Here, we develop MIJAMP (MIJAMP Is Just A MethylBED Parser), a software package that was developed to discover methylated motifs from the output of ONT’s Modkit or other data in the methylBED format. MIJAMP employs a human-driven refinement strategy that empirically validates all motifs against genome-wide methylation data, thus eliminating incorrect motifs. MIJAMP also reports methylation data on specific, user-defined motifs. Using MIJAMP, we determined the methylated motifs both in a control strain (wild-type E. coli) and in Synecococcus sp. strain PCC7002, laying the foundation for improved transformation in this organism. MIJAMP is available at https://code.ornl.gov/alexander-public/mijamp/. One Sentence Summary: Here we describe software written to discover DNA methylation motifs from nanopore sequencing data.

59 BASIC BIOLOGICAL SCIENCES

Explainable machine learning reveals that local structural motifs encode the thermodynamic state across the CuZr metallic glass-forming range

Metallic glasses derive their properties from the statistics of local atomic motifs rather than from long-range order, yet a quantitative, chemistry-specific link between motif populations and the underlying glassy state has remained elusive. In this work we combine large-scale molecular dynamics, Voronoi tessellation, deep neural networks, and SHapley Additive exPlanations (SHAP) to identify which local structural motifs define the glassy state of Cu—Zr metallic glasses. A dataset of 17,180 atomistic configurations spanning ten compositions (Cu 20 Zr 80 –Cu 80 Zr 20 ) and four quench rates (10 9 –10 12 K/s) is used to train a feed-forward neural network that regresses temperature across the 50–2000 K liquid–supercooled–glass range, achieving a mean absolute error of 19.89 K and R 2 = 0.9974, confirming that the local structural state is faithfully encoded in motif-level structure. SHAP analysis then reveals that a tightly coupled near-icosahedral family of motifs (coordination numbers (CN) 11–13, including the full icosahedron 001200 and its single-atom-perturbation sibling 10930) collectively encodes the thermodynamic state of the system across the full glass-forming range. The CN = 11–13 ordered members carry negative SHAP values at high populations, tracking the most deeply-quenched configurations, while 10930 shows the reversed signature consistent with its role as a soft-spot host whose population shrinks as the icosahedral network deepens. The analysis demonstrates that explainable machine learning can isolate the minimal motif vocabulary defining the glassy state and recovers the near-icosahedral building blocks previously identified by data-driven analyses of Cu—Zr. The approach provides a general, chemistry-specific route for characterizing the structural state of disordered materials.

36 MATERIALS SCIENCE

cWINNOWER Algorithm for Finding Fuzzy DNA Motifs

The cWINNOWER algorithm detects fuzzy motifs in DNA sequences rich in protein-binding signals. A signal is defined as any short nucleotide pattern having up to d mutations differing from a motif of length l. The algorithm finds such motifs if multiple mutated copies of the motif (i.e., the signals) are present in the DNA sequence in sufficient abundance. The cWINNOWER algorithm substantially improves the sensitivity of the winnower method of Pevzner and Sze by imposing a consensus constraint, enabling it to detect much weaker signals. We studied the minimum number of detectable motifs qc as a function of sequence length N for random sequences. We found that qc increases linearly with N for a fast version of the algorithm based on counting three-member sub-cliques. Imposing consensus constraints reduces qc, by a factor of three in this case, which makes the algorithm dramatically more sensitive. Our most sensitive algorithm, which counts four-member sub-cliques, needs a minimum of only 13 signals to detect motifs in a sequence of length N = 12000 for (l,d) = (15,4).

Liang, Shoudan

Resolving local structural motifs across the phase evolution of zinc titanates with computational x-ray absorption spectroscopy

Resolving the local structure motifs that characterize phase evolution as a function of composition is a key challenge in structure characterization of complex materials. Here, in this study, we combine first-principles simulations and x-ray absorption near-edge structures (XANES) analysis to gain insights into the structure evolution revealed by measurements across a combinatorial zinc titanate thin film, which was grown with smoothly varying composition over a wide range of the Ti:Zn ratio. Specifically, we propose a cluster blind-signal-separation (cBSS) method for XANES spectral analysis based on a library of the structures and spectra of representative local motifs. In addition to motifs from zinc titanate crystals, two types of Ti-defect models constructed in this study are key to the understanding of the structure characteristics in the Zn-rich region. The cBSS method makes use of both spectral clustering of the simulated site-XANES spectra library and the BSS procedure to construct high-fidelity spectral basis functions from an experimental spectral sequence. The method provides a rigorous measure of the spectral sensitivity and basis completeness. The results of the XANES analysis are corroborated with other experimental modalities, including x-ray diffraction and spectroscopic ellipsometry, to validate the cBSS method. The calculated motif weights resulting from fitting the XANES spectra with the cBSS basis probe the atomic structure characteristics of both crystalline and amorphous phases as a function of the Ti/Zn composition. The insights of the local structure motif evolution are pivotal to the understanding of the nonmonotonic trend in the optical gap, which may lead to potential applications through tuning the optical properties of zinc titanate. The workflow of the XANES spectral analysis developed in this work can be generalized to construct the structure-property relationship in a broad material space.

36 MATERIALS SCIENCE

cWINNOWER algorithm for finding fuzzy dna motifs

The cWINNOWER algorithm detects fuzzy motifs in DNA sequences rich in protein-binding signals. A signal is defined as any short nucleotide pattern having up to d mutations differing from a motif of length l. The algorithm finds such motifs if a clique consisting of a sufficiently large number of mutated copies of the motif (i.e., the signals) is present in the DNA sequence. The cWINNOWER algorithm substantially improves the sensitivity of the winnower method of Pevzner and Sze by imposing a consensus constraint, enabling it to detect much weaker signals. We studied the minimum detectable clique size qc as a function of sequence length N for random sequences. We found that qc increases linearly with N for a fast version of the algorithm based on counting three-member sub-cliques. Imposing consensus constraints reduces qc by a factor of three in this case, which makes the algorithm dramatically more sensitive. Our most sensitive algorithm, which counts four-member sub-cliques, needs a minimum of only 13 signals to detect motifs in a sequence of length N = 12,000 for (l, d) = (15, 4). Copyright Imperial College Press.

Evaluation Studies

Methods for Identifying Ligands that Target Nucleic Acid Molecules and Nucleic Acid Structural Motifs

Disclosed are methods for identifying a nucleic acid (e.g., RNA, DNA, etc.) motif which interacts with a ligand. The method includes providing a plurality of ligands immobilized on a support, wherein each particular ligand is immobilized at a discrete location on the support; contacting the plurality of immobilized ligands with a nucleic acid motif library under conditions effective for one or more members of the nucleic acid motif library to bind with the immobilized ligands; and identifying members of the nucleic acid motif library that are bound to a particular immobilized ligand. Also disclosed are methods for selecting, from a plurality of candidate ligands, one or more ligands that have increased likelihood of binding to a nucleic acid molecule comprising a particular nucleic acid motif, as well as methods for identifying a nucleic acid which interacts with a ligand.

Disney, Matthew D.

An evolutionarily conserved tryptophan cage promotes folding of the extended RNA recognition motif in the hnRNPR ‐like protein family

Abstract The heterogeneous nuclear ribonucleoprotein (hnRNP) R‐like family is a class of RNA binding proteins in the hnRNP superfamily with diverse functions in RNA processing. Here, we present the 1.90 Å X‐ray crystal structure and solution NMR studies of the first RNA recognition motif (RRM) of human hnRNPR. We find that this domain adopts an extended RRM (eRRM1) featuring a canonical RRM with a structured N‐terminal extension (N ext ) motif that docks against the RRM and extends the β‐sheet surface. The adjoining loop is structured and forms a tryptophan cage motif to position the N ext motif for docking to the RRM. Combining mutagenesis, solution NMR spectroscopy, and thermal denaturation studies, we evaluate the importance of residues in the N ext –RRM interface and adjoining loop on eRRM folding and conformational dynamics. We find that these sites are essential for protein solubility, conformational ordering, and thermal stability. Consistent with their importance, mutations in the N ext –RRM interface and loop are associated with several cancers in a survey of somatic mutations in cancer studies. Sequence and structure comparison of the human hnRNPR eRRM1 to experimentally verified and predicted hnRNPR‐like proteins reveals conserved features in the eRRM.

Biochemistry & Molecular Biology

The Thiamin Pyrophosphate-Motif

Using databases the authors have identified a common thiamin pyrophosphate (TPP)-motif in the family of functionally diverse TPP-dependent enzymes. This common motif consists of multimeric organization of subunits and two catalytic centers. Each catalytic center (PP:PYR) is formed at the interface of the PP-domain binding the magnesium ion, pyrophosphate and amhopyrimidine ring of TPP, and the PYR-domain binding the aminopyrimidine ring of that cofactor. A pair of these catalytic centers constitutes the catalytic core (PP:PYR)(sub 2) within these enzymes. Analysis of the structural elements of this catalytic core reveals novel definition of the common amino acid sequences, which are GXPhiX(sub 4)(G)PhiXXGQ and GDGX(sub 25-30)NN in the PP-domain, and the EX(sub 4)(G)PhiXXGPhi in the PYR-domain, where Phi corresponds to a hydrophobic amino acid. This TPP-motif provides a novel tool for annotation of TPP-dependent enzymes useful in advancing functional proteomics.

Dominiak, P.

The Thiamin Pyrophosphate-Motif

Using databases the authors have identified a common thiamin pyrophosphate (TPP)-motif in the family of functionally diverse TPP-dependent enzymes. This common motif consists of multimeric organization of subunits, two catalytic centers, common amino acid sequence, and specific contacts to provide a flip-flop, or alternate site, mechanism of action. Each catalytic center [PP:PYR] is formed at the interface of the PP-domain binding the magnesium ion, pyrophosphate and aminopyrimidine ring of TPP, and the PYR-domain binding the aminopyrimidine ring of that cofactor. A pair of these catalytic centers constitutes the catalytic core [PP:PYR]* within these enzymes. Analysis of the structural elements of this catalytic core reveals novel definition of the common amino acid sequences, which are GX@&(G)@XXGQ, and GDGX25-30 within the PP- domain, and the E&(G)@XXG@ within the PYR-domain, where Q, corresponds to a hydrophobic amino acid. This TPP-motif provides a novel tool for annotation of TPP-dependent enzymes useful in advancing functional proteomics.

Dominiak, Paulina M.

PURE mRNA display and cDNA display provide rapid detection of core epitope motif via high‐throughput sequencing

The reconstructed in vitro translation system known as the PURE system has been used in a variety of cell‐free experiments such as the expression of native and de novo proteins as well as various display methods to select for functional polypeptides. We developed a refined PURE‐based display method for the preparation of stable messenger RNA (mRNA) and complementary DNA (cDNA)‐peptide conjugates and validated its utility for in vitro selection. Our conjugate formation efficiency exceeded 40%, followed by gel purification to allow minimum carry‐over of components from the translation system to the downstream assay enabling clean and efficient random peptide sequence screening. We chose the commercially available anti‐FLAG M2 antibody as a target molecule for validation. Starting from approximately 1.7 × 10(exp 12) random sequences, a round‐by‐round high‐throughput sequencing showed clear enrichment of the FLAG epitope DYKDDD as well as revealing consensus FLAG epitope motif DYK(D/L/N)(L/Y/D/N/F)D. Enrichment of core FLAG motifs lacking one of the four key residues (DYKxxD) indicates that Tyr(Y) and Lys (K) appear as the two key residues essential for binding. Furthermore, the comparison between mRNA display and cDNA display method resulted in overall similar performance with slightly higher enrichment for mRNA display. We also show that gel purification steps in the refined PURE‐based display method improve conjugate formation efficiency and enhance the enrichment rate of FLAG epitope motifs in later rounds of selection especially for mRNA display. Overall, the generalized procedure and consistent performance of two different display methods achieved by the commercially available PURE system will be useful for future studies to explore the sequence and functional space of diverse polypeptides.

cDNA display, FLAG epitope, mRNA display, peptide

Molecular Evolution of the H5 and H7 Highly Pathogenic Avian Influenza Virus Haemagglutinin Cleavage Site Motif

ABSTRACT Avian influenza viruses are ubiquitous in the Anatinae subfamily of aquatic birds and occasionally spill over to poultry. Infection with low pathogenicity avian influenza viruses generally leads to subclinical or mild clinical disease. In contrast, highly pathogenic avian influenza viruses emerge from low pathogenic forms and can cause severe disease associated with extraordinarily high mortality rates. Here, we describe the natural history of avian influenza virus, with a focus on H5Nx and H7Nx subtypes, and the emergence of highly pathogenic forms; we review the biology of AIV; we examine cleavage of haemagglutinin by host cell enzymes with a particular emphasis on the biochemical properties of the proprotein convertases, and trypsin and trypsin‐like proteases; we describe mechanisms implicated in the functional evolution of the haemagglutinin cleavage site motif that leads to emergence of HPAIVs; and finally, we discuss the diversity of H5 and H7 haemagglutinin cleavage site sequence motifs. It is crucial to understand the molecular attributes that drive the emergence and evolution of HPAIVs with pandemic potential to inform risk assessments and mitigate the threat of HPAIVs to poultry and human populations.

Luczo, Jasmina M. [Australian Animal Health Labora

First-Principles Evaluation of Proton Hopping in Tetrahedral Oxide Motifs

Proton-conducting oxides (PCOs) are important materials used as ionic conductors for energy conversion technologies. Existing research efforts on PCO optimization and discovery generally focus on complex perovskite-based oxides that require doping and alloying to engineer oxygen deficiency and high proton conductivity. However, the variety of chemical compositions and coordination environments in oxides poses challenges for efficient materials design. In this computational study, we construct a database of simplified motifs to elucidate the relationship between fundamental materials chemistry and proton kinetics. Specifically, we focus on the zincblende crystal structure as a proxy for tetrahedral metal–oxide (M–O) coordination environments. We systematically quantified the effects of cation type, oxidation states, and M–O bond lengths on the proton hopping barrier, and found that strong M–O bonds and metal cations with large and variable oxidation states (e.g., Mo 6+ , V 5+ ) lead to smaller proton hopping barriers. By mapping the candidate cations and their preferred bond geometries onto materials databases such as the Inorganic Crystal Structure Database (ICSD) and Materials Project, we identified real materials containing the corresponding metal–oxide units. In general, we observed good agreement between the calculated proton hopping barriers obtained in real crystal structures and those predicted by our motif database. We also discuss the limitations of our model and possible future extensions to improve its predictive capabilities. Overall, our model provides a first step for the rational design and quick screening of energy-efficient PCOs.

organic

Engineering microalgal cell wall-anchored proteins using GP1 PPSPX motifs and releasing with intein-mediated fusion

AbstractHarnessing and controlling the localization of recombinant proteins is critical for advancing applications in synthetic biology, industrial biotechnology, and drug delivery. This study explores protein anchoring and controlled release inChlamydomonas reinhardtii, providing innovative tools for these fields. Using truncated variants of the GP1 glycoprotein fused to the plastic-degrading enzyme PHL7, we identified the PPSPX motif as essential for anchoring proteins to the cell wall. Constructs with increased PPSPX content exhibited reduced secretion but improved anchoring, pinpointing the potential anchor-signal sites of GP1 and highlighting the distinct roles of these motifs in protein localization. Building on the anchoring capabilities established with these glycomodules, we also demonstrated a controlled release system using a pH-sensitive intein derived from RecA fromMycobacterium tuberculosis. This intein efficiently cleaved and released PHL7 and mCherry that was fused to GP1 under acidic conditions, enabling precise temporal and environmental control. At pH 5.5, fluorescence kinetics demonstrated significant mCherry release from the pJPW4mCherry construct within 4 hours. In contrast, release was minimal under pH 8.0 conditions and negligible for the pJPW2mCherry (W2) control, irrespective of the pH. Additionally, bands on the Western blot at the expected size of mCherry also showed its efficient release from the mCherry::intein::GP1 fusion protein at pH 5.5. Conversely, at pH 8.0, no bands were detected. This anchor-release approach offers significant potential for drug delivery, biocatalysis, and environmental monitoring applications. By integrating glycomodules and pH-sensitive inteins, this study establishes a versatile framework for optimizing protein localization and release inC. reinhardtii, with broad implications for proteomics, biofilm engineering, and scalable therapeutic delivery systems.Graphical Abstract

Kang, Kalisa (ORCID:0009000619398129)

The Thiamine-Pyrophosphate-Motif

Thiamin pyrophosphate (TPP), a derivative of vitamin B1, is a cofactor for enzymes performing catalysis in pathways of energy production including the well known decarboxylation of a-keto acid dehydrogenases followed by transketolation. TPP-dependent enzymes constitute a structurally and functionally diverse group exhibiting multimeric subunit organization, multiple domains and two chemically equivalent catalytic centers. Annotation of functional TPP-dependcnt enzymes, therefore, has not been trivial due to low sequence similarity related to this complex organization. Our approach to analysis of structures of known TPP-dependent enzymes reveals for the first time features common to this group, which we have termed the TPP-motif. The TPP-motif consists of specific spatial arrangements of structural elements and their specific contacts to provide for a flip-flop, or alternate site, enzymatic mechanism of action. Analysis of structural elements entrained in the flip-flop action displayed by TPP-dependent enzymes reveals a novel definition of the common amino acid sequences. These sequences allow for annotation of TPP-dependent enzymes, thus advancing functional proteomics. Further details of three-dimensional structures of TPP-dependent enzymes will be discussed.

Ciszak, Ewa

Identification of a TAAT-containing motif required for high level expression of the COL1A1 promoter in differentiated osteoblasts of transgenic mice

Our previous studies have shown that the 49-base pair region of promoter DNA between -1719 and -1670 base pairs is necessary for transcription of the rat COL1A1 gene in transgenic mouse calvariae. In this study, we further define this element to the 13-base pair region between -1683 and -1670. This element contains a TAAT motif that binds homeodomain-containing proteins. Site-directed mutagenesis of this element in the context of a COL1A1-chloramphenicol acetyltransferase construct extending to -3518 base pairs decreased the ratio of reporter gene activity in calvariae to tendon from 3:1 to 1:1, suggesting a preferential effect on activity in calvariae. Moreover, chloramphenicol acetyltransferase-specific immunofluorescence microscopy of transgenic calvariae showed that the mutation preferentially reduced levels of chloramphenicol acetyltransferase protein in differentiated osteoblasts. Gel mobility shift assays demonstrate that differentiated osteoblasts contain a nuclear factor that binds to this site. This binding activity is not present in undifferentiated osteoblasts. We show that Msx2, a homeodomain protein, binds to this motif; however, Northern blot analysis revealed that Msx2 mRNA is present in undifferentiated bone cells but not in fully differentiated osteoblasts. In addition, cotransfection studies in ROS 17/2.8 osteosarcoma cells using an Msx2 expression vector showed that Msx2 inhibits a COL1A1 promoter-chloramphenicol acetyltransferase construct. Our results suggest that high COL1A1 expression in bone is mediated by a protein that is induced during osteoblast differentiation. This protein may contain a homeodomain; however, it is distinct from homeodomain proteins reported previously to be present in bone.

NASA Discipline Cell Biology

A calmodulin binding protein from Arabidopsis is induced by ethylene and contains a DNA-binding motif

Calmodulin (CaM), a key calcium sensor in all eukaryotes, regulates diverse cellular processes by interacting with other proteins. To isolate CaM binding proteins involved in ethylene signal transduction, we screened an expression library prepared from ethylene-treated Arabidopsis seedlings with 35S-labeled CaM. A cDNA clone, EICBP (Ethylene-Induced CaM Binding Protein), encoding a protein that interacts with activated CaM was isolated in this screening. The CaM binding domain in EICBP was mapped to the C-terminus of the protein. These results indicate that calcium, through CaM, could regulate the activity of EICBP. The EICBP is expressed in different tissues and its expression in seedlings is induced by ethylene. The EICBP contains, in addition to a CaM binding domain, several features that are typical of transcription factors. These include a DNA-binding domain at the N terminus, an acidic region at the C terminus, and nuclear localization signals. In database searches a partial cDNA (CG-1) encoding a DNA-binding motif from parsley and an ethylene up-regulated partial cDNA from tomato (ER66) showed significant similarity to EICBP. In addition, five hypothetical proteins in the Arabidopsis genome also showed a very high sequence similarity with EICBP, indicating that there are several EICBP-related proteins in Arabidopsis. The structural features of EICBP are conserved in all EICBP-related proteins in Arabidopsis, suggesting that they may constitute a new family of DNA binding proteins and are likely to be involved in modulating gene expression in the presence of ethylene.

Non-NASA Center

Discovery of Ternary Antimonides A–Al–Sb (A = Rb or Cs) with Desired Structural Motifs Guided by Machine Learning

Specific structural motifs in inorganic solids are often related to their targeted physical properties. For many classes of solids, such as Zintl phases and polar intermetallics, the crystal structures are diverse and not easy to predict. Various antimonides that are potential thermoelectric materials were proposed to be synthesizable on the basis of their estimated formation energies. Their structures were broadly classified as clathrate, channel, layered, or network through a machine learning model trained on existing ternary phases and features based on elemental properties using the sure independence screening and sparsifying operator algorithm. Through experimental validation, three new ternary antimonides were synthesized and confirmed to form layered structures: tetragonal RbAlSb 2 and CsAlSb 2 , which are isopointal but not isotypic to LiBSi 2 ; and monoclinic Rb 2 Al 2 Sb 3 , which adopts the Na 2 Al 2 Sb 3 -type structure. Finally, reinvestigation of the related compound Cs 2 In 2 Sb 3 revealed a low thermal conductivity and p-type semiconducting behavior.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH