Search NASA⌕ Search

SEARCH · Search NASA

Results for “Protein domains”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Ion-selective conformational stabilization of a disordered repeats-in-toxin protein domain

Ion-binding intrinsically disordered proteins (IDPs) recruit and bind to specific metal ions to perform critical biological functions. In proteins where ion binding and structural transitions are coupled, interactions with off-target toxic metals can dramatically disrupt protein structure and function, exemplified by lead and mercury poisoning. Understanding the complex mechanisms underlying how IDPs exclude or allow binding to different ionic species is crucial for addressing the origins of metal toxicity in biological systems. Here, we elucidate mechanisms of ion selectivity in an IDP that adopts a structure upon Ca 2+ binding. We probed ion-induced conformational changes of a repeats-in-toxin (RTX) protein domain in the presence of different ion ligands—Mg 2+ , Ca 2+ , Sr 2+ , and Ba 2+ —with chemical similarities but drastically different ionic radii. RTX adopts ion-selective conformations measured by x-ray crystallography, small-angle x-ray scattering, and circular dichroism. High-resolution x-ray structures reveal that Sr 2+ induces a nearly identical RTX structure as natively binding Ca 2+ , enabled by the intrinsic flexibility and disorder of the protein. Small-angle x-ray scattering and circular dichroism indicate that smaller Mg 2+ does not induce a significant conformational change in RTX, whereas larger Ba 2+ induces a partially folded structure. These results highlight the importance of geometric constraints imposed by protein structure in determining metal ion selectivity, yielding insights into how off-target ion binding may result in protein misfolding and malfunction.

Gudinas, Alana P. [Stanford Univ., CA (United Stat↗

Powdery mildew effectors AVR A1 and BEC1016 target the ER J‐domain protein Hv ERdj3B required for immunity in barley

Abstract The barley powdery mildew fungus, Blumeria hordei (Bh), secretes hundreds of candidate secreted effector proteins (CSEPs) to facilitate pathogen infection and colonization. One of these, CSEP0008, is directly recognized by the barley nucleotide‐binding leucine‐rich‐repeat (NLR) receptor MLA1 and therefore is designated AVR A1 . Here, we show that AVR A1 and the sequence‐unrelated Bh effector BEC1016 (CSEP0491) suppress immunity in barley. We used yeast two‐hybrid next‐generation interaction screens (Y2H‐NGIS), followed by binary Y2H and in planta protein–protein interactions studies, and identified a common barley target of AVR A1 and BEC1016, the endoplasmic reticulum (ER)‐localized J‐domain protein Hv ERdj3B. Silencing of this ER quality control (ERQC) protein increased Bh penetration. Hv ERdj3B is ER luminal, and we showed using split GFP that AVR A1 and BEC1016 translocate into the ER signal peptide‐independently. Overexpression of the two effectors impeded trafficking of a vacuolar marker through the ER; silencing of Hv ERdj3B also exhibited this same cellular phenotype, coinciding with the effectors targeting this ERQC component. Together, these results suggest that the barley innate immunity, preventing Bh entry into epidermal cells, requires ERQC. Here, the J‐domain protein Hv ERdj3B appears to be essential and can be regulated by AVR A1 and BEC1016. Plant disease resistance often occurs upon direct or indirect recognition of pathogen effectors by host NLR receptors. Previous work has shown that AVR A1 is directly recognized in the cytosol by the immune receptor MLA1. We speculate that the AVR A1 J‐domain target being inside the ER, where it is inapproachable by NLRs, has forced the plant to evolve this challenging direct recognition.

54 ENVIRONMENTAL SCIENCES↗

Normalized Cut Algorithm for Automated Assignment of Protein Domains

We present a novel computational method for automatic assignment of protein domains from structural data. At the core of our algorithm lies a recently proposed clustering technique that has been very successful for image-partitioning applications. This grap.,l-theory based clustering method uses the notion of a normalized cut to partition. an undirected graph into its strongly-connected components. Computer implementation of our method tested on the standard comparison set of proteins from the literature shows a high success rate (84%), better than most existing alternative In addition, several other features of our algorithm, such as reliance on few adjustable parameters, linear run-time with respect to the size of the protein and reduced complexity compared to other graph-theory based algorithms, would make it an attractive tool for structural biologists.

Samanta, M. P.↗

Structure-aware annotation of leucine-rich repeat domains

Protein domain annotation is typically done by predictive models such as HMMs trained on sequence motifs. However, sequence-based annotation methods are prone to error, particularly in calling domain boundaries and motifs within them. These methods are limited by a lack of structural information accessible to the model. With the advent of deep learning-based protein structure prediction, existing sequenced-based domain annotation methods can be improved by taking into account the geometry of protein structures. We develop dimensionality reduction methods to annotate repeat units of the Leucine Rich Repeat solenoid domain. The methods are able to correct mistakes made by existing machine learning-based annotation tools and enable the automated detection of hairpin loops and structural anomalies in the solenoid. The methods are applied to 127 predicted structures of LRR-containing intracellular innate immune proteins in the model plant Arabidopsis thaliana and validated against a benchmark dataset of 172 manually-annotated LRR domains.

Xu, Boyan↗

Modelling protein functional domains in signal transduction using Maude

Modelling of protein-protein interactions in signal transduction is receiving increased attention in computational biology. This paper describes recent research in the application of Maude, a symbolic language founded on rewriting logic, to the modelling of functional domains within signalling proteins. Protein functional domains (PFDs) are a critical focus of modern signal transduction research. In general, Maude models can simulate biological signalling networks and produce specific testable hypotheses at various levels of abstraction. Developing symbolic models of signalling proteins containing functional domains is important because of the potential to generate analyses of complex signalling networks based on structure-function relationships.

Signal Transduction↗

Understanding amyloids to prevent biofilm formation in space

There is a pressing need to search for novel approaches to combat biofilm formation, both in space and in medical applications. Many proteins have the ability to form ordered aggregates called amyloids. Amyloids are known to be an important part of biofilms. The use of anti-amyloid drugs is a novel venue for the development of antimicrobial agents. The ultrastructure of the amyloid aggregate shows a high packing of proteins, the second-order structure of which is dominated by β-sheets. The ability to form an amyloid aggregate is especially typical for proteins containing domains (protein fragments) with sufficient lability to arrange themselves in a tight β-sheet structure. Bioinformatics tools allow the prediction of such behavior of proteins in genomic data. We use GeneLab data of microbial populations identified aboard the International Space Station and other spacecraft to look for bacterial species that utilize amyloid aggregation in biofilm formation. We use a combined bioinformatic approach with a relatively high throughput molecular biology assay and biophysical assays to evaluate the anti-amyloid anti-biofilm approach. The significance of the research extends from understanding basic microbial community responses to spaceflight, to biofouling of the built environments in space as well as the long-term health of astronauts. Bioinformatics shows that onboard the ISS, bacterial species produce far more amyloid and prion proteins than are currently verified, hence their role in bacterial ecosystems is largely unknown. As we propose there is a link between amyloid formation in space and biofilm production, this research should lead to new paths for biofilm remediation in space.

Tomasz Zajkowski↗

Archaeal protein containing domain of unknown function 2193 undergoes oligomeric reconfiguration upon iron–sulfur cluster binding

Methanogenic archaea are particularly rich in iron–sulfur proteins, yet their roles remain largely enigmatic. Here, we characterized aMethanococcus voltae(Mvo) protein from the domain of unknown function (DUF) 2193 family, a group of proteins present primarily in archaea and characterized by a conserved cysteine‐rich C‐terminal motif.MvoDUF2193 was heterologously expressed and characterized by a range of spectroscopic and analytical methods. The results demonstrate thatMvoDUF2193 binds a single [4Fe–4S] cluster per subunit and that cluster occupancy regulates the transition from an apo tetramer to a [4Fe–4S] monomeric form. We hypothesize thatMvoDUF2193 serves a regulatory role in the cell, mediated by [Fe–S] cluster binding and changes in oligomeric state.

Biochemistry & Molecular Biology↗

Evolution of EF-hand calcium-modulated proteins. II. Domains of several subfamilies have diverse evolutionary histories

In the first report in this series we described the relationships and evolution of 152 individual proteins of the EF-hand subfamilies. Here we add 66 additional proteins and define eight (CDC, TPNV, CLNB, LPS, DGK, 1F8, VIS, TCBP) new subfamilies and seven (CAL, SQUD, CDPK, EFH5, TPP, LAV, CRGP) new unique proteins, which we assume represent new subfamilies. The main focus of this study is the classification of individual EF-hand domains. Five subfamilies--calmodulin, troponin C, essential light chain, regulatory light chain, CDC31/caltractin--and three uniques--call, squidulin, and calcium-dependent protein kinase--are congruent in that all evolved from a common four-domain precursor. In contrast calpain and sarcoplasmic calcium-binding protein (SARC) each evolved from its own one-domain precursor. The remaining 19 subfamilies and uniques appear to have evolved by translocation and splicing of genes encoding the EF-hand domains that were precursors to the congruent eight and to calpain and to SARC. The rates of evolution of the EF-hand domains are slower following formation of the subfamilies and establishment of their functions. Subfamilies are not readily classified by patterns of calcium coordination, interdomain linker stability, and glycine and proline distribution. There are many homoplasies indicating that similar variants of the EF-hand evolved by independent pathways.

Non-NASA Center↗

Calcium-binding proteins and development

The known roles for calcium-binding proteins in developmental signaling pathways are reviewed. Current information on the calcium-binding characteristics of three classes of cell-surface developmental signaling proteins (EGF-domain proteins, cadherins and integrins) is presented together with an overview of the intracellular pathways downstream of these surface receptors. The developmental roles delineated to date for the universal intracellular calcium sensor, calmodulin, and its targets, and for calcium-binding regulators of the cytoskeleton are also reviewed.

Review↗

CAHS: Context-Aware Homology Search

Protein homology search is foundational to bioinformatics: it supports annotation transfer, structure/function inference, and evolutionary analysis over rapidly expanding sequence repositories (e.g., UniProtKB). Profile hidden Markov models (pHMMs), as implemented in HMMER, remain the most widely trusted approach because they provide statistically calibrated E-values; however, their gap behavior is fixed once a profile is trained, despite biological evidence that insertion/deletion tolerance varies across flexible loops and intrinsically disordered regions. We present CAHS (Context-Aware Homology Search), a lightweight query-time adapter for pHMM search that incorporates learned and biologically motivated signals without changing HMMER's downstream search pipeline or its calibrated E-value reporting. Given a query sequence, CAHS computes per-residue representations from a protein language model and a disorder predictor, maps these to profile coordinates, and modulates only match-state transition rows (gap-open and gap-extension probabilities) while preserving Plan7 constraints. We comprehensively evaluate CAHS across six structurally diverse protein families and multi-domain architectures against a 570k-sequence target corpus. CAHS expands detection capability, retrieving thousands of additional remote homologs at relaxed thresholds by maintaining alignment quality through flexible regions. For multi-domain proteins, context-aware modulation resolves 94% of fragmented alignments. Crucially, CAHS preserves hit-set invariance at stringent operating points (E<10-10), demonstrating increased statistical confidence without inflating false positives. Furthermore, sharper statistical distinction between homologs and background noise during early filter stages yields up to a 3.87× acceleration in end-to-end wall-clock time on high-performance computing clusters. Overall, CAHS illustrates a practical AI-for-science design pattern: augmenting a trusted probabilistic model with query-specific learned signals to improve interpretable, reproducible inference in data-rich biology.

Bhattaram, Swethasree [Georgia Institute of Techno↗

The Origin and Early Evolution of Membrane Proteins

The origin and early evolution of membrane proteins, and in particular ion channels, are considered from the point of view that the transmembrane segments of membrane proteins are structurally quite simple and do not require specific sequences to fold. We argue that the transport of solute species, especially ions, required an early evolution of efficient transport mechanisms, and that the emergence of simple ion channels was protobiologically plausible. We also argue that, despite their simple structure, such channels could possess properties that, at the first sight, appear to require markedly larger complexity. These properties can be subtly modulated by local modifications to the sequence rather than global changes in molecular architecture. In order to address the evolution and development of ion channels, we focus on identifying those protein domains that are commonly associated with ion channel proteins and are conserved throughout the three main domains of life (Eukarya, Prokarya, and Archaea). We discuss the potassium-sodium-calcium superfamily of voltage-gated ion channels, mechanosensitive channels, porins, and ABC-transporters and argue that these families of membrane channels have sufficiently universal architectures that they can readily adapt to the diverse functional demands arising during evolution.

Pohorille, Andrew↗

Variation of Structure and Cellular Functions of Type IA Topoisomerases across the Tree of Life

Topoisomerases regulate the topological state of cellular genomes to prevent impediments to vital cellular processes, including replication and transcription from suboptimal supercoiling of double-stranded DNA, and to untangle topological barriers generated as replication or recombination intermediates. The subfamily of type IA topoisomerases are the only topoisomerases that can alter the interlinking of both DNA and RNA. In this article, we provide a review of the mechanisms by which four highly conserved N-terminal protein domains fold into a toroidal structure, enabling cleavage and religation of a single strand of DNA or RNA. We also explore how these conserved domains can be combined with numerous non-conserved protein sequences located in the C-terminal domains to form a diverse range of type IA topoisomerases in Archaea, Bacteria, and Eukarya. There is at least one type IA topoisomerase present in nearly every free-living organism. The variation in C-terminal domain sequences and interacting partners such as helicases enable type IA topoisomerases to conduct important cellular functions that require the passage of nucleic acids through the break of a single-strand DNA or RNA that is held by the conserved N-terminal toroidal domains. In addition, this review will exam a range of human genetic disorders that have been linked to the malfunction of type IA topoisomerase.

59 BASIC BIOLOGICAL SCIENCES↗

Molecular information theory meets protein folding

We propose an application of molecular information theory to analyze the folding of single domain proteins. We analyze results from various areas of protein science, such as sequence-based potentials, reduced amino acid alphabets, backbone configurational entropy, secondary structure content, residue burial layers, and mutational studies of protein stability changes. We found that the average information contained in the sequences of evolved proteins is very close to the average information needed to specify a fold ~2.2 ± 0.3 bits/(site operation). The effective alphabet size in evolved proteins equals the effective number of conformations of a residue in the compact unfolded state at around 5. We calculated an energy-to-information conversion efficiency upon folding of around 50%, lower than the theoretical limit of 70%, but much higher than human built macroscopic machines. We propose a simple mapping between molecular information theory and energy landscape theory and explore the connections between sequence evolution, configurational entropy and the energetics of protein folding.

Ignacio E. Sánchez↗

Cleanroom Microbes Survive Drying, Vacuum, and Proton Irradiation

Introduction : The goal of planetary protection at NASA is to mitigate the risk of contaminating sensitive target bodies with biological life. While many cleaning procedures have been put in place to reduce bioburden on spacecraft, microbes are experts at evolving to survive harsh conditions. Specifically, the dry, low-nutrient environment of a cleanroom (commonly used for assembly of spacecraft) can represent an environment where extremophiles can survive. Methods : Scientists at NASA MSFC wished to gather a snapshot of the microbial population within a variety of cleanrooms on site. A study was undertaken to collect air, surface, and floor samples from clean-rooms and isolate unique morphologies. From this study, 95 isolates were collected and saved in a microbial library. About 86% of these were identified at least to a genus level. Following identification, 24 microbes were selected, based on a literature review, as potential extremophiles. These were grown in liquid cultures, diluted to a set optical density, washed with water, and then applied to a sterilized Kapton coupon. Droplets were allowed to dry overnight in a biosafety cabinet. Coupons were then installed in a pelletron and pumped down to high vacuum (~1E-6 Torr). Samples were then subjected 100 keV protons at a fluence of 2x10 15 p+/cm 2 up to 4x10 15 p+/cm 2 . Following exposure, samples were returned to the microbiology lab where they were pro-cessed by submerging in water, vortexing, and then plating either droplets or spread plates. Recovery data collected was qualitative with a ranking or +, minor, or – for growth. Some selected radiotolerant strains were sequenced using the Illumina sequencing platform. The resulting genomes were annotated with the Rapid Annotations using Subsystems Technology (RAST) server and analyzed for conserved and unique stress response relevant genomic signatures to identify clues related to specific tolerances. Results and Discussion : After five rounds of proton radiation, we narrowed our isolates to five, non-spore forming bacteria that demonstrated survival: Arthrobacter koreensis, Paenarthrobacter nitroguajacolicus, Mycetocola manganoxydans , and an Erwinia sp. Furthermore, we exposed these four microbes to 254 nm wavelength light at an intensity of 80 W/m 2 at a distance of ~18 cm for 10 minutes. Only A. koreensis demonstrated survival following UV exposure. Finally, we performed whole genome sequencing on the four strains to look for genetic markers of stress resistance. When we compared the genomes of the four strains, we found that genes coding for GGDEF and EAL domains with PAS/PAC sensors were only found in A. koreensis . These domains, modulated by PAS/PAC sensors, are hypothesized to facilitate survival under drying, desiccation, and proton irradiation. Drying and Desiccation : PAS domains sense hydration changes and modulate GGDEF and EAL domain activity to adjust c-di-GMP levels, enhancing resistance to desiccation. For instance, in Pseudomonas aeruginosa , the PAS domain of RbdA modulates activity under varying hydration conditions, affecting stress responses [1]. Proton Irradiation : Proton irradiation causes oxidative stress, leading to ROS generation. PAS domains detect this stress and modulate GGDEF and EAL domains to manage oxidative stress responses. In Shewanella , EAL domain proteins modulated by PAS sensors help bacteria adapt to extreme conditions [2]. These genes upregulate other stress response genes, protecting membrane function, protein stability, DNA repair, and antioxidant defenses. The modulation of c-di-GMP by PAS domains is crucial for bacterial adaptation to stress conditions, enabling dynamic physio-logical adjustments [3]. Understanding these mechanisms provides insights into bacterial stress responses and strategies for controlling bacterial growth [4]. Conclusions : These findings indicate that clean-rooms harbor extremophile microbes that may be able to survive conditions in deep space. Furthermore, while we identified certain stress-response genes that may be at least partly responsible for the phenotypes observed in this study, there are likely unidentified genes or characteristics about A. koreensis , and other bacteria, that may allow them to survive in harsh environments. Future studies will focus on identifying these unknown genes and characteristics, further elucidating the mechanisms of extremophile survival and potentially informing the development of new biotechnologies for space exploration and other extreme environments.

Chelsi Cassilly↗

Bipartite chromatin recognition by Hop1 from two diverged Holozoa

In meiosis, ploidy reduction is driven by a complex series of DNA breakage and recombination events between homologous chromosomes, orchestrated by meiotic HORMA domain proteins (HORMADs). Meiotic HORMADs possess a central chromatin binding region (CBR) whose architecture varies across eukaryotic groups. Here, we determine high-resolution crystal structures of the meiotic HORMAD CBR from two diverged aquatic Holozoa,Schistosoma mansoniandPatiria miniata, which reveal tightly associated plant homeodomain (PHD) and winged helix-turn-helix (wHTH) domains. We show that PHD–wHTH CBRs bind duplex DNA through their wHTH domains, and identify key residues that disrupt this interaction. Combining experimental and predicted structures, we show that the CBRs’ PHDs likely interact with the tail of histone H3, and may discriminate between unmethylated and trimethylated H3 lysine 4. Finally, we show that Holozoa Hop1 CBRs bind nucleosomes in vitro in a bipartite manner involving both the PHD and wHTH domain. Our data reveal how meiotic HORMADs with PHD–wHTH CBRs can bind chromatin and potentially discriminate between chromatin states to drive meiotic recombination to specific chromosomal regions.

Life Sciences & Biomedicine - Other Topics↗

MarK, a Novosphingobium aromaticivorans kinase required for catabolism of multiple aromatic monomers

The aromatic compounds used in a variety of industrial products are currently obtained from nonrenewable petroleum sources. Alternatively, the plant polymer lignin is an abundant renewable source of aromatics, and its depolymerization generates a variety of products that can include acetovanillone, a vanillin derivative containing an acetyl side chain. The Alphaproteobacterium Novosphingobium aromaticivorans DSM12444 can metabolize several chemically modified aromatics in deconstructed lignin, but not acetovanillone. In this work, adaptive laboratory evolution identified a single amino acid change in the previously uncharacterized gene product Saro_1862 that is necessary and sufficient for N. aromaticivorans growth with acetovanillone as a sole growth substrate, as well as other aromatic monomers not metabolized by wild-type cells. We show that a glutamate (E) to lysine (K) substitution at amino acid residue 16 of Saro_1862 results in a ~1600-fold increase in the rate of ATP-dependent acetovanillone phosphorylation. We also find that recombinant Saro_1862 E16K phosphorylates several other aromatic compounds in vitro , defining the first reported catalytic activity for the widespread UPF0261 protein domain contained in Saro_1862. Thus, we propose naming Saro_1862 MarK, for multiple aromatic kinase. A 1.57 Å crystal structure of MarK E16K predicts that the E16K substitution lies in a potential ATP binding site, suggesting how this amino acid change increased catalytic activity. A search for homologs of MarK and other proteins required for acetovanillone degradation predicts that this pathway for aromatic metabolism exists throughout the bacterial phylogeny.

Novosphingobium↗

Generating Protein Structures for Pathway Discovery Using Deep Learning

Resolving the intricate details of biological phenomena at the molecular level is fundamentally limited by both length- and time scales that can be probed experimentally. Molecular dynamics (MD) simulations at various scales are powerful tools frequently employed to offer valuable biological insights beyond experimental resolution. However, while it is relatively simple to observe long-lived, stable configurations of, for example, proteins, at the required spatial resolution, simulating the more interesting rare transitions between such states often takes orders of magnitude longer than what is feasible even on the largest supercomputers available today. One common aspect of this challenge is pathway discovery, where the start and end states of a scientific phenomenon are known or can be approximated, but the mechanistic details in between are unknown. Here, we propose a representation-learning-based solution that uses interpolation and extrapolation in an abstract representation space to synthesize potential transition states, which are automatically validated using MD simulations. The new simulations of the synthesized transition states are subsequently incorporated into the representation learning, leading to an iterative framework for targeted path sampling. Our approach is demonstrated by recovering the transition of a RAS-RAF protein domain (CRD) from membrane-free to interacting with the membrane using coarse-grain MD simulations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗