Search NASA⌕ Search

SEARCH · Search NASA

Results for “Protein structure predictions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

The spatial and temporal expression of Ch-en, the engrailed gene in the polychaete Chaetopterus, does not support a role in body axis segmentation

We are interested in understanding whether the annelids and arthropods shared a common segmented ancestor and have approached this question by characterizing the expression pattern of the segment polarity gene engrailed (en) in a basal annelid, the polychaete Chaetopterus. We have isolated an en gene, Ch-en, from a Chaetopterus cDNA library. Genomic Southern blotting suggests that this is the only en class gene in this animal. The predicted protein sequence of the 1.2-kb cDNA clone contains all five domains characteristic of en proteins in other taxa, including the en class homeobox. Whole-mount in situ hybridization reveals that Ch-en is expressed throughout larval life in a complex spatial and temporal pattern. The Ch-en transcript is initially detected in a small number of neurons associated with the apical organ and in the posterior portion of the prototrochophore. At later stages, Ch-en is expressed in distinct patterns in the three segmented body regions (A, B, and C) of Chaetopterus. In all segments, Ch-en is expressed in a small set of segmentally iterated cells in the CNS. In the A region, Ch-en is also expressed in a small group of mesodermal cells at the base of the chaetal sacs. In the B region, Ch-en is initially expressed broadly in the mesoderm that then resolves into one band/segment coincident with morphological segmentation. The mesodermal expression in the B region is located in the anterior region of each segment, as defined by the position of ganglia in the ventral nerve cord, and is involved in the morphogenesis of segment-specific feeding structures late in larval life. We observe banded mesodermal and ectodermal staining in an anterior-posterior sequence in the C region. We do not observe a segment polarity pattern of expression of Ch-en in the ectoderm, as is observed in arthropods. Copyright 2001 Academic Press.

Non-NASA Center↗

Leveraging High-resolution Molecular Composition of Soil Organic Matter to Enhance Carbon Cycling Modeling

Soils store more carbon than the atmosphere and vegetation combined, yet Earth system models still struggle to predict how this vast reservoir will respond to environmental change. A central limitation is that most soil biogeochemical models represent organic matter using bulk conceptual pools or chemically homogeneous fractions, preventing direct use of rapidly expanding molecular-scale datasets. Here we develop and test a new soil decomposition framework that explicitly integrates high-resolution information on organic matter composition. First, we construct a molecularly informed litter decomposition module in which plant inputs are partitioned into five functional compound classes—carbohydrates, proteins, lignin-like aromatics, lipids, and carbonyls—using a molecular mixing model calibrated to solid-state 13 C Nuclear Magnetic Resonance (NMR) spectra. Class-specific kinetics, lignin-dependent physical protection, and substrate-driven microbial carbon use efficiency allow the module to capture metabolic tradeoffs associated with enzyme production and nutrient limitation. We then embed this litter module within a microbially explicit whole-soil model that tracks the transformation of these compound classes through particulate organic matter, dissolved organic matter, mineral-associated organic matter, and microbial biomass. High-resolution Fourier Transform Ion Cyclotron Resonance mass spectrometry (FTICR-MS) data are used to link internal pools to measurable soil organic matter fractions and to constrain key process parameters. Applications at soil-core and ecosystem scales demonstrate that the new model reproduces observed soil respiration dynamics while providing mechanistic attribution of CO 2 fluxes to specific chemical classes and pools. Compared to existing frameworks such as the Community Land Model soil biogeochemistry module and the Millennial model, our approach maintains competitive predictive skill while substantially improving interpretability and opportunities for data–model integration. This work illustrates a viable pathway for leveraging molecular-scale observations to reduce structural uncertainty in soil carbon–climate feedback projections.

54 ENVIRONMENTAL SCIENCES↗

Evolution and folding of repeat proteins

Repeat proteins are made with tandem copies of similar amino acid stretches that fold into elongated architectures. These proteins constitute excellent model systems to investigate how evolution relates to structure, folding, and function. Here, we propose a scheme to map evolutionary information at the sequence level to a coarse-grained model for repeat-protein folding and use it to investigate the folding of thousands of repeat proteins. We model the energetics by a combination of an inverse Potts-model scheme with an explicit mechanistic model of duplications and deletions of repeats to calculate the evolutionary parameters of the system at the single-residue level. These parameters are used to inform an Ising-like model that allows for the generation of folding curves, apparent domain emergence, and occupation of intermediate states that are highly compatible with experimental data in specific case studies. We analyzed the folding of thousands of natural Ankyrin repeat proteins and found that a multiplicity of folding mechanisms are possible. Fully cooperative all-or-none transitions are obtained for arrays with enough sequence-similar elements and strong interactions between them, while noncooperative element-by-element intermittent folding arose if the elements are dissimilar and the interactions between them are energetically weak. Additionally, we characterized nucleation-propagation and multidomain folding mechanisms. We show that the global stability and cooperativity of the repeating arrays can be predicted from simple sequence scores.

Ezequiel A. Galpern↗

Mechanical behavior in living cells consistent with the tensegrity model

Alternative models of cell mechanics depict the living cell as a simple mechanical continuum, porous filament gel, tensed cortical membrane, or tensegrity network that maintains a stabilizing prestress through incorporation of discrete structural elements that bear compression. Real-time microscopic analysis of cells containing GFP-labeled microtubules and associated mitochondria revealed that living cells behave like discrete structures composed of an interconnected network of actin microfilaments and microtubules when mechanical stresses are applied to cell surface integrin receptors. Quantitation of cell tractional forces and cellular prestress by using traction force microscopy confirmed that microtubules bear compression and are responsible for a significant portion of the cytoskeletal prestress that determines cell shape stability under conditions in which myosin light chain phosphorylation and intracellular calcium remained unchanged. Quantitative measurements of both static and dynamic mechanical behaviors in cells also were consistent with specific a priori predictions of the tensegrity model. These findings suggest that tensegrity represents a unified model of cell mechanics that may help to explain how mechanical behaviors emerge through collective interactions among different cytoskeletal filaments and extracellular adhesions in living cells.

NASA Discipline Cell Biology↗

Machine learning prediction of enzyme optimum pH

The relationship between pH and enzyme catalytic activity, especially the optimal pH (pH opt ) at which enzymes function, is critical for biotechnological applications. Hence, computational methods to predict pH opt will enhance enzyme discovery and design by facilitating accurate identification of enzymes that function optimally at specific pH levels, and by elucidating sequence-function relationships. Here, in this study, we proposed and evaluated various machine learning methods for predicting pH opt , conducting extensive hyperparameter optimization and training over 11,000 model instances. Our results demonstrate that models utilizing language model embeddings markedly outperform other methods in predicting pHopt. We present EpHod, the best-performing model, to predict pHopt, making it publicly available to researchers. From sequence data, EpHod directly learns structural and biophysical features that relate to pH opt , including proximity of residues to the catalytic centre and the accessibility of solvent molecules. Overall, EpHod presents a promising advancement in pH opt prediction and will potentially speed up the development of enzyme technologies.

97 MATHEMATICS AND COMPUTING↗

Blocking C-terminal processing of KRAS4b via a direct covalent attack on the CaaX-box cysteine

RAS is the most frequently mutated oncogene in cancer. RAS proteins show high sequence similarities in their G-domains but are significantly different in their C-terminal hypervariable regions (HVR). These regions interact with the cell membrane via lipid anchors that result from posttranslational modifications (PTM) of cysteine residues. KRAS4b is unique as it has only one cysteine that undergoes PTM, C185. Small molecule covalent modification of C185 would block any form of prenylation and subsequently inhibit attachment of KRAS4b to the cell membrane, blocking its biological activity. We translated this concept to the discovery and development of disulfide tethering screen hits into irreversible covalent modifiers of C185. These compounds inhibited proliferation of KRAS4b-driven mouse embryonic fibroblasts, but not cells driven by N-myristoylated KRAS4b that harbor a C185S mutation and are not dependent on C185 prenylation. Top–down proteomics was used to confirm target engagement in cells. These compounds bind in a pocket formed when the HVR folds back between helix 3 and 4 in the G-domain (HVR-α3-α4). This interaction can happen in the absence of small molecules as predicted by molecular dynamics simulations and is stabilized in the presence of C185 binders as confirmed by small-angle X-ray scattering and solution NMR. NOESY-HSQC, an NMR approach that measures internuclear distances of 6 Å or less, and structure analysis identified the critical residues and interactions that define the HVR-α3-α4 pocket. Further development of compounds that bind to this pocket could be the basis of a new approach to targeting KRAS cancers.

C185↗

Identification and Classification of Fungal GPCR Gene Families

G protein-coupled receptors (GPCRs) are transmembrane proteins crucial for signal transduction in eukaryotes, responding to diverse extracellular signals. Researchers have found and systematically summarized 14 distinct types of GPCRs in fungi but their distribution among numerous fungal species remained largely unexamined. Additionally, three families of mammalian homologs (Rhodopsin, Glutamate, and Frizzled) have been found in previous studies, but they are not included in the systematic classification of fungal GPCRs. Our study establishes a unified classification of 17 GPCR classes in fungi, combining 14 fungal and 3 mammalian previously recognized groups, and classifies 28,294 GPCRs across 1357 fungal species, significantly expanding the scale of GPCRs in fungi and demonstrating their broader distribution. We found that mammalian homologs are notably more prevalent in Early Diverging Fungi (EDF), whereas the previous 14 classes are predominantly found in Ascomycota and Basidiomycota. The most abundant class detected in fungi was Pth11-like GPCRs, exclusively found in Pezizomycotina and involved in fungal pathogenicity. Our analysis suggested that Pezizomycotina ancestor possessed an extensive array of Pth11-like GPCRs, but over time, some species underwent considerable reductions in these GPCRs in conjunction with genome contractions. Utilizing a custom-built convolutional neural network (CNN) for the identification of fungal GPCRs, we identified several putative novel fungal GPCRs. Predicted interactions between these prospective new GPCRs and G-alpha proteins, as simulated by AlphaFold Multimer, provided additional support for their functional relevance. In conclusion, our work defines the first large-scale, unified classification of fungal GPCRs, reveals lineage-specific expansions and contractions, and uncovers previously unrecognized GPCR candidates with potential functional roles in fungal signaling.

G protein-coupled receptors↗

Integrative SP3 Workflow for Multi-PTM Proteomics Profiling (TZ-DP0)

The goal of the experiment was to demonstrate that the optimized multiplexed multi-PTM profiling workflow can comprehensively and quantitatively capture dynamic changes in protein abundance, cysteine oxidation, phosphorylation, and acetylation in cytokine-induced inflammatory stress in mouse pancreatic ß-cells. Global proteomic, redox proteomic, phosphoproteomic, and acetylomic were data collected from mouse Beta-TC-6 pancreatic Beta-cells, untreated (mock) and cytokine-treated Beta-cells at 4, 8, and 24 hours with 4 biological replicates. Samples were digested with trypsin and Lys-C, then analyzed by LC-MS/MS. Data were searched with MS-GF+, MASIC, and MaxQuant using PNNL's DMS processing pipeline.

59 BASIC BIOLOGICAL SCIENCES↗

Metadynamics simulations reveal mechanisms of Na + and Ca 2+ transport in two open states of the channelrhodopsin chimera, C1C2

Cation conducting channelrhodopsins (ChRs) are a popular tool used in optogenetics to control the activity of excitable cells and tissues using light. ChRs with altered ion selectivity are in high demand for use in different cell types and for other specialized applications. However, a detailed mechanism of ion permeation in ChRs is not fully resolved. Here, we use complementary experimental and computational methods to uncover the mechanisms of cation transport and valence selectivity through the channelrhodopsin chimera, C1C2, in the high- and low-conducting open states. Electrophysiology measurements identified a single-residue substitution within the central gate, N297D, that increased Ca 2+ permeability vs. Na + by nearly two-fold at peak current, but less so at stationary current. We then developed molecular models of dimeric wild-type C1C2 and N297D mutant channels in both open states and calculated the PMF profiles for Na + and Ca 2+ permeation through each protein using well-tempered/multiple-walker metadynamics. Results of these studies agree well with experimental measurements and demonstrate that the pore entrance on the extracellular side differs from original predictions and is actually located in a gap between helices I and II. Cation transport occurs via a relay mechanism where cations are passed between flexible carboxylate sidechains lining the full length of the pore by sidechain swinging, like a monkey swinging on vines. In the mutant channel, residue D297 enhances Ca 2+ permeability by mediating the handoff between the central and cytosolic binding sites via direct coordination and sidechain swinging. We also found that altered cation binding affinities at both the extracellular entrance and central binding sites underly the distinct transport properties of the low-conducting open state. This work significantly advances our understanding of ion selectivity and permeation in cation channelrhodopsins and provides the insights needed for successful development of new ion-selective optogenetic tools.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Tutorial: Machine-Learning-Based CREASE-2D Analysis of 2D SAXS Profiles to Characterize Anisotropic Nanostructures in Soft Materials

We present a tutorial to guide users on how to extend the Computational Reverse Engineering Analysis of Scattering Experiments-2D (CREASE-2D) framework to interpret their experimental two-dimensional small-angle scattering (SAS) data from soft materials (e.g., polymers, peptide amphiphiles, biomolecular fibrils). Unlike most traditional SAS analysis approaches, which typically rely on azimuthally averaged onedimensional (1D) profiles, CREASE-2D utilizes the complete 2D scattering profile to reveal information about anisotropy in the structure. In past applications, CREASE has provided insights into complex structural features, including the cross-sectional shapes of assembled nanostructures and dispersity in these features, which are difficult to discern with existing analytical models. While (1D- ) CREASE has been applied to SANS and SAXS data, this tutorial shares the steps for implementing CREASE-2D using an example of a dipeptide solution system, for which we have SAXS data. We present details for these steps involved in using CREASE-2D to interpret SAXS profiles: how to preprocess SAXS data, define relevant structural features, generate three-dimensional real-space structures for specific values of these features, train a machine learning (ML) surrogate model to predict scattering profiles for given structural features, and optimize these features using genetic algorithms (GA). Then, we use these steps to interpret complex 2DSAXS data collected from dipeptide solutions that, in microscopy images, exhibit nanoscale structures that could be elliptical tubes/ flat tapes/cylinders or a combination of these cross sections. Open-source codes, computational hardware, and software requirements, as well as the strengths and limitations of this protocol, are also presented. We expect researchers working with (soft) biomaterials, peptide amphiphiles, amphiphilic polymer solutions, polymer nanocomposites, and blends of particles/polymers will find this CREASE-2D method and this tutorial of use.

CREASE↗

The Analysis of the Patterns of Radiation-Induced DNA Damage Foci by a Stochastic Monte Carlo Model of DNA Double Strand Breaks Induction by Heavy Ions and Image Segmentation Software

To create a generalized mechanistic model of DNA damage in human cells that will generate analytical and image data corresponding to experimentally observed DNA damage foci and will help to improve the experimental foci yields by simulating spatial foci patterns and resolving problems with quantitative image analysis. Material and Methods: The analysis of patterns of RIFs (radiation-induced foci) produced by low- and high-LET (linear energy transfer) radiation was conducted by using a Monte Carlo model that combines the heavy ion track structure with characteristics of the human genome on the level of chromosomes. The foci patterns were also simulated in the maximum projection plane for flat nuclei. Some data analysis was done with the help of image segmentation software that identifies individual classes of RIFs and colocolized RIFs, which is of importance to some experimental assays that assign DNA damage a dual phosphorescent signal. Results: The model predicts the spatial and genomic distributions of DNA DSBs (double strand breaks) and associated RIFs in a human cell nucleus for a particular dose of either low- or high-LET radiation. We used the model to do analyses for different irradiation scenarios. In the beam-parallel-to-the-disk-of-a-flattened-nucleus scenario we found that the foci appeared to be merged due to their high density, while, in the perpendicular-beam scenario, the foci appeared as one bright spot per hit. The statistics and spatial distribution of regions of densely arranged foci, termed DNA foci chains, were predicted numerically using this model. Another analysis was done to evaluate the number of ion hits per nucleus, which were visible from streaks of closely located foci. In another analysis, our image segmentaiton software determined foci yields directly from images with single-class or colocolized foci. Conclusions: We showed that DSB clustering needs to be taken into account to determine the true DNA damage foci yield, which helps to determine the DSB yield. Using the model analysis, a researcher can refine the DSB yield per nucleus per particle. We showed that purely geometric artifacts, present in the experimental images, can be analytically resolved with the model, and that the quantization of track hits and DSB yields can be provided to the experimentalists who use enumeration of radiation-induced foci in immunofluorescence experiments using proteins that detect DNA damage. An automated image segmentaiton software can prove useful in a faster and more precise object counting for colocolized foci images.

Ponomarev, Artem↗

Neutrons in Structural Biology: Challenges and Opportunities (Workshop Report)

Gaining a thorough understanding of biological systems requires building our knowledge about biological processes from the level of atoms and electrons, and up to whole organisms. Such comprehensive knowledge will allow for a predictive understanding of complex biological systems behavior. It will guide us in the design and development of novel therapeutics and vaccines to tackle existing health threats and to prepare for future pandemics, and it will provide information necessary to create new biomaterials and bio-inspired technologies through manipulation of biological macromolecules, their assemblies, single cells and even microorganisms. Reaching these goals will require a synergistic combination of multiple experimental techniques with molecular calculations and predictive simulations, and the design and development of new techniques and capabilities that bridge current knowledge and technology gaps. Neutron scattering provides unique information about the biomacromolecular structure and function and can play a major role in achieving these goals. A workshop was held to engage the scientific community in identifying pressing challenges in biochemistry, structural biology, enzymology and structure-guided drug design not solved with the current neutron scattering technologies or utilizing other structural biology techniques such as X-ray crystallography, NMR, and cryo-EM. The workshop brought together structural biology, biochemistry and computational experts, as well as early career researchers and students, creating a forum for discussing scientific advancement and collaboration. The workshop included a one-day satellite training workshop where graduate students and postdoctoral researchers were educated in the application of neutron crystallography and small-angle scattering in structural biology. Furthermore, the Instrument Scientific Advisory Board (ISAB) for the development of a macromolecular neutron diffractometer at ORNL’s Second Target Station was introduced at the workshop. The major outcome was that neutrons can provide atomic-level understanding of biomacromolecular structure, function and dynamics which is of paramount importance for addressing the identified challenges. Neutron crystallography, in particular, can resolve long-standing biochemical issues regarding enzyme function by delineating the underlying chemistry and can have a major impact on the design of small-molecule therapeutics, especially in combination with molecular computation (quantum chemistry and molecular dynamics simulations) and the emerging artificial intelligence (AI)-assisted drug design technologies. The unique properties of neutrons, including their high sensitivity to hydrogen and their non-destructive nature, make them ideal probes of biological matter. There is a palpable need in the scientific community to expand and enhance the impact of neutron sciences on biology. Neutron crystallography is the only structural biology method capable of determining positions of all hydrogen atoms in proteins, nucleic acids and their complexes at near-physiological temperatures and of unstable species at cryogenic temperatures. Moreover, neutron analysis is non-ionizing, non-destructive and does not perturb the structure or redox chemistry of active site metal centers and clusters in proteins, which can be invaluable for studying radiation-sensitive metalloprotein complexes. Further, neutron energies used in scattering applications are similar to atomic motions, permitting neutron spectroscopies to characterize the dynamics of biomacromolecules on the picosecond to microsecond timescales. The different sensitivities of neutrons to protium (H) and deuterium (D) isotopes of hydrogen allow enhanced visibility of specific parts of biological complexes through isotopic labeling. The impact of neutrons will be most powerful when neutron scattering is combined with complementary experimental techniques that use photons and electrons, and with high-performance computing. The interconnection and mutuality of the experimental and theoretical capabilities will drive discoveries in biological and health sciences to generate more complete picture of complex biological systems. The major limitation in the field of biological neutron crystallography has been signal-to-noise, demanding large samples that are difficult to produce for the majority of biomacromolecules and limiting the applicability of this technique in biological sciences. A neutron crystallography instrument at the Second Target Station will revolutionize biological science with neutrons by engaging a large scientific community of structural biologists, enabling successful neutron diffraction experiments from radically smaller biomacromolecular crystals, resolving unanswered biochemical questions, and meaningfully contributing to rational drug design. The meeting highlighted 10 grand challenges that will be addressed with this advanced capability over the next decade and beyond, and the recommendations required to help address them are given below.

59 BASIC BIOLOGICAL SCIENCES↗

Using Strong Magnetic Fields to Control Solutal Convection

An important component in biotechnology, particularly in the area of protein engineering and rational drug design is the knowledge of the precise three-dimensional molecular structure of proteins. The quality of structural information obtained from X-ray diffraction methods is directly dependent on the degree of perfection of the protein crystals. As a consequence, the growth of high quality macromolecular crystals for diffraction analyses has been the central focus for biochemists, biologists, and bioengineers. Macromolecular crystals are obtained from solutions that contain the crystallizing species in equilibrium with higher aggregates, ions, precipitants, other possible phases of the protein, foreign particles, the walls of the container, and a likely host of other impurities. By changing transport modes in general, i.e., reduction of convection and sedimentation, as is achieved in microgravity , we have been able to dramatically affect the movement and distribution of macromolecules in the fluid, and thus their transport, formation of crystal nuclei, and adsorption to the crystal surface. While a limited number of high quality crystals from space flights have been obtained, as the recent National Research Council (NRC) review of the NASA microgravity crystallization program pointed out, the scientific approach and research in crystallization of proteins has been mainly empirical yielding inconclusive results. We postulate that we can reduce convection in ground-based experiments and we can understand the different aspects of convection control through the use of strong magnetic fields and field gradients. We postulate that limited convection in a magnetic field will provide the environment for the growth of high quality crystals. The approach exploits the variation of fluid magnetic susceptibility with concentration for this purpose and the convective damping is realized by appropriately positioning the crystal growth cell so that the magnetic susceptibility force counteracts terrestrial gravity. The general objective is to test the hypothesis of convective control using a strong magnetic field and magnetic field gradient and to understand the nature of the various forces that come into play. Specifically we aim to delineate causative factors and to quantify them through experiments, analysis and numerical modeling. The paper will report on the experimental results using paramagnetic salts and solutions in magnetic fields and compare them to analytical predictions.

Ramachandran, N.↗

Substrate-Directed Dimensional and Phase Control of Peptide Assemblies on Two-Dimensional van der Waals Materials

Understanding and controlling biomolecular self-assembly on van der Waals (vdW) materials has the potential to advance hybrid bioelectronic devices by enabling precise tuning of the interface and modulation of the resulting electronic properties of the biomolecule-vdW heterostructure. However, how surface properties of vdW materials direct biomolecule assembly remains poorly understood. To fill this knowledge gap, we investigated the assembly of a peptide known to assemble into two-dimensional (2D) crystalline films on MoS 2 on three representative vdW surfaces: WS 2 , MoS 2 , and highly oriented pyrolytic graphite (HOPG). Using in situ atomic force microscopy (AFM), we find that assembly is substrate-dependent, resulting in multilayers on WS 2 , monolayers on MoS 2 , and multiple coexisting phases on HOPG. WS 2 exhibits a higher negative charge, strong long-range electrostatic interactions, and extensive hydration layering that may promote multilayer stacking. In contrast, MoS 2 has stronger short-range interactions with the peptides but much weaker long-range interactions and hydration structure, which may favor monolayer formation. Molecular dynamics simulations predict a corresponding switch from monolayer to multilayer aggregates of the adsorbed monomers, reflected in their relative mobilities. On hydrophobic HOPG, the peptides bind most strongly and remain as monomers with high surface mobility. The peptide dimers comprising the basic unit of the crystals are more compact on HOPG, which has a smaller lattice constant than WS 2 or MoS 2 , suggesting strain contributes to stabilizing multiple phases. Our results provide mechanistic insights into how surface charge and hydration structure, and the lattice structure of the substrates governs peptide assembly on vdW materials, offering a framework to rationally control the 2D peptide-vdW heterostructures.

Molecular dynamics simulations↗

A Structurally Diverse Compound Screening Library to Identify Substrates for Diamine, Polyamine, and Related Acetyltransferases

Spermidine/spermine N-acetyltransferases (SSATs) and other types of polyamine acetyltransferases (PAATs) acetylate diamines and/or polyamines. These enzymes are evolutionarily related and belong to the Gcn5-related N-acetyltransferase (GNAT) superfamily, yet we lack a fundamental understanding of their substrate specificity and/or promiscuity toward different compounds. Many of these enzymes are known or are predicted to acetylate polyamines, but in the cell there are other types of compounds that contain moieties derived from polyamines that may be the native substrates for these enzymes. To learn more about the identity of substrates that are acetylated, we selected and screened 17 different GNAT enzymes for activity toward a set of structurally diverse compounds that contained different types of amine moieties (e.g., aminopropyl, aminobutyl, etc.). These compounds included diamines, triamines, and polyamines containing primary amino groups, and they had structural diversity with variation of the chain length and presence or absence of internal amino groups and other functional groups. We found 12 of the 17 enzymes acetylated at least one of the compounds. Some enzymes were selective toward acetylating only one compound while others exhibited substrate promiscuity toward numerous compounds. Our experimental results ultimately allowed us to pinpoint specific substrates that could be further investigated to more fully understand substrate specificity versus promiscuity of GNAT enzymes and the role of acetylated small molecules in cells.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Genes encoding calmodulin-binding proteins in the Arabidopsis genome

Analysis of the recently completed Arabidopsis genome sequence indicates that approximately 31% of the predicted genes could not be assigned to functional categories, as they do not show any sequence similarity with proteins of known function from other organisms. Calmodulin (CaM), a ubiquitous and multifunctional Ca(2+) sensor, interacts with a wide variety of cellular proteins and modulates their activity/function in regulating diverse cellular processes. However, the primary amino acid sequence of the CaM-binding domain in different CaM-binding proteins (CBPs) is not conserved. One way to identify most of the CBPs in the Arabidopsis genome is by protein-protein interaction-based screening of expression libraries with CaM. Here, using a mixture of radiolabeled CaM isoforms from Arabidopsis, we screened several expression libraries prepared from flower meristem, seedlings, or tissues treated with hormones, an elicitor, or a pathogen. Sequence analysis of 77 positive clones that interact with CaM in a Ca(2+)-dependent manner revealed 20 CBPs, including 14 previously unknown CBPs. In addition, by searching the Arabidopsis genome sequence with the newly identified and known plant or animal CBPs, we identified a total of 27 CBPs. Among these, 16 CBPs are represented by families with 2-20 members in each family. Gene expression analysis revealed that CBPs and CBP paralogs are expressed differentially. Our data suggest that Arabidopsis has a large number of CBPs including several plant-specific ones. Although CaM is highly conserved between plants and animals, only a few CBPs are common to both plants and animals. Analysis of Arabidopsis CBPs revealed the presence of a variety of interesting domains. Our analyses identified several hypothetical proteins in the Arabidopsis genome as CaM targets, suggesting their involvement in Ca(2+)-mediated signaling networks.

NASA Discipline Plant Biology↗

Empirical evidence that glucan-interacting amino acid side chains within the transmembrane channel collectively facilitate cellulose synthase function

The fundamental mechanism of cellulose synthesis is widely conserved across Kingdoms and depends on cellulose synthases, which are processive, dual-function, family 2 glycosyltransferases (GT-2). These enzymes polymerize glucose on the cytoplasmic side of the plasma membrane and export the glucan chain to the cell surface through an integral transmembrane (TM) channel. Structural studies of active plant cellulose synthases (CESAs) have revealed interactions between the nascent glucan chain and the side chains of polar, charged, and aromatic amino acid residues that line the TM channel. However, the functional consequences of modifying these side chains have not been tested in vivo in CESAs or other processive GT-2s. To test this, we used an established in vivo assay based on genetic complementation of CESA5 in the moss, Physcomitrium patens. For accurate prediction of glucan-interacting amino acid residues, we generated a complete homotrimeric molecular model of PpCESA5 using a combination of homology and de novo modeling. All-atom molecular dynamics-based analyses of contact metrics and interaction energy identified 23 amino acid residues with high propensity to interact with the nascent glucan chain within the TM channel or on the apoplastic surface of PpCESA5. Mutating any one of 18 of these amino acid residues to alanine, thereby removing their side chains, abolished or impaired CESA function, with the strongest effects observed upon the loss of charged amino acid side chains. This provides direct evidence to support the hypothesis that multiple amino acid residues collectively maintain a smooth energy landscape within the TM channel to facilitate glucan translocation.

59 BASIC BIOLOGICAL SCIENCES↗

Integrative analysis of the 3D genome and epigenome in mouse embryonic tissues

While a rich set of putative cis-regulatory sequences involved in mouse fetal development have been annotated recently on the basis of chromatin accessibility and histone modification patterns, delineating their role in developmentally regulated gene expression continues to be challenging. To fill this gap, here we mapped chromatin contacts between gene promoters and distal sequences across the genome in seven mouse fetal tissues and across six developmental stages of the forebrain. We identified 248,620 long-range chromatin interactions centered at 14,138 protein-coding genes and characterized their tissue-to-tissue variations and developmental dynamics. Integrative analysis of the interactome with previous epigenome and transcriptome datasets from the same tissues revealed a strong correlation between the chromatin contacts and chromatin state at distal enhancers, as well as gene expression patterns at predicted target genes. We predicted target genes of 15,098 candidate enhancers and used them to annotate target genes of homologous candidate enhancers in the human genome that harbor risk variants of human diseases. We present evidence that schizophrenia and other adult disease risk variants are frequently found in fetal enhancers, providing support for the hypothesis of fetal origins of adult diseases.

59 BASIC BIOLOGICAL SCIENCES↗