Search NASA⌕ Search

SEARCH · Search NASA

Results for “Structural Bioinformatics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

ZMPY3D: accelerating protein structure volume analysis through vectorized 3D Zernike moments and Python-based GPU integration

Abstract Motivation Volumetric 3D object analyses are being applied in research fields such as structural bioinformatics, biophysics, and structural biology, with potential integration of artificial intelligence/machine learning (AI/ML) techniques. One such method, 3D Zernike moments, has proven valuable in analyzing protein structures (e.g., protein fold classification, protein–protein interaction analysis, and molecular dynamics simulations). Their compactness and efficiency make them amenable to large-scale analyses. Established methods for deriving 3D Zernike moments, however, can be inefficient, particularly when higher order terms are required, hindering broader applications. As the volume of experimental and computationally-predicted protein structure information continues to increase, structural biology has become a “big data” science requiring more efficient analysis tools. Results This application note presents a Python-based software package, ZMPY3D, to accelerate computation of 3D Zernike moments by vectorizing the mathematical formulae and using graphical processing units (GPUs). The package offers popular GPU-supported libraries such as CuPy and TensorFlow together with NumPy implementations, aiming to improve computational efficiency, adaptability, and flexibility in future algorithm development. The ZMPY3D package can be installed via PyPI, and the source code is available from GitHub. Volumetric-based protein 3D structural similarity scores and transform matrix of superposition functionalities have both been implemented, creating a powerful computational tool that will allow the research community to amalgamate 3D Zernike moments with existing AI/ML tools, to advance research and education in protein structure bioinformatics. Availability and implementation ZMPY3D, implemented in Python, is available on GitHub (https://github.com/tawssie/ZMPY3D) and PyPI, released under the GPL License.

Lai, Jhih-Siang (ORCID:0000000156775890)↗

Library Screening, In Vivo Confirmation, and Structural and Bioinformatic Analysis of Pentapeptide Sequences as Substrates for Protein Farnesyltransferase

Protein farnesylation is a post-translational modification where a 15-carbon farnesyl isoprenoid is appended to the C-terminal end of a protein by farnesyltransferase (FTase). This process often causes proteins to associate with the membrane and participate in signal transduction pathways. The most common substrates of FTase are proteins that have C-terminal tetrapeptide CaaX box sequences where the cysteine is the site of modification. However, recent work has shown that five amino acid sequences can also be recognized, including the pentapeptides CMIIM and CSLMQ. In this work, peptide libraries were initially used to systematically vary the residues in those two parental sequences using an assay based on Matrix Assisted Laser Desorption Ionization–Mass Spectrometry (MALDI-MS). In addition, 192 pentapeptide sequences from the human proteome were screened using that assay to discover additional extended CaaaX-box motifs. Selected hits from that screening effort were rescreened using an in vivo yeast reporter protein assay. The X-ray crystal structure of CMIIM bound to FTase was also solved, showing that the C-terminal tripeptide of that sequence interacted with the enzyme in a similar manner as the C-terminal tripeptide of CVVM, suggesting that the tripeptide comprises a common structural element for substrate recognition in both tetrapeptide and pentapeptide sequences. Molecular dynamics simulation of CMIIM bound to FTase further shed light on the molecular interactions involved, showing that a putative catalytically competent Zn(II)-thiolate species was able to form. Bioinformatic predictions of tetrapeptide (CaaX-box) reactivity correlated well with the reactivity of pentapeptides obtained from in vivo analysis, reinforcing the importance of the C-terminal tripeptide motif. This analysis provides a structural framework for understanding the reactivity of extended CaaaX-box motifs and a method that may be useful for predicting the reactivity of additional FTase substrates bearing CaaaX-box sequences.

59 BASIC BIOLOGICAL SCIENCES↗

RCSB Protein Data Bank: supporting research and education worldwide through explorations of experimentally determined and computationally predicted atomic level 3D biostructures

The Protein Data Bank (PDB) was established as the first open-access digital data resource in biology and medicine in 1971 with seven X-ray crystal structures of proteins. Today, the PDB houses >210 000 experimentally determined, atomic level, 3D structures of proteins and nucleic acids as well as their complexes with one another and small molecules ( e.g. approved drugs, enzyme cofactors). These data provide insights into fundamental biology, biomedicine, bioenergy and biotechnology. They proved particularly important for understanding the SARS-CoV-2 global pandemic. The US-funded Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) and other members of the Worldwide Protein Data Bank (wwPDB) partnership jointly manage the PDB archive and support >60 000 `data depositors' (structural biologists) around the world. wwPDB ensures the quality and integrity of the data in the ever-expanding PDB archive and supports global open access without limitations on data usage. The RCSB PDB research-focused web portal at https://www.rcsb.org/ (RCSB.org) supports millions of users worldwide, representing a broad range of expertise and interests. In addition to retrieving 3D structure data, PDB `data consumers' access comparative data and external annotations, such as information about disease-causing point mutations and genetic variations. RCSB.org also provides access to >1 000 000 computed structure models (CSMs) generated using artificial intelligence/machine-learning methods. To avoid doubt, the provenance and reliability of experimentally determined PDB structures and CSMs are identified. Related training materials are available to support users in their RCSB.org explorations.

59 BASIC BIOLOGICAL SCIENCES↗

Updated resources for exploring experimentally-determined PDB structures and Computed Structure Models at the RCSB Protein Data Bank

The Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB, RCSB.org), the US Worldwide Protein Data Bank (wwPDB, wwPDB.org) data center for the global PDB archive, provides access to the PDB data via its RCSB.org research-focused web portal. We report substantial additions to the tools and visualization features available at RCSB.org, which now delivers more than 227000 experimentally determined atomic-level three-dimensional (3D) biostructures stored in the global PDB archive alongside more than 1 million Computed Structure Models (CSMs) of proteins (including models for human, model organisms, select human pathogens, crop plants and organisms important for addressing climate change). In addition to providing support for 3D structure motif searches with user-provided coordinates, new features highlighted herein include query results organized by redundancy-reduced Groups and summary pages that facilitate exploration of groups of similar proteins. Newly released programmatic tools are also described, as are enhanced training opportunities.

Burley, Stephen K.↗

Environmental Contributions to Proton Sharing in Protein Low-Barrier Hydrogen Bonds

Hydrogen bonds (H-bonds) are central to biomolecular structure and dynamics. Although H-bonds are typically characterized by well-defined proton positions, proton delocalization has been proposed to play a role in facilitating enzyme catalysis and allostery in some systems. Experimentally locating protons is difficult, hampering the study of proton mobility in H-bonds. We used neutron crystallography, atomic resolution X-ray bond length analysis, and large quantum mechanics/molecular mechanics-Born–Oppenheimer molecular dynamics (QM/MM-BOMD) simulations to comprehensively characterize the shared proton/deuteron in a Glu–Asp low-barrier hydrogen bond (LBHB) in the bacterial protein YajL that is a conventional H-bond in the homologous disease-associated human protein DJ-1. X-ray bond length analysis of protiated and perdeuterated DJ-1 and YajL shows no significant effect of deuteron substitution on these carboxylic acid-carboxylate H-bonds but does reveal an effect at the active site glutamic acid near a cysteine thiolate. Residues in an H-bonded network that might favor LBHB formation in YajL were interrogated by the mutation of homologous residues in DJ-1. A distal DJ-1 substitution increases proton delocalization in the Glu–Asp H-bond, demonstrating that mutations within extended H-bond networks can modulate proton transfer barriers in carboxylic acid-carboxylate H-bonds. In addition, proton mobility in the H-bond is correlated with dimer-spanning motions in the QM/MM-BOMD simulations of YajL and DJ-1. Our results show that proton delocalization can be tuned using combined bioinformatic, structural, and computational information, opening the possibility of using engineered proton delocalization as a probe of H-bonding environments and as a tool to test hypotheses about LBHB function.

Lin, Jiusheng [University of Nebraska, Lincoln, NE↗

The influence of protein electrostatics on potential inversion in flavoproteins

Biology uses relatively few electron-transfer cofactors, tuning their potentials, electronic couplings, and reorganization energies to carry out the required chemistry. It is remarkable that the potential ordering of two-electron transfer active flavins can be normal (first oxidation at low potential and second oxidation at high potential) or inverted, and the gap between the potentials can be as large as one volt. Analysis based on structural bioinformatics and electrostatics indicates that the ordering of the flavin redox potential is influenced by protein electrostatics. In all 36 flavoproteins examined, the introduction of a negative charge near the flavin in silico increases the extent of potential inversion (by lowering the electrochemical potential of the second electron-transfer step); the introduction of a positive charge near the flavin favors normally ordered potentials. We also find that the addition of positive charges increases the electrochemical potential for the naturally occurring one-electron transition in flavodoxins (between deprotonated hydroquinone and neutral semiquinone) and also increases the second one-electron transition in bifurcating flavins (between anionic semiquinone and fully oxidized flavin). Finally, we find that proximity of a proton acceptor, notably conserved arginine, supports proton-coupled electron transfer because it may act as a proton acceptor, promoting potential inversion. This key arginine residue may enable two-electron transfer chemistry by promoting the proton-coupled electron transfer process over the pure electron transfer process, suggesting how a protein's flavin environment may influence one- or two-electron chemistry in flavoproteins.

Singh, Niven [Duke Univ., Durham, NC (United State↗

Autotrophic biofilms sustained by deeply sourced groundwater host diverse bacteria implicated in sulfur and hydrogen metabolism

Abstract Background Biofilms in sulfide-rich springs present intricate microbial communities that play pivotal roles in biogeochemical cycling. We studied chemoautotrophically based biofilms that host diverse CPR bacteria and grow in sulfide-rich springs to investigate microbial controls on biogeochemical cycling. Results Sulfide springs biofilms were investigated using bulk geochemical analysis, genome-resolved metagenomics, and scanning transmission X-ray microscopy (STXM) at room temperature and 87 K. Chemolithotrophic sulfur-oxidizing bacteria, including Thiothrix and Beggiatoa , dominate the biofilms, which also contain CPR Gracilibacteria, Absconditabacteria, Saccharibacteria, Peregrinibacteria, Berkelbacteria, Microgenomates, and Parcubacteria. STXM imaging revealed ultra-small cells near the surfaces of filamentous bacteria that may be CPR bacterial episymbionts. STXM and NEXAFS spectroscopy at carbon K and sulfur L 2,3 edges show that filamentous bacteria contain protein-encapsulated spherical elemental sulfur granules, indicating that they are sulfur oxidizers, likely Thiothrix . Berkelbacteria and Moranbacteria in the same biofilm sample are predicted to have a novel electron bifurcating group 3b [NiFe]-hydrogenase, putatively a sulfhydrogenase, potentially linked to sulfur metabolism via redox cofactors. This complex could potentially contribute to symbioses, for example, with sulfur-oxidizing bacteria such as Thiothrix that is based on cryptic sulfur cycling. One Doudnabacteria genome encodes adjacent sulfur dioxygenase and rhodanese genes that may convert thiosulfate to sulfite. We find similar conserved genomic architecture associated with CPR bacteria from other sulfur-rich subsurface ecosystems. Conclusions Our combined metagenomic, geochemical, spectromicroscopic, and structural bioinformatics analyses of biofilms growing in sulfide-rich springs revealed consortia that contain CPR bacteria and sulfur-oxidizing Proteobacteria, including Thiothrix , and bacteria from a new family within Beggiatoales. We infer roles for CPR bacteria in sulfur and hydrogen cycling.

59 BASIC BIOLOGICAL SCIENCES↗

RCSB protein data Bank: Next‐generation advanced search for exploration of experimental structures and computed structure models

Abstract The Protein Data Bank (PDB), established in 1971, is the primary global, open‐access archive for experimentally determined 3D macromolecular structures (proteins, RNA, DNA). The research‐focused RCSB.org web‐portal provides access to these data alongside more than one million machine‐learning‐predicted structure models, greatly expanding the available structural landscape. Rapid growth of both experimental and computational structures has increased the need for powerful yet accessible search tools that serve a broad and diverse scientific community. Herein, we describe a redesigned RCSB Protein Data Bank RCSB.org Advanced Search capability that supports intuitive discovery of 3D structures through a unified interface. This interface integrates annotation‐, sequence‐, and 3D structure‐based searches, embeds an interactive 3D viewer, and incorporates curated biological knowledge, such as catalytic site definitions from Mechanism and Catalytic Site Atlas and ligand‐guided structural motifs, for constructing geometry‐driven queries. A new Chemical Search tool allows definition of chemical queries via an integrated drawing tool or standard identifiers, seamlessly combining them with annotation filters. By allowing query definition directly within spatial and chemical contexts, these search interfaces reduce the need for detailed knowledge of residue numbering, chain identifiers, or external cheminformatics software. This capability enables efficient exploration of structures, chemical diversity, and structure–function relationships across all life domains. The redesigned interfaces can be accessed directly at rcsb.org/search/advanced for Advanced Search and rcsb.org/search/chemical for Chemical Search.

Rose, Yana [Research Collaboratory for Structural ↗

rcsb-api : Python Toolkit for Streamlining Access to RCSB Protein Data Bank APIs

The Protein Data Bank (PDB) was founded in 1971 as the first open-access digital data resource in biology to serve as the single global archive for three-dimensional (3D) macromolecular structure data. Current PDB holdings exceed 230,000 experimentally determined structures of proteins, nucleic acids, viruses, and macromolecular machines. The RCSB Protein Data Bank RCSB.org research-focused web portal facilitates search, analyses, and visualization of every PDB structure along with more than one million Computed Structure Models from AlphaFold DB and the ModelArchive. It is powered by a set of publicly available Application Programming Interfaces (APIs) that both support RCSB.org users and provide programmatic access to PDB data. Given the breadth and levels of granularity encompassed in this rich data collection, efficiently accessing the information programmatically may be challenging for new users. RCSB PDB has developed a Python software package, rcsb-api , that facilitates easy and efficient use of RCSB PDB APIs within a Python environment. This software tool is designed to streamline access to the extensive corpus of data housed within the PDB, enabling researchers to search, retrieve, and analyze 3D biostructure data seamlessly. Its use will accelerate research in structural biology, molecular biology and biochemistry, drug discovery, and bioinformatics by providing more efficient tools for data integration and analysis. The new toolkit is available on GitHub (github.com/rcsb/py-rcsb-api) and published to the public Python package repository (PyPI) to foster wider usage and support basic and applied research in fundamental biology, biomedicine, and the energy sciences.

FAIR principles↗

Geometry-complete perceptron networks for 3D molecular graphs

Abstract Motivation The field of geometric deep learning has recently had a profound impact on several scientific domains such as protein structure prediction and design, leading to methodological advancements within and outside of the realm of traditional machine learning. Within this spirit, in this work, we introduce GCPNet, a new chirality-aware SE(3)-equivariant graph neural network designed for representation learning of 3D biomolecular graphs. We show that GCPNet, unlike previous representation learning methods for 3D biomolecules, is widely applicable to a variety of invariant or equivariant node-level, edge-level, and graph-level tasks on biomolecular structures while being able to (1) learn important chiral properties of 3D molecules and (2) detect external force fields. Results Across four distinct molecular-geometric tasks, we demonstrate that GCPNet’s predictions (1) for protein–ligand binding affinity achieve a statistically significant correlation of 0.608, more than 5%, greater than current state-of-the-art methods; (2) for protein structure ranking achieve statistically significant target-local and dataset-global correlations of 0.616 and 0.871, respectively; (3) for Newtownian many-body systems modeling achieve a task-averaged mean squared error less than 0.01, more than 15% better than current methods; and (4) for molecular chirality recognition achieve a state-of-the-art prediction accuracy of 98.7%, better than any other machine learning method to date. Availability and implementation The source code, data, and instructions to train new models or reproduce our results are freely available at https://github.com/BioinfoMachineLearning/GCPNet.

59 BASIC BIOLOGICAL SCIENCES↗

Identification and preliminary characterization of conserved uncharacterized proteins from Chlamydomonas reinhardtii , Arabidopsis thaliana , and Setaria viridis

Abstract The rapid accumulation of sequenced plant genomes in the past decade has outpaced the still difficult problem of genome‐wide protein‐coding gene annotation. A substantial fraction of protein‐coding genes in all plant genomes are poorly annotated or unannotated and remain functionally uncharacterized. We identified unannotated proteins in three model organisms representing distinct branches of the green lineage (Viridiplantae): Arabidopsis thaliana (eudicot), Setaria viridis (monocot), and Chlamydomonas reinhardtii (Chlorophyte alga). Using similarity searching, we identified a subset of unannotated proteins that were conserved between these species and defined them as Deep Green proteins. Bioinformatic, genomic, and structural predictions were performed to begin classifying Deep Green genes and proteins. Compared to whole proteomes for each species, the Deep Green set was enriched for proteins with predicted chloroplast targeting signals predictive of photosynthetic or plastid functions, a result that was consistent with enrichment for daylight phase diurnal expression patterning. Structural predictions using AlphaFold and comparisons to known structures showed that a significant proportion of Deep Green proteins may possess novel folds. Though only available for three organisms, the Deep Green genes and proteins provide a starting resource of high‐value targets for further investigation of potentially new protein structures and functions conserved across the green lineage.

59 BASIC BIOLOGICAL SCIENCES↗

Visualizing and analyzing 3D biomolecular structures using Mol* at RCSB.org: Influenza A H5N1 virus proteome case study

The easiest and often most useful way to work with experimentally determined or computationally predicted structures of biomolecules is by viewing their three-dimensional (3D) shapes using a molecular visualization tool. Mol* was collaboratively developed by RCSB Protein Data Bank (RCSB PDB, RCSB.org) and Protein Data Bank in Europe (PDBe, PDBe.org) as an open-source, web-based, 3D visualization software suite for examination and analyses of biostructures. It is capable of displaying atomic coordinates and related experimental data of biomolecular structures together with a variety of annotations, facilitating basic and applied research, training, education, and information dissemination. Across RCSB.org, the RCSB PDB research-focused web portal, Mol* has been implemented to support single-mouse-click atomic-level visualization of biomolecules (e.g., proteins, nucleic acids, carbohydrates) with bound cofactors, small-molecule ligands, ions, water molecules, or other macromolecules. RCSB.org Mol* can seamlessly display 3D structures from various sources, allowing structure interrogation, superimposition, and comparison. Using influenza A H5N1 virus as a topical case study of an important pathogen, we exemplify how Mol* has been embedded within various RCSB.org tools—allowing users to view polymer sequence and structure-based annotations integrated from trusted bioinformatics data resources, assess patterns and trends in groups of structures, and view structures of any size and compositional complexity. In addition to being linked to every experimentally determined biostructure and Computed Structure Model made available at RCSB.org, Standalone Mol* is freely available for visualizing any atomic-level or multi-scale biostructure at rcsb.org/3d-view.

3D biostructure↗

Chemical classification program synthesis using generative artificial intelligence

Accurately classifying chemical structures is essential for cheminformatics and bioinformatics, including tasks such as identifying bioactive compounds of interest, screening molecules for toxicity to humans, finding non-organic compounds with desirable material properties, or organizing large chemical libraries for drug discovery or environmental monitoring. However, manual classification is labor-intensive and difficult to scale to large chemical databases. Existing automated approaches either rely on manually constructed classification rules, or are deep learning methods that lack explainability. This work presents an approach that uses generative artificial intelligence to automatically write chemical classifier programs for classes in the Chemical Entities of Biological Interest (ChEBI) database. These programs can be used for efficient deterministic run-time classification of SMILES structures, with natural language explanations. The programs themselves constitute an explainable computable ontological model of chemical class nomenclature, which we call the ChEBI Chemical Class Program Ontology (C3PO). We validated our approach against the ChEBI database, and compared our results against deep learning models and a naive SMARTS pattern based classifier. C3PO outperforms the naive classifier, but does not reach the performance of state of the art deep learning methods. However, C3PO has a number of strengths that complement deep learning methods, including explainability and reduced data dependence. C3PO can be used alongside deep learning classifiers to provide an explanation of the classification, where both methods agree. The programs can be used as part of the ontology development process, and iteratively refined by expert human curators.

Artificial Intelligence↗

Linking Spatiotemporal Biological Data to Predict Harmful Algal Blooms

Cyanobacterial Harmful Algal Blooms (cHABs) have significant impacts on an affected region’s economy, ecology, and human health. The blooms can release toxins that kill fish and poison water for people and animals. The global adverse effects of cHABs are exacerbated by the consequences of climate change and increased pollution. Though the phenomena are well documented, scientists’ efforts to mitigate the damage are hampered by insufficient predictive models and incomplete granular knowledge of cHAB community structure. With a goal of leveraging bioinformatics and machine learning tools to better understand and predict cHABs, we are first exploring water sample data sets. Using nearly four thousand samples from the National Center for Biotechnology Information Sequence Read Archive (NCBI-SRA) across 16 years with latitude and longitude embedded in the metadata, we mapped the location of the samples onto a Lake Erie shape file. We combined information about location, date, and community taxa in the NCBI samples to discover factors that determine cHAB features. The data are separated into three distinct zones, with the majority pooled at the southwest end of the lake and occurring in 2017. The samples are rich in biological data; our next steps are to carry out whole genome sequence analysis and use the community profiles as part of our predictive machine learning model.

59 BASIC BIOLOGICAL SCIENCES↗

Specificity determinants revealed by the structure of glycosyltransferase Campylobacter concisus PglA

Abstract In selected Campylobacter species, the biosynthesis of N‐linked glycoconjugates via the pgl pathway is essential for pathogenicity and survival. However, most of the membrane‐associated GT‐B fold glycosyltransferases responsible for diversifying glycans in this pathway have not been structurally characterized which hinders the understanding of the structural factors that govern substrate specificity and prediction of resulting glycan composition. Herein, we report the 1.8 Å resolution structure of Campylobacter concisus PglA, the glycosyltransferase responsible for the transfer of N ‐acetylgalatosamine (GalNAc) from uridine 5′‐diphospho‐ N ‐acetylgalactosamine (UDP‐GalNAc) to undecaprenyl‐diphospho‐ N , N ′‐diacetylbacillosamine (UndPP‐diNAcBac) in complex with the sugar donor GalNAc. This study identifies distinguishing characteristics that set PglA apart within the GT4 enzyme family. Computational docking of the structure in the membrane in comparison to homologs points to differences in interactions with the membrane‐embedded acceptor and the structural analysis of the complex together with bioinformatics and site‐directed mutagenesis identifies donor sugar binding motifs. Notably, E113, conserved solely among PglA enzymes, forms a hydrogen bond with the GalNAc C6″‐OH. Mutagenesis of E113 reveals activity consistent with this role in substrate binding, rather than stabilization of the oxocarbenium ion transition state, a function sometimes ascribed to the corresponding residue in GT4 homologs. The bioinformatic analyses reveal a substrate‐specificity motif, showing that Pro281 in a substrate binding loop of PglA directs configurational preference for GalNAc over GlcNAc. This proline is replaced by a conformationally flexible glycine, even in distant homologs, which favor substrates with the same stereochemistry at C4, such as glucose. The signature loop is conserved across all Campylobacter PglA enzymes, emphasizing its importance in substrate specificity.

Vuksanovic, Nemanja↗