Search NASA⌕ Search

SEARCH · Search NASA

Results for “protein data bank”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

rcsb-api : Python Toolkit for Streamlining Access to RCSB Protein Data Bank APIs

The Protein Data Bank (PDB) was founded in 1971 as the first open-access digital data resource in biology to serve as the single global archive for three-dimensional (3D) macromolecular structure data. Current PDB holdings exceed 230,000 experimentally determined structures of proteins, nucleic acids, viruses, and macromolecular machines. The RCSB Protein Data Bank RCSB.org research-focused web portal facilitates search, analyses, and visualization of every PDB structure along with more than one million Computed Structure Models from AlphaFold DB and the ModelArchive. It is powered by a set of publicly available Application Programming Interfaces (APIs) that both support RCSB.org users and provide programmatic access to PDB data. Given the breadth and levels of granularity encompassed in this rich data collection, efficiently accessing the information programmatically may be challenging for new users. RCSB PDB has developed a Python software package, rcsb-api , that facilitates easy and efficient use of RCSB PDB APIs within a Python environment. This software tool is designed to streamline access to the extensive corpus of data housed within the PDB, enabling researchers to search, retrieve, and analyze 3D biostructure data seamlessly. Its use will accelerate research in structural biology, molecular biology and biochemistry, drug discovery, and bioinformatics by providing more efficient tools for data integration and analysis. The new toolkit is available on GitHub (github.com/rcsb/py-rcsb-api) and published to the public Python package repository (PyPI) to foster wider usage and support basic and applied research in fundamental biology, biomedicine, and the energy sciences.

FAIR principles↗

Updated resources for exploring experimentally-determined PDB structures and Computed Structure Models at the RCSB Protein Data Bank

The Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB, RCSB.org), the US Worldwide Protein Data Bank (wwPDB, wwPDB.org) data center for the global PDB archive, provides access to the PDB data via its RCSB.org research-focused web portal. We report substantial additions to the tools and visualization features available at RCSB.org, which now delivers more than 227000 experimentally determined atomic-level three-dimensional (3D) biostructures stored in the global PDB archive alongside more than 1 million Computed Structure Models (CSMs) of proteins (including models for human, model organisms, select human pathogens, crop plants and organisms important for addressing climate change). In addition to providing support for 3D structure motif searches with user-provided coordinates, new features highlighted herein include query results organized by redundancy-reduced Groups and summary pages that facilitate exploration of groups of similar proteins. Newly released programmatic tools are also described, as are enhanced training opportunities.

Burley, Stephen K.↗

RCSB Protein Data Bank: supporting research and education worldwide through explorations of experimentally determined and computationally predicted atomic level 3D biostructures

The Protein Data Bank (PDB) was established as the first open-access digital data resource in biology and medicine in 1971 with seven X-ray crystal structures of proteins. Today, the PDB houses >210 000 experimentally determined, atomic level, 3D structures of proteins and nucleic acids as well as their complexes with one another and small molecules ( e.g. approved drugs, enzyme cofactors). These data provide insights into fundamental biology, biomedicine, bioenergy and biotechnology. They proved particularly important for understanding the SARS-CoV-2 global pandemic. The US-funded Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) and other members of the Worldwide Protein Data Bank (wwPDB) partnership jointly manage the PDB archive and support >60 000 `data depositors' (structural biologists) around the world. wwPDB ensures the quality and integrity of the data in the ever-expanding PDB archive and supports global open access without limitations on data usage. The RCSB PDB research-focused web portal at https://www.rcsb.org/ (RCSB.org) supports millions of users worldwide, representing a broad range of expertise and interests. In addition to retrieving 3D structure data, PDB `data consumers' access comparative data and external annotations, such as information about disease-causing point mutations and genetic variations. RCSB.org also provides access to >1 000 000 computed structure models (CSMs) generated using artificial intelligence/machine-learning methods. To avoid doubt, the provenance and reliability of experimentally determined PDB structures and CSMs are identified. Related training materials are available to support users in their RCSB.org explorations.

59 BASIC BIOLOGICAL SCIENCES↗

RCSB protein data Bank: Next‐generation advanced search for exploration of experimental structures and computed structure models

Abstract The Protein Data Bank (PDB), established in 1971, is the primary global, open‐access archive for experimentally determined 3D macromolecular structures (proteins, RNA, DNA). The research‐focused RCSB.org web‐portal provides access to these data alongside more than one million machine‐learning‐predicted structure models, greatly expanding the available structural landscape. Rapid growth of both experimental and computational structures has increased the need for powerful yet accessible search tools that serve a broad and diverse scientific community. Herein, we describe a redesigned RCSB Protein Data Bank RCSB.org Advanced Search capability that supports intuitive discovery of 3D structures through a unified interface. This interface integrates annotation‐, sequence‐, and 3D structure‐based searches, embeds an interactive 3D viewer, and incorporates curated biological knowledge, such as catalytic site definitions from Mechanism and Catalytic Site Atlas and ligand‐guided structural motifs, for constructing geometry‐driven queries. A new Chemical Search tool allows definition of chemical queries via an integrated drawing tool or standard identifiers, seamlessly combining them with annotation filters. By allowing query definition directly within spatial and chemical contexts, these search interfaces reduce the need for detailed knowledge of residue numbering, chain identifiers, or external cheminformatics software. This capability enables efficient exploration of structures, chemical diversity, and structure–function relationships across all life domains. The redesigned interfaces can be accessed directly at rcsb.org/search/advanced for Advanced Search and rcsb.org/search/chemical for Chemical Search.

Rose, Yana [Research Collaboratory for Structural ↗

Protein data bank: From two epidemics to the global pandemic to mRNA vaccines and Paxlovid

Structural biologists and the open-access Protein Data Bank (PDB) played decisive roles in combating the COVID-19 pandemic. Global biostructure data were turned into global knowledge, allowing scientists and engineers to understand the inner workings of coronaviruses and develop effective countermeasures. Two mRNA vaccines, initially designed with guidance from PDB structures of the SARS-CoV-1 and MERS-CoV spike proteins, prevented infections entirely or reduced the likelihood of morbidity and mortality for more than five billion individual recipients worldwide. Structure-guided drug discovery by Pfizer, Inc (facilitated by PDB structures), initiated in the 2000s in response to SARS-CoV-1 and resumed in 2020, yielded nirmatrelvir (the active ingredient of Paxlovid) -- a potent, orally-bioavailable inhibitor of the SARS-CoV-2 main protease. You've got to love the Protein Data Bank!

Burley, Stephen K.↗

Impact of structural biology and the protein data bank on us fda new drug approvals of low molecular weight antineoplastic agents 2019–2023

Abstract Open access to three-dimensional atomic-level biostructure information from the Protein Data Bank (PDB) facilitated discovery/development of 100% of the 34 new low molecular weight, protein-targeted, antineoplastic agents approved by the US FDA 2019–2023. Analyses of PDB holdings, the scientific literature, and related documents for each drug-target combination revealed that the impact of structural biologists and public-domain 3D biostructure data was broad and substantial, ranging from understanding target biology (100% of all drug targets), to identifying a given target as likely druggable (100% of all targets), to structure-guided drug discovery (>80% of all new small-molecule drugs, made up of 50% confirmed and >30% probable cases). In addition to aggregate impact assessments, illustrative case studies are presented for six first-in-class small-molecule anti-cancer drugs, including a selective inhibitor of nuclear export targeting Exportin 1 (selinexor, Xpovio), an ATP-competitive CSF-1R receptor tyrosine kinase inhibitor (pexidartinib,Turalia), a non-ATP-competitive inhibitor of the BCR-Abl fusion protein targeting the myristoyl binding pocket within the kinase catalytic domain of Abl (asciminib, Scemblix), a covalently-acting G12C KRAS inhibitor (sotorasib, Lumakras or Lumykras), an EZH2 methyltransferase inhibitor (tazemostat, Tazverik), and an agent targeting the basic-Helix-Loop-Helix transcription factor HIF-2α (belzutifan, Welireg).

60 APPLIED LIFE SCIENCES↗

Assessing Metal Ion Assignment Accuracy in Protein Data Bank Models via Elemental Spectroscopy

Accurate representation of metal ions in macromolecular structures is critical for chemical interpretation, computational modeling, and machine-learning methods that rely on Protein Data Bank (PDB) entries. However, the elemental identity of metals modeled in crystallographic structures is often inferred indirectly and rarely validated experimentally. Here, we combine Particle Induced X-ray Emission (PIXE) and X-ray Fluorescence Spectroscopy (XRFS) to determine the elemental composition of protein samples used to generate 70 deposited metalloprotein crystal structures. By analyzing the original protein material employed for crystallization, but before the addition of crystallization buffer solutions, we assess whether the modeled metal ions in deposited structures are consistent with experimentally detectable elemental content. We find that in a majority of cases, the metals modeled in the corresponding PDB entries are inconsistent with the metals present in the protein samples before crystallization, or that additional metals are present but not represented in the structural models. Spectroscopic results were integrated with automated crystallographic validation metrics, including real-space Z-difference (RSZD) analysis and systematic rerefinement, to evaluate atomic-number mismatch at metal sites. PIXE and XRFS show strong agreement for dominant elemental signals and provide complementary, scalable approaches for identifying suspect metal assignments. This work does not address physiological or functional metalation but instead highlights a widespread data integrity issue in deposited macromolecular structures, PDB-wide. These results establish an experimentally corroborated link between elemental identity and crystallographic validation metrics, enabling the large-scale detection of chemically inconsistent annotations in structural databases used for computational modeling and machine learning.

Crystallization↗

A new chapter for RCSB Protein Data Bank Molecule of the Month in 2025

The online Molecule of the Month series authored by David S. Goodsell and published by the Research Collaboratory for Structural Biology Protein Data Bank at PDB101.RCSB.org has highlighted stories about the biomolecular structures driving fundamental biology, biomedicine, bioenergy, and biotechnology since January 2000. A new chapter begins in 2025: Janet Iwasa has taken over as the series creator of stories about critically important biological macromolecules in a rapidly changing world.

Bioenergy↗

Protein Data Bank (PDB): Fifty-three years young and having a transformative impact on science and society

This review article describes the co-evolution of structural biology as a discipline and the Protein Data Bank (PDB), established in 1971 as the first open-access data resource in biology by like-minded structural scientists. As the PDB archive grew in size and scope to encompass macromolecular crystallography, NMR spectroscopy, and cryo-electron microscopy, new technologies were developed to ingest, validate, curate, store, and distribute the information. Community engagement ensured that the needs of structural biologists (data depositors) and data consumers were met. Today, the archive houses more than 230,000 experimentally determined structures of proteins, nucleic acids, and macromolecular machines and their complexes with one another and small-molecule ligands. Aggregate costs of PDB data preservation are ~1% of the cost of structure determination. The enormous impact of PDB data on basic and applied research and education across the natural and medical sciences is presented and highlighted with illustrative examples. Enablement of de novo protein structure prediction (AlphaFold2, RoseTTAfold, OpenFold, etc.) is the most widely appreciated benefit of having a corpus of rigorously validated, expertly curated 3D biostructure data.

bioinformatics↗

Visualizing and analyzing 3D biomolecular structures using Mol* at RCSB.org: Influenza A H5N1 virus proteome case study

The easiest and often most useful way to work with experimentally determined or computationally predicted structures of biomolecules is by viewing their three-dimensional (3D) shapes using a molecular visualization tool. Mol* was collaboratively developed by RCSB Protein Data Bank (RCSB PDB, RCSB.org) and Protein Data Bank in Europe (PDBe, PDBe.org) as an open-source, web-based, 3D visualization software suite for examination and analyses of biostructures. It is capable of displaying atomic coordinates and related experimental data of biomolecular structures together with a variety of annotations, facilitating basic and applied research, training, education, and information dissemination. Across RCSB.org, the RCSB PDB research-focused web portal, Mol* has been implemented to support single-mouse-click atomic-level visualization of biomolecules (e.g., proteins, nucleic acids, carbohydrates) with bound cofactors, small-molecule ligands, ions, water molecules, or other macromolecules. RCSB.org Mol* can seamlessly display 3D structures from various sources, allowing structure interrogation, superimposition, and comparison. Using influenza A H5N1 virus as a topical case study of an important pathogen, we exemplify how Mol* has been embedded within various RCSB.org tools—allowing users to view polymer sequence and structure-based annotations integrated from trusted bioinformatics data resources, assess patterns and trends in groups of structures, and view structures of any size and compositional complexity. In addition to being linked to every experimentally determined biostructure and Computed Structure Model made available at RCSB.org, Standalone Mol* is freely available for visualizing any atomic-level or multi-scale biostructure at rcsb.org/3d-view.

3D biostructure↗

Community recommendations on cryoEM data archiving and validation

In January 2020, a workshop was held at EMBL-EBI (Hinxton, UK) to discuss data requirements for the deposition and validation of cryoEM structures, with a focus on single-particle analysis. The meeting was attended by 47 experts in data processing, model building and refinement, validation, and archiving of such structures. This report describes the workshop's motivation and history, the topics discussed, and the resulting consensus recommendations. Some challenges for future methods-development efforts in this area are also highlighted, as is the implementation to date of some of the recommendations.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

SparcleQC: Automated Input File Creation for QM/MM Studies of Protein:Ligand Complexes

SparcleQC is a Python package that, given a protein:ligand complex in the Protein Data Bank (PDB) file format, can create quantum mechanics/molecular mechanics (QM/MM)-like input files for the electronic structure theory packages PSI4, QChem, and NWChem. The resulting input files include quantum mechanical representations of the ligand and a small section of the protein, surrounded by point charges that represent the rest of the protein. Creation of these QM/MM input files includes cutting and capping the QM subregion, obtaining point charges for the protein, and adjusting charges at the QM/MM boundary; and each of these tasks are automated by the software. In this article, we describe the details of SparcleQC’s procedure, show examples of the Python API, and explain additional features that are helpful in protein:ligand interaction studies. Finally, we show that SparcleQC enables automated preparation of input files for QM/MM calculations, which can return can return accurate interaction energies in minutes, while a fully quantum mechanical computation on the protein:ligand complex could take days, if it is even possible.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Predicting the Functional State of Protein Kinases Using Interpretable Graph Neural Networks

Kinases are a family of proteins that function as molecular switches, regulating several essential cellular activities such as cell proliferation. Dysfunctional kinases are implicated in several types of cancers and hence they are actively pursued as drug targets. Given the vast number of complex kinase structures that are available in the protein data bank (PDB), there is a necessity to develop methodologies that can identify structurally important moieties of the kinases in an automated fashion, for such techniques can be instrumental in identifying novel drug targets. In this work, we develop a graph neural network (GNN) based deep learning framework for classifying the functionally active and inactive states of a large set of eukaryotic protein kinases, making use of their 3D structure from the PDB. We show that GNN based machine learning models can classify protein states with an accuracy greater than 97%. We further use the GNN models to automatically identify regions of the kinases that are important for its function. For this purpose, Gradient-weighted Class Activation Mapping (Grad-CAM) was implemented on the protein graphs. Remarkably, Grad-CAM consistently identifies the highly conserved DFG motif as the most important part of the protein across the entire kinome, without any prior input. Other regions of the hydrophobic core such as the HRD motif were also identified by the interpretable GNN framework, consistent with the literature. We discuss the significance of each of these regions in detail.

Ashwin Ravichandran↗

Deciphering the Cofilin Oligomers via Intermolecular Disulfide Bond Formation: A Coarse-Grained Molecular Dynamics Approach to Understanding Cofilin’s Regulation on Actin Filaments

Cofilin, a key actin-binding protein, orchestrates the dynamics of the actomyosin network through its actin-severing activity and by promoting the recycling of actin monomers. Recent experimental work suggests that cofilin also forms functionally distinct oligomers through thiol post-translational modification (PTM) that encourages actin nucleation and assembly. Despite these advances, the structural conformations of cofilin oligomers that modulate actin activity remain elusive because there are combinatorial ways to oxidize thiols in cysteines to form disulfide bonds rapidly. This study employs molecular dynamics simulations to investigate human cofilin 1 as a case study for exploring cofilin dimers via disulfide bond formation. Using the free energy profiling, our simulations unveil a range of probable cofilin dimer structures not represented in current Protein Data Bank entries. These candidate dimers are characterized by their distinct population distributions and relative free energies. Of particular note is a dimer featuring an interface between cysteines 139 and 147 residues, which demonstrates stable free energy characteristics and intriguingly symmetrical geometry. In contrast, the experimentally proposed dimer structure exhibits a less stable free energy profile. Here, we also evaluate frustration quantification based on the energy landscape theory in the protein-protein interactions at the dimer interfaces. Notably, the 39-39 dimer configuration emerges as a promising candidate for forming cofilin tetramers, as substantiated by frustration analysis. Additionally, docking simulations with actin filaments further evaluate the stability of these cofilin dimer-actin complexes. Our findings thus offer a computational framework for understanding the role of thiol post-translational modification cofilin proteins in regulating oligomerization, and the subsequent cofilin-mediated actin dynamics in the actomyosin network.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Geometry, spin coupling, and dielectric control of redox potentials in [4Fe–4S] Clusters

Iron–sulfur (Fe–S) clusters are common biological cofactors that facilitate vital redox reactions. Despite extensive research, the molecular basis of redox potential tuning in ferredoxin-like proteins remains an active area of debate. In this study, we combine statistical analysis of over one thousand [4Fe–4S]-containing protein structures from the Protein Data Bank (PDB) with broken-symmetry and extended broken-symmetry density functional theory to examine how cysteine ligand orientations and environmental screening affect redox properties of the clusters. We identified five main ligand configurations, three of which are predominant in natural structures. Among these, the adiabatic electron affinity differs by less than 0.1 V, indicating that, while geometry plays a secondary role, it allows localized fine-tuning of redox properties. In contrast, electrostatic and solvation effects primarily determine the overall potential range.

Computational Chemistry↗

Development of a TSR-based method for understanding structural relationships of cofactors and local environments in photosystem I

All chemical forms of energy and oxygen on Earth are generated via photosynthesis where light energy is converted into redox energy by two photosystems (PS I and PS II). There is an increasing number of PS I 3D structures deposited in the Protein Data Bank (PDB). The Triangular Spatial Relationship (TSR)-based algorithm converts 3D structures into integers (TSR keys). A comprehensive study was conducted, by taking advantage of the PS I 3D structures and the TSR-based algorithm, to answer three questions: (i) Are electron cofactors including P700, A -1 and A 0 , which are chemically identical chlorophylls, structurally different? (ii) There are two electron transfer chains (A and B branches) in PS I. Are the cofactors on both branches structurally different? (iii) Are the amino acids in cofactor binding sites structurally different from those not in cofactor binding sites? The key contributions and important findings include: (i) a novel TSR-based method for representing 3D structures of pigments as well as for quantifying pigment structures was developed; (ii) the results revealed that the redox cofactor, P700, are structurally conserved and different from other redox factors. Similar situations were also observed for both A -1 and A 0 ; (iii) the results demonstrated structural differences between A and B branches for the redox cofactors P700, A -1 , A 0 and A 1 as well as their cofactor binding sites; (iv) the tryptophan residues close to A 0 and A 1 are structurally conserved; (v) The TSR-based method outperforms the Root Mean Square Deviation (RMSD) and the Ultrafast Shape Recognition (USR) methods. The structural analyses of redox cofactors and their binding sites provide a foundation for understanding the unique chemical and physical properties of each redox cofactor in PS I, which are essential for modulating the rate and direction of energy and electron transfers.

59 BASIC BIOLOGICAL SCIENCES↗