Search NASA⌕ Search

SEARCH · Search NASA

Results for “protein structure file”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

CHARMM-GUI Drude prepper for molecular dynamics simulation using the classical Drude polarizable force field

Explicit treatment of electronic polarizability in empirical force fields (FFs) represents an extension over a traditional additive or pairwise FF and provides a more realistic model of the variations in electronic structure in condensed phase, macromolecular simulations. To facilitate utilization of the polarizable FF based on the classical Drude oscillator model, Drude Prepper has been developed in CHARMM-GUI. Drude Prepper ingests additive CHARMM protein structures file (PSF) and pre-equilibrated coordinates in CHARMM, PDB, or NAMD format, from which the molecular components of the system are identified. These include all residues and patches connecting those residues along with water, ions, and other solute molecules. This information is then used to construct the Drude FF-based PSF using molecular generation capabilities in CHARMM, followed by minimization and equilibration. In addition, inputs are generated for molecular dynamics (MD) simulations using CHARMM, GROMACS, NAMD, and OpenMM. Validation of the Drude Prepper protocol and inputs is performed through conversion and MD simulations of various heterogeneous systems that include proteins, nucleic acids, lipids, polysaccharides, and atomic ions using the aforementioned simulation packages. Stable simulations are obtained in all studied systems, including 5 μs simulation of ubiquitin, verifying the integrity of the generated Drude PSFs. Additionally, the ability of the Drude FF to model variations in electronic structure is shown through dipole moment analysis in selected systems. Finally, the capabilities and availability of Drude Prepper in CHARMM-GUI is anticipated to greatly facilitate the application of the Drude FF to a range of condensed phase, macromolecular systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Proteome-scale Structure Prediction Data - Pseudodesulfovibrio mercurii

The number of proteins predicted for Pseudodesulfovibrio mercurii is 3,446, each of which have five predicted structures from an AlphaFold run, as well as structural alignment results using the TMscore-based structural alignment method within the APoc program. Specifically, AlphaFold outputs the atoms and coordinates of the protein model in human-readable PDB files and quantitative prediction metrics in Python PICKLE files. The 5 models have been ranked based on the predicted TM-score (pTMS), a quantitative confidence metric output by AlphaFold that reports on protein model quality. The top ranked model has undergone an energy minimization calculation to relax and remove any potential clashes in the atomic coordinates. Structural alignment results are stored in two files for each protein; the top ranked model (as discussed above) is used for all alignment analyses. Both are compressed gzip files that, once unpacked, are human readable. The first file is the TMalign score results and contains the quantitative metrics for the top alignments between the predicted structure and experimental structures from the PDB70, a curated non-redundant database of about 80,000 experimental structures developed by the Soding lab. Each data point in this file is directly associated with one experimental structure; PDB ID and brief meta-data about the protein taken from the PDB70 file are reported alongside the quantitative metrics. The second results file contains the raw results associated with each alignment reported in the score results file. Specifically, the translation and rotation arrays for each alignment are provided so that the structural alignment can be recreated. Additionally, residue-level scores are reported to quantify the closeness of the aligned residues between the predicted and experimental models.

59 BASIC BIOLOGICAL SCIENCES↗

Structural Models and Sequence Alignment Results of the Rhodospirillum rubrum Proteome

This dataset contains the structural models for the primary transcripts of the Rhodospirillum rubrum proteome as well as sequence alignment results for a subset of the encoded proteins. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the R. rubrum proteome to those available in the AlphaFold Protein Structure Database. For proteins that have been annotated as hypothetical, sequence alignment results from the HHblits and SAdLSA alignment methods are provided. These methods are often more capable to resolve sequence homology than other methods. Therefore, the results from both HHblits and SAdLSA are provided to identify possible homologs for these challenging proteins. Numerous sequence databases are utilized for these alignments. References AlphaFold v2 Multimer: https://doi.org/10.1101/2021.10.04.463034. References HHBlits: https://doi.org/10.1186/s12859-019-3019-7. References SAdLSA: https://doi.org/10.3389/fbinf.2021.689960.

59 BASIC BIOLOGICAL SCIENCES↗

Structural Models and Sequence Alignment Results of the Desulfovibrio vulgaris Proteome

This dataset contains the structural models for the primary transcripts of the Desulfovibrio vulgaris proteome as well as sequence alignment results for a subset of the encoded proteins. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the D. vulgaris proteome to those available in the AlphaFold Protein Structure Database (AFDB). This is a bit more complicated since the proteins reporting in the AFDB originate from an outdated form of the D. vulgaris sequence. The different versions of the D. vulgaris gene annotation are collected in the Chronology subdirectory; further consideration of these changes on the structural space of the proteome are currently underway. For proteins that have been annotated as hypothetical, sequence alignment results from the HHblits and SAdLSA alignment methods are provided. These methods are often more capable to resolve sequence homology than other methods. Therefore, the results from both HHblits and SAdLSA are provided to identify possible homologs for these challenging proteins. Numerous sequence databases are utilized for these alignments. References AlphaFold v2 Multimer: https://doi.org/10.1101/2021.10.04.463034. References HHblits: hhtps://doi.org/10.1186/s12859-019-3019-7. References SAdLSA: hhtps://doi.org/10.3389/fbinf.2021.689960.

59 BASIC BIOLOGICAL SCIENCES↗

Structural Models of the Rhodopseudomonas palustris Proteome

This dataset contains the structural models for the primary transcripts of the Rhodopseudomonas palustris proteome. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the R. palustris proteome to those available in the AlphaFold Protein Structure Database.

59 BASIC BIOLOGICAL SCIENCES↗

SparcleQC: Automated Input File Creation for QM/MM Studies of Protein:Ligand Complexes

SparcleQC is a Python package that, given a protein:ligand complex in the Protein Data Bank (PDB) file format, can create quantum mechanics/molecular mechanics (QM/MM)-like input files for the electronic structure theory packages PSI4, QChem, and NWChem. The resulting input files include quantum mechanical representations of the ligand and a small section of the protein, surrounded by point charges that represent the rest of the protein. Creation of these QM/MM input files includes cutting and capping the QM subregion, obtaining point charges for the protein, and adjusting charges at the QM/MM boundary; and each of these tasks are automated by the software. In this article, we describe the details of SparcleQC’s procedure, show examples of the Python API, and explain additional features that are helpful in protein:ligand interaction studies. Finally, we show that SparcleQC enables automated preparation of input files for QM/MM calculations, which can return can return accurate interaction energies in minutes, while a fully quantum mechanical computation on the protein:ligand complex could take days, if it is even possible.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Algorithms and file structures to extend and enhance liquid chromatography and ion mobility mass spectrometry workflows (CRADA Final Report)

The purpose of this project was to continue supporting customizations of algorithms and raw data file structures to enhance software workflows for liquid chromatography (LC), mass spectrometry (MS) and ion mobility mass spectrometry (IM-MS)-based protein and metabolite characterization. PNNL worked with Agilent to design, implement, evaluate, and demonstrate new algorithms and integrated them as functionalities into the PNNL-PreProcessor software. The project augmented PNNL’s capabilities to analyze complex proteomics and metabolomics samples. These capabilities are directly beneficial to DOE and PNNL efforts to characterize and analyze these compounds in microbial and plant communities. The project assisted Agilent in further developing improved instrument-software solutions combining liquid chromatography and ion mobility with mass spectrometry for widespread applications in life sciences and other fields.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Validated ligand geometries for macromolecular refinement restraints and molecular-mechanics force fields

In macromolecular structure refinement, the low observation-to-parameter ratio and the lack of high-resolution data are countered by using a priori information in the form of restraints. Having accurate geometries of the chemical entities in the sample is paramount for generating accurate chemical restraints and, therefore, accurate macromolecular structures. In particular, it is desirable to have accurate restraints for known and novel ligand entities. Quantum mechanics (QM) can minimize the energy of a ligand by adjusting its geometry, and these geometries can be used to generate restraints for macromolecular refinement. This article describes a library of approximately 37 000 small molecules extracted from the Chemical Component Dictionary in the Protein Data Bank and minimized by density-functional QM. The library includes restraint files for use in crystallography or cryo-EM refinement, along with files suitable for molecular-dynamics simulation. Because the geometries are validated using the Cambridge Structural Database, the restraints library provides users with both functional restraints and minimized geometries. This work also provides procedures for generating new and accurate restraints.

Amber↗

Unified Nanotechnology Format: One Way to Store Them All

The domains of DNA and RNA nanotechnology are steadily gaining in popularity while proving their value with various successful results, including biosensing robots and drug delivery cages. Nowadays, the nanotechnology design pipeline usually relies on computer-based design (CAD) approaches to design and simulate the desired structure before the wet lab assembly. To aid with these tasks, various software tools exist and are often used in conjunction. However, their interoperability is hindered by a lack of a common file format that is fully descriptive of the many design paradigms. Therefore, in this paper, we propose a Unified Nanotechnology Format (UNF) designed specifically for the biomimetic nanotechnology field. UNF allows storage of both design and simulation data in a single file, including free-form and lattice-based DNA structures. By defining a logical and versatile format, we hope it will become a widely accepted and used file format for the nucleic acid nanotechnology community, facilitating the future work of researchers and software developers. Together with the format description and publicly available documentation, we provide a set of converters from existing file formats to simplify the transition. Finally, we present several use cases visualizing example structures stored in UNF, showcasing the various types of data UNF can handle.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

phosaa14SB and phosaa19SB: Updated Amber Force Field Parameters for Phosphorylated Amino Acids

Phosphorylated amino acids are involved in many cell regulatory networks; proteins containing these post-translational modifications are widely studied both experimentally and computationally. Simulations are used to investigate a wide range of structural and dynamic properties of biomolecules, such as ligand binding, enzyme-reaction mechanisms, and protein folding. However, the development of force field parameters for the simulation of proteins containing phosphorylated amino acids using the Amber program has not kept pace with the development of parameters for standard amino acids, and it is challenging to model these modified amino acids with accuracy comparable to proteins containing only standard amino acids. In particular, the popular ff14SB and ff19SB models do not contain parameters for phosphorylated amino acids. Here, the dihedral parameters for the side chains of the most common phosphorylated amino acids are trained against reference data from QM calculations adopting the ff14SB approach, followed by validation against experimental data. Finally, library files and corresponding parameter files are provided, with versions that are compatible with both ff14SB and ff19SB.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

PPI DataHub Project Data Package: High-density Lipoprotein (HDL) Structure and Function Proteomics

The purpose of this experiment was to investigate how the interactions between APOA1 and APOA2 on the surface of high-density lipoproteins (HDL) impact particle function. Interactions were investigated on HDL isolated from human blood plasma using structural proteomics tools such as chemical cross-linking and limited proteolysis (LiP). The structural proteomics data was acquired using a Q-Exactive HF-X mass spectrometer and data was processed and compiled using MaxQuant sofware (v.1.6.17.0). Processed datasets are openly accessible from the download button (~2.8 GB) and contain secondary processed LiP and global proteomic results files and supporting metadata materials. Processed data downloads include a sample naming key, processed MaxQuant results/parameters, and protein annotated relative abundance files.

59 BASIC BIOLOGICAL SCIENCES↗

PPI DataHub Project Data Package: S. elongatus PCC 7942 Limited Proteolysis and Thermal Proteome Profiling Structural Proteomics (JM-PB-DP3)

The purpose of this experiment was to investigate structural alterations in proteins involved in central carbon metabolism and photosynthetic electron transfer pathways in Synechococcus elongatus PCC 7942. Sample data was obtained from S. elongatus cell lysates using three complementary mass spectrometry (MS) techniques using limited proteolysis (LiP-MS), thermal proteome profiling (TPP-MS), and redox enrichment (Redox-MS) in evaluating alterations solvent accessibility and structural stability caused by light perturbation at the molecular level. Experimentally processed sample data for LiP and TPP proteomic datasets were derived from the same cell culture stock, prepared simultaneously in parallel, and acquired by mass spectrometry. Processed datasets are openly accessible from the download button and contain secondary processed proteomic results files, computed outputs, and supporting metadata materials. Experimental samples processed for LiP-MS label-free quantification (LFQ) or TPP-MS tandem mass tag (TMT) 10-plex were acquired using a Q-Exactive HF-X mass spectrometer and processed/compiled using either MSGF+ (v2024.03.26) or ​​​​PlexedPiper for proteome evaluation. Additional software supporting downstream proteomic analysis include FragPipe (v.4.0), MSFragger (v.22.1), and an adapted Microbial Isolate LiP Analysis Workflow (located at Zenodo). Processed proteomic data downloads include a sample naming key, normalized quantification results files, and processed protein annotated abundance files.

59 BASIC BIOLOGICAL SCIENCES↗

pKPDB: a protein data bank extension database of p Ka and pI theoretical values

Abstract Summary pKa values of ionizable residues and isoelectric points of proteins provide valuable local and global insights about their structure and function. These properties can be estimated with reasonably good accuracy using Poisson–Boltzmann and Monte Carlo calculations at a considerable computational cost (from some minutes to several hours). pKPDB is a database of over 12 M theoretical pKa values calculated over 120k protein structures deposited in the Protein Data Bank. By providing precomputed pKa and pI values, users can retrieve results instantaneously for their protein(s) of interest while also saving countless hours and resources that would be spent on repeated calculations. Furthermore, there is an ever-growing imbalance between experimental pKa and pI values and the number of resolved structures. This database will complement the experimental and computational data already available and can also provide crucial information regarding buried residues that are under-represented in experimental measurements. Availability and implementation Gzipped csv files containing p Ka and isoelectric point values can be downloaded from https://pypka.org/pKPDB. To query a single PDB code please use the PypKa free server at https://pypka.org. The pKPDB source code can be found at https://github.com/mms-fcul/pKPDB. Supplementary information Supplementary data are available at Bioinformatics online.

Reis, Pedro B. P. S. (ORCID:0000000335636239)↗

TeraChem protocol buffers ( TCPB ): Accelerating QM and QM/MM simulations with a client–server model

The routine use of electronic structures in many chemical simulation applications calls for efficient and easy ways to access electronic structure programs. Here, we describe how the graphics processing unit (GPU) accelerated electronic structure program TeraChem can be set up as an electronic structure server, to be easily accessed by third-party client programs. We exploit Google’s protocol buffer framework for data serialization and communication. The client interface, called TeraChem protocol buffers (TCPB), has been designed for ease of use and compatibility with multiple programming languages, such as C++, Fortran, and Python. To demonstrate the ease of coupling third-party programs with electronic structures using TCPB, we have incorporated the TCPB client into Amber for quantum mechanics/molecular mechanics (QM/MM) simulations. The TCPB interface saves time with GPU initialization and I/O operations, achieving a speedup of more than 2× compared to a prior file-based implementation for a QM region with ~250 basis functions. We demonstrate the practical application of TCPB by computing the free energy profile of p-hydroxybenzylidene-2,3-dimethylimidazolinone (p-HBDI - )—a model chromophore in green fluorescent proteins—on the first excited singlet state using Hamiltonian replica exchange for enhanced sampling. All calculations in this work have been performed with the non-commercial freely-available version of TeraChem, which is sufficient for many QM region sizes in common use.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

ezAlign: A Tool for Converting Coarse-Grained Molecular Dynamics Structures to Atomistic Resolution for Multiscale Modeling

Soft condensed matter is challenging to study due to the vast time and length scales that are necessary to accurately represent complex systems and capture their underlying physics. Multiscale simulations are necessary to study processes that have disparate time and/or length scales, which abound throughout biology and other complex systems. Herein we present ezAlign, an open-source software for converting coarse-grained molecular dynamics structures to atomistic representation, allowing multiscale modeling of biomolecular systems. The ezAlign v1.1 software package is publicly available for download at github.com/LLNL/ezAlign. Its underlying methodology is based on a simple alignment of an atomistic template molecule, followed by position-restraint energy minimization, which forces the atomistic molecule to adopt a conformation consistent with the coarse-grained molecule. The molecules are then combined, solvated, minimized, and equilibrated with position restraints. Validation of the process was conducted on a pure POPC membrane and compared with other popular methods to construct atomistic membranes. Additional examples, including surfactant self-assembly, membrane proteins, and more complex bacterial and human plasma membrane models, are also presented. By providing these examples, parameter files, code, and an easy-to-follow recipe to add new molecules, this work will aid future multiscale modeling efforts.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

PDB‐101: Molecular Explorations through Biology and Medicine

PDB‐101 is an online portal for teachers, students, and the general public to promote exploration of the structural biology of proteins and nucleic acids ( pdb101.rcsb.org ). Learning about the diverse shapes and functions of these biological macromolecules helps to understand all aspects of biomedicine and agriculture, from protein synthesis to health and disease to biological energy. Why PDB‐101? Researchers around the world are studying these molecules at the atomic level. These 3D structures are freely available at the Protein Data Bank (PDB), the central storehouse of biomolecular structures. This website builds introductory materials to help beginners get started in the basics of biomolecular structure and function (“101”, as in an entry level course) as well as resources for extended learning. Since 2011, PDB‐101 has been developed by the RCSB PDB , a global resource for the advancement of research and education in biology and medicine. Along with our Worldwide PDB collaborators, RCSB PDB curates, annotates, and makes publicly available the PDB data deposited by scientists around the globe. The RCSB PDB then provides a window to these data through a rich online resource with powerful searching, reporting, and visualization tools for researchers. This information is then streamlined for students and teachers at PDB‐101. Features include the ongoing Molecule of the Month series, educational materials such as paper models, posters, molecular animations, educational curricula and more. The section “Guide to Understanding PDB Data” is a primer for detailed PDB‐specific information: PDB Data, Visualizing Structures, Reading Coordinate Files, scientific methods for structure determination, and more. PDB‐101 also runs annual Video Challenges for high school students. Participants create short videos that tell molecular stories that connect structural biology and medicine. Previous topics have included HIV/AIDS, diabetes, and antimicrobial resistance. The 2022 challenge will focus on Molecular Mechanisms of Cancer. PDB‐101 activities are evaluated using user surveys, feedback from in‐person activities, and website analytics. In 2020, PDB‐101 hosted >850,000 users and >2.6 million page views.

Zardecki, Christine↗

SAXS Assistant: Automated SAXS analysis for structural discovery in biologics and polymeric nanoparticles

Small-angle x-ray scattering (SAXS) is a powerful technique for assessing macromolecular structure. High-throughput SAXS is limited by the time-consuming and, at times, subjective nature of SAXS data interpretation. Here, we present SAXS Assistant, a Python-based script that streamlines SAXS data analysis to extract features for machine learning (ML) and key structural parameters, including the Guinier radius of gyration (R g ), pair distance distribution function (PDDF)-derived R g , maximum particle dimension (D max ), and Kratky plots. The script builds upon BioXTAS RAW and validates reliability via Guinier/PDDF R g agreement, an important indicator of well-measured data sets. For assistance in D max estimation, a multilayer perceptron regressor was trained with 1940 data files from the Small Angle Scattering Biological Data Bank. The model achieved a test set performance R 2 = 0.90 and mean absolute error = 11.7 Å. Training exclusively with experimental data translates analyses from researchers, including experts in the field, to the ML model, which helps assess D max estimations from PDDF. Gaussian mixture model clustering was implemented to classify profiles into structural classes based on entries in the Small Angle Scattering Biological Data Bank. Users may therefore assess the similarity between experimental samples and known biomolecular shapes within the mapped repository entries. This probabilistic clustering aids in quantifying information from Kratky and generating shape-descriptive features. SAXS Assistant accelerates SAXS data analysis through enforced quality control, ML-ready outputs, and flags for low-confidence results. In addition to providing the ability to analyze large data sets at high throughput, this tool is versatile and may serve researchers in both biological and synthetic polymer research fields.

36 MATERIALS SCIENCE↗