Search NASA⌕ Search

SEARCH · Search NASA

Results for “protein structure file”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

CHARMM-GUI Drude prepper for molecular dynamics simulation using the classical Drude polarizable force field

Explicit treatment of electronic polarizability in empirical force fields (FFs) represents an extension over a traditional additive or pairwise FF and provides a more realistic model of the variations in electronic structure in condensed phase, macromolecular simulations. To facilitate utilization of the polarizable FF based on the classical Drude oscillator model, Drude Prepper has been developed in CHARMM-GUI. Drude Prepper ingests additive CHARMM protein structures file (PSF) and pre-equilibrated coordinates in CHARMM, PDB, or NAMD format, from which the molecular components of the system are identified. These include all residues and patches connecting those residues along with water, ions, and other solute molecules. This information is then used to construct the Drude FF-based PSF using molecular generation capabilities in CHARMM, followed by minimization and equilibration. In addition, inputs are generated for molecular dynamics (MD) simulations using CHARMM, GROMACS, NAMD, and OpenMM. Validation of the Drude Prepper protocol and inputs is performed through conversion and MD simulations of various heterogeneous systems that include proteins, nucleic acids, lipids, polysaccharides, and atomic ions using the aforementioned simulation packages. Stable simulations are obtained in all studied systems, including 5 μs simulation of ubiquitin, verifying the integrity of the generated Drude PSFs. Additionally, the ability of the Drude FF to model variations in electronic structure is shown through dipole moment analysis in selected systems. Finally, the capabilities and availability of Drude Prepper in CHARMM-GUI is anticipated to greatly facilitate the application of the Drude FF to a range of condensed phase, macromolecular systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

SARS-CoV2 Docking Dataset

Description: Small-molecule conformations and docking scores for 1.4 billion molecules docked against 6 protein targets from SARS-CoV2: MPro 5R84, MPro 6WQF, NSP15 6WLC, PLPro 7JIR, Spike 6M0J, and a hand-optimized model of the RNA-dependent RNA polymerase. Docking was carried out using the Autodock-GPU program performing 20 independent structure minimizations per dock - saving 3 results per molecule. Scores reported include the Autodock free energy estimate as well as RF3 and VS-DUD-E v2 machine-learned rescoring models. Protein structure files and maps in the format input to Autodock-GPU are included. Literature Ref: Supercomputer-Based Ensemble Docking Drug Discovery Pipeline with Application to Covid-19, J. Chem. Inf. Model. 2020, 60(12): 5832–5852.

36 MATERIALS SCIENCE↗

Proteome-scale Structure Prediction Data - Pseudodesulfovibrio mercurii

The number of proteins predicted for Pseudodesulfovibrio mercurii is 3,446, each of which have five predicted structures from an AlphaFold run, as well as structural alignment results using the TMscore-based structural alignment method within the APoc program. Specifically, AlphaFold outputs the atoms and coordinates of the protein model in human-readable PDB files and quantitative prediction metrics in Python PICKLE files. The 5 models have been ranked based on the predicted TM-score (pTMS), a quantitative confidence metric output by AlphaFold that reports on protein model quality. The top ranked model has undergone an energy minimization calculation to relax and remove any potential clashes in the atomic coordinates. Structural alignment results are stored in two files for each protein; the top ranked model (as discussed above) is used for all alignment analyses. Both are compressed gzip files that, once unpacked, are human readable. The first file is the TMalign score results and contains the quantitative metrics for the top alignments between the predicted structure and experimental structures from the PDB70, a curated non-redundant database of about 80,000 experimental structures developed by the Soding lab. Each data point in this file is directly associated with one experimental structure; PDB ID and brief meta-data about the protein taken from the PDB70 file are reported alongside the quantitative metrics. The second results file contains the raw results associated with each alignment reported in the score results file. Specifically, the translation and rotation arrays for each alignment are provided so that the structural alignment can be recreated. Additionally, residue-level scores are reported to quantify the closeness of the aligned residues between the predicted and experimental models.

59 BASIC BIOLOGICAL SCIENCES↗

Structural Models and Sequence Alignment Results of the Rhodospirillum rubrum Proteome

This dataset contains the structural models for the primary transcripts of the Rhodospirillum rubrum proteome as well as sequence alignment results for a subset of the encoded proteins. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the R. rubrum proteome to those available in the AlphaFold Protein Structure Database. For proteins that have been annotated as hypothetical, sequence alignment results from the HHblits and SAdLSA alignment methods are provided. These methods are often more capable to resolve sequence homology than other methods. Therefore, the results from both HHblits and SAdLSA are provided to identify possible homologs for these challenging proteins. Numerous sequence databases are utilized for these alignments. References AlphaFold v2 Multimer: https://doi.org/10.1101/2021.10.04.463034. References HHBlits: https://doi.org/10.1186/s12859-019-3019-7. References SAdLSA: https://doi.org/10.3389/fbinf.2021.689960.

59 BASIC BIOLOGICAL SCIENCES↗

Structural Models and Sequence Alignment Results of the Desulfovibrio vulgaris Proteome

This dataset contains the structural models for the primary transcripts of the Desulfovibrio vulgaris proteome as well as sequence alignment results for a subset of the encoded proteins. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the D. vulgaris proteome to those available in the AlphaFold Protein Structure Database (AFDB). This is a bit more complicated since the proteins reporting in the AFDB originate from an outdated form of the D. vulgaris sequence. The different versions of the D. vulgaris gene annotation are collected in the Chronology subdirectory; further consideration of these changes on the structural space of the proteome are currently underway. For proteins that have been annotated as hypothetical, sequence alignment results from the HHblits and SAdLSA alignment methods are provided. These methods are often more capable to resolve sequence homology than other methods. Therefore, the results from both HHblits and SAdLSA are provided to identify possible homologs for these challenging proteins. Numerous sequence databases are utilized for these alignments. References AlphaFold v2 Multimer: https://doi.org/10.1101/2021.10.04.463034. References HHblits: hhtps://doi.org/10.1186/s12859-019-3019-7. References SAdLSA: hhtps://doi.org/10.3389/fbinf.2021.689960.

59 BASIC BIOLOGICAL SCIENCES↗

Structural Models of the Rhodopseudomonas palustris Proteome

This dataset contains the structural models for the primary transcripts of the Rhodopseudomonas palustris proteome. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the R. palustris proteome to those available in the AlphaFold Protein Structure Database.

59 BASIC BIOLOGICAL SCIENCES↗

Native glycosylated HIV-1 Env in silico ensemble

The computationally modeled datasets generated in this study, namely the natively glycosylated SOSIP Env and the uniform mannose-9 glycosylated SOSIP Env ensembles will be submitted as Mendeley dataset server as required by the journal publishers, for public availability. These include two sets of Protein Data Bank (PDB)ensembles of 1000 structures each. The files are in the standardized PDB format. Each of the constituent ensemble structures are separated by the ‘END’ tag within the files. Snapshots from the beginning and end of each data set are given here. The column descriptions are as follows: Col1 –ATOM tag; Col 2 –atom number; COL 3 –atom name; COL4 –residue name; COL 5 –chain identifier; COL 6 –residue number; COL. 7,8,9 –X,Y,Z coordinates of atom in 3Dspace; COL 10 –Occupancy Factor; COL 11 -Temperature factor; COL 12 –segment identifier; COL 13 –element symbol.

59 BASIC BIOLOGICAL SCIENCES↗

SparcleQC: Automated Input File Creation for QM/MM Studies of Protein:Ligand Complexes

SparcleQC is a Python package that, given a protein:ligand complex in the Protein Data Bank (PDB) file format, can create quantum mechanics/molecular mechanics (QM/MM)-like input files for the electronic structure theory packages PSI4, QChem, and NWChem. The resulting input files include quantum mechanical representations of the ligand and a small section of the protein, surrounded by point charges that represent the rest of the protein. Creation of these QM/MM input files includes cutting and capping the QM subregion, obtaining point charges for the protein, and adjusting charges at the QM/MM boundary; and each of these tasks are automated by the software. In this article, we describe the details of SparcleQC’s procedure, show examples of the Python API, and explain additional features that are helpful in protein:ligand interaction studies. Finally, we show that SparcleQC enables automated preparation of input files for QM/MM calculations, which can return can return accurate interaction energies in minutes, while a fully quantum mechanical computation on the protein:ligand complex could take days, if it is even possible.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Algorithms and file structures to extend and enhance liquid chromatography and ion mobility mass spectrometry workflows (CRADA Final Report)

The purpose of this project was to continue supporting customizations of algorithms and raw data file structures to enhance software workflows for liquid chromatography (LC), mass spectrometry (MS) and ion mobility mass spectrometry (IM-MS)-based protein and metabolite characterization. PNNL worked with Agilent to design, implement, evaluate, and demonstrate new algorithms and integrated them as functionalities into the PNNL-PreProcessor software. The project augmented PNNL’s capabilities to analyze complex proteomics and metabolomics samples. These capabilities are directly beneficial to DOE and PNNL efforts to characterize and analyze these compounds in microbial and plant communities. The project assisted Agilent in further developing improved instrument-software solutions combining liquid chromatography and ion mobility with mass spectrometry for widespread applications in life sciences and other fields.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Validated ligand geometries for macromolecular refinement restraints and molecular-mechanics force fields

In macromolecular structure refinement, the low observation-to-parameter ratio and the lack of high-resolution data are countered by using a priori information in the form of restraints. Having accurate geometries of the chemical entities in the sample is paramount for generating accurate chemical restraints and, therefore, accurate macromolecular structures. In particular, it is desirable to have accurate restraints for known and novel ligand entities. Quantum mechanics (QM) can minimize the energy of a ligand by adjusting its geometry, and these geometries can be used to generate restraints for macromolecular refinement. This article describes a library of approximately 37 000 small molecules extracted from the Chemical Component Dictionary in the Protein Data Bank and minimized by density-functional QM. The library includes restraint files for use in crystallography or cryo-EM refinement, along with files suitable for molecular-dynamics simulation. Because the geometries are validated using the Cambridge Structural Database, the restraints library provides users with both functional restraints and minimized geometries. This work also provides procedures for generating new and accurate restraints.

Amber↗

SARS-CoV2 Protein-Ligand Simulation Dataset: Layer 1 (Simulation Initial Conditions and Parameters)

A set of 24 protein structures/complexes from the SARS CoV-2 proteome and inputs prepared for simulation using the CHARMM36m forcefield in PDB and gromacs formats. Each system contains a pdb and gromacs top and related input files necessary for running a temperature replica-exchange simulation. In addition, we also include: charmm PSF files (generated from the gromacs topology), a list of temperatures at which replica-exchange simulations were done (tempRamp), example gromacs run input mdp files, and initial minimized structures where available (minimized.pdb).

60 APPLIED LIFE SCIENCES↗

Unified Nanotechnology Format: One Way to Store Them All

The domains of DNA and RNA nanotechnology are steadily gaining in popularity while proving their value with various successful results, including biosensing robots and drug delivery cages. Nowadays, the nanotechnology design pipeline usually relies on computer-based design (CAD) approaches to design and simulate the desired structure before the wet lab assembly. To aid with these tasks, various software tools exist and are often used in conjunction. However, their interoperability is hindered by a lack of a common file format that is fully descriptive of the many design paradigms. Therefore, in this paper, we propose a Unified Nanotechnology Format (UNF) designed specifically for the biomimetic nanotechnology field. UNF allows storage of both design and simulation data in a single file, including free-form and lattice-based DNA structures. By defining a logical and versatile format, we hope it will become a widely accepted and used file format for the nucleic acid nanotechnology community, facilitating the future work of researchers and software developers. Together with the format description and publicly available documentation, we provide a set of converters from existing file formats to simplify the transition. Finally, we present several use cases visualizing example structures stored in UNF, showcasing the various types of data UNF can handle.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

phosaa14SB and phosaa19SB: Updated Amber Force Field Parameters for Phosphorylated Amino Acids

Phosphorylated amino acids are involved in many cell regulatory networks; proteins containing these post-translational modifications are widely studied both experimentally and computationally. Simulations are used to investigate a wide range of structural and dynamic properties of biomolecules, such as ligand binding, enzyme-reaction mechanisms, and protein folding. However, the development of force field parameters for the simulation of proteins containing phosphorylated amino acids using the Amber program has not kept pace with the development of parameters for standard amino acids, and it is challenging to model these modified amino acids with accuracy comparable to proteins containing only standard amino acids. In particular, the popular ff14SB and ff19SB models do not contain parameters for phosphorylated amino acids. Here, the dihedral parameters for the side chains of the most common phosphorylated amino acids are trained against reference data from QM calculations adopting the ff14SB approach, followed by validation against experimental data. Finally, library files and corresponding parameter files are provided, with versions that are compatible with both ff14SB and ff19SB.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

PPI DataHub Project Data Package: High-density Lipoprotein (HDL) Structure and Function Proteomics

The purpose of this experiment was to investigate how the interactions between APOA1 and APOA2 on the surface of high-density lipoproteins (HDL) impact particle function. Interactions were investigated on HDL isolated from human blood plasma using structural proteomics tools such as chemical cross-linking and limited proteolysis (LiP). The structural proteomics data was acquired using a Q-Exactive HF-X mass spectrometer and data was processed and compiled using MaxQuant sofware (v.1.6.17.0). Processed datasets are openly accessible from the download button (~2.8 GB) and contain secondary processed LiP and global proteomic results files and supporting metadata materials. Processed data downloads include a sample naming key, processed MaxQuant results/parameters, and protein annotated relative abundance files.

59 BASIC BIOLOGICAL SCIENCES↗

PPI DataHub Project Data Package: S. elongatus PCC 7942 Limited Proteolysis and Thermal Proteome Profiling Structural Proteomics (JM-PB-DP3)

The purpose of this experiment was to investigate structural alterations in proteins involved in central carbon metabolism and photosynthetic electron transfer pathways in Synechococcus elongatus PCC 7942. Sample data was obtained from S. elongatus cell lysates using three complementary mass spectrometry (MS) techniques using limited proteolysis (LiP-MS), thermal proteome profiling (TPP-MS), and redox enrichment (Redox-MS) in evaluating alterations solvent accessibility and structural stability caused by light perturbation at the molecular level. Experimentally processed sample data for LiP and TPP proteomic datasets were derived from the same cell culture stock, prepared simultaneously in parallel, and acquired by mass spectrometry. Processed datasets are openly accessible from the download button and contain secondary processed proteomic results files, computed outputs, and supporting metadata materials. Experimental samples processed for LiP-MS label-free quantification (LFQ) or TPP-MS tandem mass tag (TMT) 10-plex were acquired using a Q-Exactive HF-X mass spectrometer and processed/compiled using either MSGF+ (v2024.03.26) or ​​​​PlexedPiper for proteome evaluation. Additional software supporting downstream proteomic analysis include FragPipe (v.4.0), MSFragger (v.22.1), and an adapted Microbial Isolate LiP Analysis Workflow (located at Zenodo). Processed proteomic data downloads include a sample naming key, normalized quantification results files, and processed protein annotated abundance files.

59 BASIC BIOLOGICAL SCIENCES↗

pKPDB: a protein data bank extension database of p Ka and pI theoretical values

Abstract Summary pKa values of ionizable residues and isoelectric points of proteins provide valuable local and global insights about their structure and function. These properties can be estimated with reasonably good accuracy using Poisson–Boltzmann and Monte Carlo calculations at a considerable computational cost (from some minutes to several hours). pKPDB is a database of over 12 M theoretical pKa values calculated over 120k protein structures deposited in the Protein Data Bank. By providing precomputed pKa and pI values, users can retrieve results instantaneously for their protein(s) of interest while also saving countless hours and resources that would be spent on repeated calculations. Furthermore, there is an ever-growing imbalance between experimental pKa and pI values and the number of resolved structures. This database will complement the experimental and computational data already available and can also provide crucial information regarding buried residues that are under-represented in experimental measurements. Availability and implementation Gzipped csv files containing p Ka and isoelectric point values can be downloaded from https://pypka.org/pKPDB. To query a single PDB code please use the PypKa free server at https://pypka.org. The pKPDB source code can be found at https://github.com/mms-fcul/pKPDB. Supplementary information Supplementary data are available at Bioinformatics online.

Reis, Pedro B. P. S. (ORCID:0000000335636239)↗