It's a Scheduling Affair: GROMACS in the Cloud with the KubeFlux Scheduler
Explore the source record for details and available documents.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
A set of 24 protein structures/complexes from the SARS CoV-2 proteome and inputs prepared for simulation using the CHARMM36m forcefield in PDB and gromacs formats. Each system contains a pdb and gromacs top and related input files necessary for running a temperature replica-exchange simulation. In addition, we also include: charmm PSF files (generated from the gromacs topology), a list of temperatures at which replica-exchange simulations were done (tempRamp), example gromacs run input mdp files, and initial minimized structures where available (minimized.pdb).
Raw temperature replica-exchange molecular dyanmics trajectories of 23 different SARS CoV-2 Systems including S Protein ACE2-receptor binding-domain, MPro, PLPro, NSP3 ADRP (X-domain/macrodomain/phosphatase), NSP15 (endoribonuclease), NSP9, NSP10, NSP16, and N-protein N-terminus. Systems were prepared with charmm-gui and simulated using the GROMACS simulation software suite. Non-demuxed, i.e. discontinuous/constant temperature window, trajectories are provided in compressed GROMACS xtc and full trr formats. This data supplements the data releases DOI: 10.13139/OLCF/160650 'SARS-CoV2 Protein-Ligand Simulation Dataset: Layer 1 (Simulation Initial Conditions and Parameters)' and DOI: 10.13139/OLCF/1657844 'SARS-CoV2 Protein-Ligand Simulation Dataset: Layer 2 (Extracted Protein Coordinate Trajectories)'
The ezAlign is aimed at clustering coarse grain simulation to find common or uncommon occurrences and convert coarse (CG) grained coordinate and topology files to atomistic formats. We use a PointNet based approach to map individual frames of simulation to points in a latent space. These points are then clustered using a variety of clustering methods. Clusters are analyzed to associate them with states in the simulation. Frames can be chosen from these clusters based on proximity to cluster centers. ezAlign takes CG coordinate and topology files and converts and outputs their corresponding atomistic formats using an alignment and relaxation procedure. ezAlign is designed to convert complex, solvated biological systems including lipid membranes with drug-like molecules using GROMACS. A GROMACS checkpoint (.cpt) file is also outputted to enable continuation simulations that retain the equilibrated atomic velocities. Independent atomistic coordinates and topologies for every molecule must already be included in ezAlign/files. A number of commonly simulated biological molecules are currently provided.
Protein-only trajectories of the lowest-temperature window (310K) extracted from temperature replica-exchange molecular dynamics simulations of 23 different SARS CoV-2 systems, including S Protein ACE2-receptor binding-domain, MPro, PLPro, NSP3 ADRP (X-domain/macrodomain/phosphatase), NSP15 (endoribonuclease), NSP9, NSP10, NSP16, and N-protein N-terminus. Systems were prepared with charmm-gui and simulated using the GROMACS simulation software suite. Trajectories are provided in compressed dcd format with accompanying coordinate/topology files in pdb and psf formats. This data supplements the data release, DOI: 10.13139/OLCF/1650650 ('SARS-CoV2 Protein-Ligand Simulation Dataset: Layer 1 (Simulation Initial Conditions and Parameters)')
Anhydrous Hydrogen Fluoride (HF) at high temperatures and pressures is used to process and manufacture nuclear fuel. As HF is often used directly with uranium, correct neutron thermal scattering cross sections are crucial to criticality safety applications. Classical molecular dynamics (CMD) simulation of the flexible HF system was used to create the thermal scattering law (TSL) and cross sections. The initial 2-site model is used in LAMMPS, and it can not capture the H-bond. To correctly represent the H-bond effects, a second, 3-site model was constructed in GROMACS. The 3-site model handled H-bonds by connecting a massless charge to the molecule. Key model parameters were compared to experimental data to verify the approach and models. To get the normalized VACF, the model was compared using hydrogen and fluorine bond length, density, potential energy, and diffusion coefficient. The phonon DOSs for both models were derived from the normalized VACF. DOSs were used to estimate the TSL ( S ( α, β )) and neutron thermal scattering cross sections for hydrogen in HF. The TSLs were evaluated using the FLASSH code with the Schofield diffusion model. It was observed that the representation of the hydrogen bonding changes the TSL's diffusional contributions. This is represented in the low energy scattering cross section, where intermolecular binding effects shift the cross section.
Explicit treatment of electronic polarizability in empirical force fields (FFs) represents an extension over a traditional additive or pairwise FF and provides a more realistic model of the variations in electronic structure in condensed phase, macromolecular simulations. To facilitate utilization of the polarizable FF based on the classical Drude oscillator model, Drude Prepper has been developed in CHARMM-GUI. Drude Prepper ingests additive CHARMM protein structures file (PSF) and pre-equilibrated coordinates in CHARMM, PDB, or NAMD format, from which the molecular components of the system are identified. These include all residues and patches connecting those residues along with water, ions, and other solute molecules. This information is then used to construct the Drude FF-based PSF using molecular generation capabilities in CHARMM, followed by minimization and equilibration. In addition, inputs are generated for molecular dynamics (MD) simulations using CHARMM, GROMACS, NAMD, and OpenMM. Validation of the Drude Prepper protocol and inputs is performed through conversion and MD simulations of various heterogeneous systems that include proteins, nucleic acids, lipids, polysaccharides, and atomic ions using the aforementioned simulation packages. Stable simulations are obtained in all studied systems, including 5 μs simulation of ubiquitin, verifying the integrity of the generated Drude PSFs. Additionally, the ability of the Drude FF to model variations in electronic structure is shown through dipole moment analysis in selected systems. Finally, the capabilities and availability of Drude Prepper in CHARMM-GUI is anticipated to greatly facilitate the application of the Drude FF to a range of condensed phase, macromolecular systems.
Performance evaluation is crucial to understanding the behavior of scientific workflows. In this study, we target an emerging type of workflow, called in situ workflows. These workflows tightly couple components such as simulation and analysis to improve overall workflow performance. To understand the tradeoffs of various configurable parameters for coupling these heterogeneous tasks, namely simulation stride, and component placement, separately monitoring each component is insufficient to gain insights into the entire workflow behavior. Through an analysis of the state-of-the-art research, we propose a lightweight metric, derived from a defined in situ step, for assessing resource usage efficiency of an in situ workflow execution. By applying this metric to a synthetic workflow, which is parameterized to emulate behaviors of a molecular dynamics simulation, we explore two possible scenarios (Idle Simulation and Idle Analyzer) for the characterization of in situ workflow execution. In addition to preliminary results from a recently published study [11], we further exploit the proposed metric to evaluate a practical in situ workflow with a real molecular dynamics application, i.e., GROMACS. Here, experimental results show that the in transit placement (analytics on dedicated nodes) sustains a higher frequency for performing in situ analysis compared to the helper-core configuration (analytics co-allocated with simulation).
An implementation of the replica exchange with dynamical scaling (REDS) method in the commonly used molecular dynamics program GROMACS is presented. REDS is a replica exchange method that requires fewer replicas than conventional replica exchange while still providing data over a range of temperatures and can be used in either constant volume or constant pressure ensembles. Details for running REDS simulations are given, and an application to the human islet amyloid polypeptide (hIAPP) 11-25 fragment shows that the model efficiently samples conformational space.
Here, we introduce Nanopolysaccharide Builder (NPB), a user-friendly software tool designed to construct polysaccharide nanostructures─mainly those based on cellulose, chitin, and chitosan─using experimental data or user-defined parameters. NPB enables the generation of cellulose and chitin allomorphs with customizable biochemical topologies and also facilitates the construction of large bundles that replicate nanostructures found in biological support systems, including plant cell walls and arthropod cuticles. The software outputs atomic Cartesian coordinates in Protein Data Bank (PDB) format and also provides atom connectivity files in PSF and PARM formats, ensuring seamless integration with major molecular dynamics (MD) engines such as NAMD, CHARMM, GROMACS, AMBER, OpenMM, and LAMMPS. Built on an interactive visualization framework, NPB features a graphical user interface (GUI) and supports both macOS and Linux operating systems. By enabling detailed atomic-scale studies of polysaccharide evolution in extracellular matrices and cell walls of algae, bacteria, fungi, and plants, NPB is poised to advance AI-guided research in sustainable chemical development and biomass utilization.
Thermodynamic properties of liquid mixtures govern processes that range from drug delivery to energy storage, yet extracting these properties from molecular simulations remains challenging. Kirkwood–Buff (KB) theory offers a rigorous route by linking microscopic pair distribution functions to macroscopic free energies, but practical use of the theory has been hindered by two obstacles: (i) the long simulations needed to obtain well-converged Kirkwood-Buff integrals (KBIs) and (ii) the specialized corrections required to translate finite-size data to the thermodynamic limit. $\texttt{KBKit}$ is an open-source Python package that removes these barriers. It automatically computes KBIs and derived thermodynamic quantities from GROMACS input files, applies state-of-the-art finite-size corrections, and provides built-in diagnostic tools to quantify statistical uncertainty. Written with modern software-engineering practices—continuous integration, extensive unit testing, and thorough documentation—$\texttt{KBKit}$ is both reliable and easy to extend. By condensing complex KBI analysis into a few intuitive commands, $\texttt{KBKit}$ enables researchers to incorporate KB theory into routine simulation workflows and accelerate the discovery of solution-phase thermodynamics.
The importance of ensemble computing is well established. However, executing ensembles at scale introduces interesting performance fluctuations that have not been well investigated. In this paper, we trace our experience uncovering performance fluctuations of ensemble applications (primarily constituting a workflow of GROMACS tasks), and unsuccessful attempts, so far, at trying to discern the underlying cause(s) of performance fluctuations. Is the failure to discern the causative or contributing factors a failure of capability? Or imagination? Do the fluctuations have their genesis in some inscrutable aspect of the system or software? Does it warrant a fundamental reassessment and rethinking of how we assume and conceptualize performance reproducibility? Answers to these questions are not straightforward, nor are they immediate or obvious. We conclude with a discussion about the performance of ensemble applications and ruminate over the implications for how we define and measure application performance.
PTM-Psi is a Python 3 package that combines several capabilities to streamline the workflow to interrogate the impact of PTMs on proteins using well-established software packages. The workflow of the PTM-Psi software package includes input files and launch instances from standard packages such as AlphaFold, NWChem, GROMACS, and the Autodock Suite
The Multiscale Machine-Learned Modeling Infrastructure (MuMMI) is a multiscale workflow management infrastructure that can concurrently orchestrate thousands of molecular dynamics (MD) simulations operating at different time and/or length scales, spanning nanoseconds to seconds and nanometers to micrometers. MuMMI uses machine learning (backed by biology experiments) to guide a massive ensemble of MD simulations that capture biologically relevant time and length scales with unprecedented resolution. MuMMI supports multiple MD codes such as GROMACS and ddcMD and can be fully deployed using the HPC package manager Spack. MuMMI has been used in many publications to run hundreds of thousands simulations, leading to significant biology breakthroughs.
Colicins are antimicrobial proteins produced by bacteria for the purpose of destroying neighboring bacteria. Colicin activity is neutralized by a specific cognate immunity protein in order to protect the host. This study investigates the structural and binding mechanisms underlying the interaction of colicin-D, -E3 and -E8 to their respective immunity proteins (ImD, Im3 and Im8) using structure prediction, molecular dynamics (MD) simulations and MM-PBSA approach of free energy calculations. High-confidence colicin-immunity (Col-Im) complex structures predicted using AlphaFold2 were subjected to MD simulations of 150 ns with GROMACS and were analyzed for the binding free energy calculation using gmx_MMPBSA. Results showed that the complex of Col_E3-Im3 exhibited the most favorable binding free energy, driven by strong van der Waals and electrostatic interactions. Col_D-ImD and Col_E8-Im8 also showed the favorable binding. Electrostatics and hydrogen bonding emerged as a key factor driving binding and stability, while polar solvation acted as a destabilizing factor across all systems. These outcomes provide an understanding of the molecular mechanisms of Col-Im systems, with potential applications for developing natural antimicrobials for food safety.
Contains the initial configurations and output results for the simulations reported in the paper Chemical and Morphological Structure of Transgenic Switchgrass Organosolv Lignin Extracted by Ethanol, Tetrahydrofuran, and γ-Valerolactone Pretreatments