Search NASA⌕ Search

SEARCH · Search NASA

Results for “Virtual drug screening”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

An artificial intelligence accelerated virtual screening platform for drug discovery

Abstract Structure-based virtual screening is a key tool in early drug discovery, with growing interest in the screening of multi-billion chemical compound libraries. However, the success of virtual screening crucially depends on the accuracy of the binding pose and binding affinity predicted by computational docking. Here we develop a highly accurate structure-based virtual screen method, RosettaVS, for predicting docking poses and binding affinities. Our approach outperforms other state-of-the-art methods on a wide range of benchmarks, partially due to our ability to model receptor flexibility. We incorporate this into a new open-source artificial intelligence accelerated virtual screening platform for drug discovery. Using this platform, we screen multi-billion compound libraries against two unrelated targets, a ubiquitin ligase target KLHDC2 and the human voltage-gated sodium channel Na V 1.7. For both targets, we discover hit compounds, including seven hits (14% hit rate) to KLHDC2 and four hits (44% hit rate) to Na V 1.7, all with single digit micromolar binding affinities. Screening in both cases is completed in less than seven days. Finally, a high resolution X-ray crystallographic structure validates the predicted docking pose for the KLHDC2 ligand complex, demonstrating the effectiveness of our method in lead discovery.

Science & Technology - Other Topics↗

Ensemble transfer learning for the prediction of anti-cancer drug response

Abstract Transfer learning, which transfers patterns learned on a source dataset to a related target dataset for constructing prediction models, has been shown effective in many applications. In this paper, we investigate whether transfer learning can be used to improve the performance of anti-cancer drug response prediction models. Previous transfer learning studies for drug response prediction focused on building models to predict the response of tumor cells to a specific drug treatment. We target the more challenging task of building general prediction models that can make predictions for both new tumor cells and new drugs. Uniquely, we investigate the power of transfer learning for three drug response prediction applications including drug repurposing, precision oncology, and new drug development, through different data partition schemes in cross-validation. We extend the classic transfer learning framework through ensemble and demonstrate its general utility with three representative prediction algorithms including a gradient boosting model and two deep neural networks. The ensemble transfer learning framework is tested on benchmark in vitro drug screening datasets. The results demonstrate that our framework broadly improves the prediction performance in all three drug response prediction applications with all three prediction algorithms.

60 APPLIED LIFE SCIENCES↗

Do Molecular Fingerprints Identify Diverse Active Drugs in Large-Scale Virtual Screening? (No)

Computational approaches for small-molecule drug discovery now regularly scale to the consideration of libraries containing billions of candidate small molecules. One promising approach to increased the speed of evaluating billion-molecule libraries is to develop succinct representations of each molecule that enable the rapid identification of molecules with similar properties. Molecular fingerprints are thought to provide a mechanism for producing such representations. Here, we explore the utility of commonly used fingerprints in the context of predicting similar molecular activity. We show that fingerprint similarity provides little discriminative power between active and inactive molecules for a target protein based on a known active—while they may sometimes provide some enrichment for active molecules in a drug screen, a screened data set will still be dominated by inactive molecules. We also demonstrate that high-similarity actives appear to share a scaffold with the query active, meaning that they could more easily be identified by structural enumeration. Furthermore, even when limited to only active molecules, fingerprint similarity values do not correlate with compound potency. In sum, these results highlight the need for a new wave of molecular representations that will improve the capacity to detect biologically active molecules based on their similarity to other such molecules.

59 BASIC BIOLOGICAL SCIENCES↗

Converting tabular data into images for deep learning with convolutional neural networks

Abstract Convolutional neural networks (CNNs) have been successfully used in many applications where important information about data is embedded in the order of features, such as speech and imaging. However, most tabular data do not assume a spatial relationship between features, and thus are unsuitable for modeling using CNNs. To meet this challenge, we develop a novel algorithm, image generator for tabular data (IGTD), to transform tabular data into images by assigning features to pixel positions so that similar features are close to each other in the image. The algorithm searches for an optimized assignment by minimizing the difference between the ranking of distances between features and the ranking of distances between their assigned pixels in the image. We apply IGTD to transform gene expression profiles of cancer cell lines (CCLs) and molecular descriptors of drugs into their respective image representations. Compared with existing transformation methods, IGTD generates compact image representations with better preservation of feature neighborhood structure. Evaluated on benchmark drug screening datasets, CNNs trained on IGTD image representations of CCLs and drugs exhibit a better performance of predicting anti-cancer drug response than both CNNs trained on alternative image representations and prediction models trained on the original tabular data.

59 BASIC BIOLOGICAL SCIENCES↗

Supercomputing Pipelines Search for Therapeutics Against COVID-19

The urgent search for drugs to combat SARS-CoV-2 has included the use of supercomputers. The use of general-purpose graphical processing units (GPUs), massive parallelism, and new software for high-performance computing (HPC) has allowed researchers to search the vast chemical space of potential drugs faster than ever before. We developed a new drug discovery pipeline using the Summit supercomputer at Oak Ridge National Laboratory to help pioneer this effort, with new platforms that incorporate GPU-accelerated simulation and allow for the virtual screening of billions of potential drug compounds in days compared to weeks or months for their ability to inhibit SARS-COV-2 proteins. Here, this effort will accelerate the process of developing drugs to combat the current COVID-19 pandemic and other diseases.

60 APPLIED LIFE SCIENCES↗

High-throughput virtual laboratory for drug discovery using massive datasets

Time-to-solution for structure-based screening of massive chemical databases for COVID-19 drug discovery has been decreased by an order of magnitude, and a virtual laboratory has been deployed at scale on up to 27,612 GPUs on the Summit supercomputer, allowing an average molecular docking of 19,028 compounds per second. Over one billion compounds were docked to two SARS-CoV-2 protein structures with full optimization of ligand position and 20 poses per docking, each in under 24 hours. GPU acceleration and high-throughput optimizations of the docking program produced 350× mean speedup over the CPU version (50× speedup per node). GPU acceleration of both feature calculation for machine-learning based scoring and distributed database queries reduced processing of the 2.4 TB output by orders of magnitude. The resulting 50× speedup for the full pipeline reduces an initial 43 day runtime to 21 hours per protein for providing high-scoring compounds to experimental collaborators for validation assays.

97 MATHEMATICS AND COMPUTING↗

Discovery of Novel Rhizoctonia solani DHFR Inhibitors as Fungicides Using Virtual Screening

Dihydrofolate reductase (DHFR) is an essential enzyme in the folate pathway and has been recognized as a well-known target for antibacterial and antifungal drugs. We discovered eight compounds from the ZINC database using virtual screening to inhibit Rhizoctonia solani (R. solani), a fungal pathogen in crops. These compounds were evaluated with in vitro assays for enzymatic and antifungal activity. Among these, compound Hit8 is the most active R. solani DHFR inhibitor, with the IC 50 of 10.2 μM. The selectivity of inhibition is 22.3 against human DHFR with the IC 50 of 227.7 μM. Moreover, Hit8 has higher antifungal activity against R. solani (EC 50 of 38.2 mg L –1 ) compared with validamycin A (EC 50 of 67.6 mg L –1 ), a well-documented fungicide. These results suggest that Hit8 may be a potential fungicide. Finally, our study exemplifies a computer-aided method to discover novel inhibitors that could target plant pathogenic fungi.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Drugsniffer: An Open Source Workflow for Virtually Screening Billions of Molecules for Binding Affinity to Protein Targets

The SARS-CoV2 pandemic has highlighted the importance of efficient and effective methods for identification of therapeutic drugs, and in particular has laid bare the need for methods that allow exploration of the full diversity of synthesizable small molecules. While classical high-throughput screening methods may consider up to millions of molecules, virtual screening methods hold the promise of enabling appraisal of billions of candidate molecules, thus expanding the search space while concurrently reducing costs and speeding discovery. Here, we describe a new screening pipeline, called drugsniffer, that is capable of rapidly exploring drug candidates from a library of billions of molecules, and is designed to support distributed computation on cluster and cloud resources. As an example of performance, our pipeline required ~40,000 total compute hours to screen for potential drugs targeting three SARS-CoV2 proteins among a library of ~3.7 billion candidate molecules.

59 BASIC BIOLOGICAL SCIENCES↗

Iterative computational design and crystallographic screening identifies potent inhibitors targeting the Nsp3 macrodomain of SARS-CoV-2

The nonstructural protein 3 (NSP3) of the severe acute respiratory syndrome-coronavirus-2 (SARS-CoV-2) contains a conserved macrodomain enzyme (Mac1) that is critical for pathogenesis and lethality. While small-molecule inhibitors of Mac1 have great therapeutic potential, at the outset of the COVID-19 pandemic, there were no well-validated inhibitors for this protein nor, indeed, the macrodomain enzyme family, making this target a pharmacological orphan. Here, we report the structure-based discovery and development of several different chemical scaffolds exhibiting low- to sub-micromolar affinity for Mac1 through iterations of computer-aided design, structural characterization by ultra-high-resolution protein crystallography, and binding evaluation. Potent scaffolds were designed with in silico fragment linkage and by ultra-large library docking of over 450 million molecules. Both techniques leverage the computational exploration of tangible chemical space and are applicable to other pharmacological orphans. Overall, 160 ligands in 119 different scaffolds were discovered, and 153 Mac1-ligand complex crystal structures were determined, typically to 1 Å resolution or better. Our analyses discovered selective and cell-permeable molecules, unexpected ligand-mediated conformational changes within the active site, and key inhibitor motifs that will template future drug development against Mac1.

60 APPLIED LIFE SCIENCES↗

Predicting receptor-ligand pairing preferences in plant-microbe interfaces via molecular dynamics and machine learning

Microbiome assembly, structure, and dynamics significantly influence plant health. Secreted microbial signaling molecules initiate and mediate symbiosis by binding to structurally compatible plant receptors. For example, lipo-chitooligosaccharides (LCOs), produced by nitrogen-fixing rhizobial bacteria and various fungi, are recognized by plant lysin motif receptor-like kinases (LysM-RLKs), which activate the common symbiotic pathway. Accurately predicting these molecular interactions could reveal complementary signatures underlying the initial stages of endosymbiosis. Despite the breakthrough in protein-ligand structure prediction with deep learning-based tools, such as AlphaFold3, the large size and highly flexible nature of signaling compounds like LCOs present major challenges for detailed structural characterization and binding-affinity prediction. Typical structure-/physics-based methods of ligand virtual screening are designed for small, drug-like molecules, often rely on high-resolution, experimentally determined structures of the protein receptors, and rarely achieve sufficient sampling to obtain converged thermodynamic quantities with large ligands. In this study, we developed a hybrid molecular dynamics/machine learning (MD/ML) approach capable of predicting binding affinity rankings with high accuracy in systems involving large, flexible ligands, despite limited experimental structural information. Using coarse initial structural models, the predictions using the MD/ML workflow achieved strong alignment with experimental trends, particularly in the top-affinity tier for four legume LysM-RLKs (LYR3) binding to LCOs and a chitooligosaccharide. Furthermore, the MD-based conformation selection protocol provided critical structural insights into substrate specificity and binding mechanisms. This study demonstrates a powerful method to screen for challenging cognate ligand-receptors and advance our understanding of the molecular basis of microbial colonization in plants.

Lipo-chitooligosaccharides↗

Identification of inhibitors against SARS-CoV-2 variants of concern using virtual screening and metadynamics-based enhanced sampling

Among the variants of SARS-CoV-2, some are more infectious than the Wild-type. Interestingly, these mutations enable the virus to evade the therapeutic efforts. Hence, there is a need for candidate drug molecules that can potently bind with all the variants. Here we have adopted a strategy combining virtual screening, molecular docking followed by rigorous sampling by metadynamics simulations to find candidate molecules. From our results we found four highly potent drug candidates that can bind to the Spike-RBD of all the variants of the virus. Additionally, we also found that certain signature residues on the RBM region commonly bind to each of these inhibitors. Thus, our study not only gives information on the chemical compounds, but also residues on the proteins which could be targeted for future drug and vaccine development studies.

60 APPLIED LIFE SCIENCES↗

Structure based virtual screening identifies small molecule effectors for the sialoglycan binding protein Hsa

Infective endocarditis (IE) is a cardiovascular disease often caused by bacteria of the viridans group of streptococci, which includes Streptococcus gordonii and Streptococcus sanguinis. Previous research has found that serine-rich repeat (SRR) proteins on the S. gordonii bacterial surface play a critical role in pathogenesis by facilitating bacterial attachment to sialylated glycans displayed on human platelets. Despite their important role in disease progression, there are currently no anti-adhesive drugs available on the market. Here, we performed structure-based virtual screening using an ensemble docking approach followed by consensus scoring to identify novel small molecule effectors against the sialoglycan binding domain of the SRR adhesin protein Hsa from the S. gordonii strain DL1. The screening successfully predicted nine compounds which were able to displace the native ligand (sialyl-T antigen) in an in vitro assay and bind competitively to Hsa. Furthermore, hierarchical clustering based on the MACCS fingerprints showed that eight of these small molecules do not share a common scaffold with the native ligand. This study indicates that SRR family of adhesin proteins can be inhibited by diverse small molecules and thus prevent the interaction of the protein with the sialoglycans. This opens new avenues for discovering potential drugs against IE.

60 APPLIED LIFE SCIENCES↗

Crystal structure of the human PRPK–TPRKB complex

Mutations of the p53-related protein kinase (PRPK) and TP53RK-binding protein (TPRKB) cause Galloway-Mowat syndrome (GAMOS) and are found in various human cancers. We have previously shown that small compounds targeting PRPK showed anti-cancer activity against colon and skin cancer. Here we present the 2.53 Å crystal structure of the human PRPK-TPRKB-AMPPNP (adenylyl-imidodiphosphate) complex. The structure reveals details in PRPK-AMPPNP coordination and PRPK-TPRKB interaction. PRPK appears in an active conformation, albeit lacking the conventional kinase activation loop. We constructed a structural model of the human EKC/KEOPS complex, composed of PRPK, TPRKB, OSGEP, LAGE3, and GON7. Disease mutations in PRPK and TPRKB are mapped into the structure, and we show that one mutation, PRPK K238Nfs*2, lost the binding to OSGEP. Our structure also makes the virtual screening possible and paves the way for more rational drug design.

59 BASIC BIOLOGICAL SCIENCES↗

A goldilocks computational protocol for inhibitor discovery targeting DNA damage responses including replication-repair functions

While many researchers can design knockdown and knockout methodologies to remove a gene product, this is mainly untrue for new chemical inhibitor designs that empower multifunctional DNA Damage Response (DDR) networks. Here, we present a robust Goldilocks (GL) computational discovery protocol to efficiently innovate inhibitor tools and preclinical drug candidates for cellular and structural biologists without requiring extensive virtual screen (VS) and chemical synthesis expertise. By computationally targeting DDR replication and repair proteins, we exemplify the identification of DDR target sites and compounds to probe cancer biology. Our GL pipeline integrates experimental and predicted structures to efficiently discover leads, allowing early-structure and early-testing (ESET) experiments by many laboratories. By employing an efficient VS protocol to examine protein-protein interfaces (PPIs) and allosteric interactions, we identify ligand binding sites beyond active sites, leveraging in silico advances for molecular docking and modeling to screen PPIs and multiple targets. A diverse 3,174 compound ESET library combines Diamond Light Source DSI-poised, Protein Data Bank fragments, and FDA-approved drugs to span relevant chemotypes and facilitate downstream hit evaluation efficiency for academic laboratories. Two VS per library and multiple ranked ligand binding poses enable target testing for several DDR targets. This GL library and protocol can thus strategically probe multiple DDR network targets and identify readily available compounds for early structural and activity testing to overcome bottlenecks that can limit timely breakthrough drug discoveries. By testing accessible compounds to dissect multi-functional DDRs and suggesting inhibitor mechanisms from initial docking, the GL approach may enable more groups to help accelerate discovery, suggest new sites and compounds for challenging targets including emerging biothreats and advance cancer biology for future precision medicine clinical trials.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative Assessment of Pose Prediction Accuracy in RNA–Ligand Docking

Structure-based virtual high-throughput screening is used in early-stage drug discovery. Over the years, docking protocols and scoring functions for protein–ligand complexes have evolved to improve the accuracy in the computation of binding strengths and poses. In the past decade, RNA has also emerged as a target class for new small-molecule drugs. However, most ligand docking programs have been validated and tested for proteins and not RNA. Here, we test the docking power (pose prediction accuracy) of three state-of-the-art docking protocols on 173 RNA–small molecule crystal structures. The programs are AutoDock4 (AD4) and AutoDock Vina (Vina), which were designed for protein targets, and rDock, which was designed for both protein and nucleic acid targets. AD4 performed relatively poorly. For RNA targets for which a crystal structure of a bound ligand used to limit the docking search space is available and for which the goal is to identify new molecules for the same pocket, rDock performs slightly better than Vina, with success rates of 48% and 63%, respectively. However, in the more common type of early-stage drug discovery setting, in which no structure of a ligand–target complex is known and for which a larger search space is defined, rDock performed similarly to Vina, with a low success rate of ~27%. Further, Vina was found to have bias for ligands with certain physicochemical properties, whereas rDock performs similarly for all ligand properties. Thus, for projects where no ligand–protein structure already exists, Vina and rDock are both applicable. However, the relatively poor performance of all methods relative to protein–target docking illustrates a need for further methods refinement.

59 BASIC BIOLOGICAL SCIENCES↗

Performance Portability of Molecular Docking Miniapp On Leadership Computing Platforms

Rapidly changing computer architectures, such as those found at high-performance computing (HPC) facilities, present the need for mini-applications (miniapps) that capture essential algorithms used in large applications to test program performance and portability, aiding transitions to new systems. The COVID-19 pandemic has fueled a flurry of activity in computational drug discovery, including the use of supercomputers and GPU acceleration for massive virtual screens for therapeutics. Recent work targeting COVID-19 at the Oak Ridge Leadership Computing Facility (OLCF) used the GPU-accelerated program AutoDock-GPU to screen billions of compounds on the Summit supercomputer. In this paper we present the development of a new miniapp, miniAutoDock-GPU, that can be used to evaluate the performance and portability of GPU-accelerated protein-ligand docking programs on different computer architectures. These tests are especially relevant as facilities transition from petascale systems and prepare for upcoming exascale systems that will use a variety of GPU vendors. The key calculations, namely, the Lamarckian genetic algorithm combined with a local search using a Solis-Wets based random optimization algorithm, are implemented. We developed versions of the miniapp using several different programming models for GPU acceleration, including a version using the CUDA runtime API for NVIDIA GPUs, and the Kokkos middle-ware API which is facilitated by C++ template libraries. A third version, currently in progress, uses the HIP programming model. These efforts will help facilitate the transition to exascale systems for this important emerging HPC application, as well as its use on a wide range of heterogeneous platforms.

Thavappiragasam, Mathialakan↗

Multivariate chemogenomic screening prioritizes new macrofilaricidal leads

Development of direct acting macrofilaricides for the treatment of human filariases is hampered by limitations in screening throughput imposed by the parasite life cycle. In vitro adult screens typically assess single phenotypes without prior enrichment for chemicals with antifilarial potential. We developed a multivariate screen that identified dozens of compounds with submicromolar macrofilaricidal activity, achieving a hit rate of >50% by leveraging abundantly accessible microfilariae. Adult assays were multiplexed to thoroughly characterize compound activity across relevant parasite fitness traits, including neuromuscular control, fecundity, metabolism, and viability. Seventeen compounds from a diverse chemogenomic library elicited strong effects on at least one adult trait, with differential potency against microfilariae and adults. Our screen identified five compounds with high potency against adults but low potency or slow-acting microfilaricidal effects, at least one of which acts through a novel mechanism. We show that the use of microfilariae in a primary screen outperforms model nematode developmental assays and virtual screening of protein structures inferred with deep learning. These data provide new leads for drug development, and the high-content and multiplex assays set a new foundation for antifilarial discovery.

59 BASIC BIOLOGICAL SCIENCES↗