Search NASA⌕ Search

SEARCH · Search NASA

Results for “Drug Discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Fingerprinting Interactions between Proteins and Ligands for Facilitating Machine Learning in Drug Discovery

Molecular recognition is fundamental in biology, underpinning intricate processes through specific protein–ligand interactions. This understanding is pivotal in drug discovery, yet traditional experimental methods face limitations in exploring the vast chemical space. Computational approaches, notably quantitative structure–activity/property relationship analysis, have gained prominence. Molecular fingerprints encode molecular structures and serve as property profiles, which are essential in drug discovery. While two-dimensional (2D) fingerprints are commonly used, three-dimensional (3D) structural interaction fingerprints offer enhanced structural features specific to target proteins. Machine learning models trained on interaction fingerprints enable precise binding prediction. Recent focus has shifted to structure-based predictive modeling, with machine-learning scoring functions excelling due to feature engineering guided by key interactions. Notably, 3D interaction fingerprints are gaining ground due to their robustness. Various structural interaction fingerprints have been developed and used in drug discovery, each with unique capabilities. This review recapitulates the developed structural interaction fingerprints and provides two case studies to illustrate the power of interaction fingerprint-driven machine learning. The first elucidates structure–activity relationships in β2 adrenoceptor ligands, demonstrating the ability to differentiate agonists and antagonists. The second employs a retrosynthesis-based pre-trained molecular representation to predict protein–ligand dissociation rates, offering insights into binding kinetics. Despite remarkable progress, challenges persist in interpreting complex machine learning models built on 3D fingerprints, emphasizing the need for strategies to make predictions interpretable. Binding site plasticity and induced fit effects pose additional complexities. Interaction fingerprints are promising but require continued research to harness their full potential.

3D structural interaction fingerprints↗

Deep generative molecular design reshapes drug discovery

Recent advances and accomplishments of artificial intelligence (AI) and deep generative models have established their usefulness in medicinal applications, especially in drug discovery and development. To correctly apply AI, the developer and user face questions such as which protocols to consider, which factors to scrutinize, and how the deep generative models can integrate the relevant disciplines. This review summarizes classical and newly developed AI approaches, providing an updated and accessible guide to the broad computational drug discovery and development community. We introduce deep generative models from different standpoints and describe the theoretical frameworks for representing chemical and biological structures and their applications. We discuss the data and technical challenges and highlight future directions of multimodal deep generative models for accelerating drug discovery.

59 BASIC BIOLOGICAL SCIENCES↗

An artificial intelligence accelerated virtual screening platform for drug discovery

Abstract Structure-based virtual screening is a key tool in early drug discovery, with growing interest in the screening of multi-billion chemical compound libraries. However, the success of virtual screening crucially depends on the accuracy of the binding pose and binding affinity predicted by computational docking. Here we develop a highly accurate structure-based virtual screen method, RosettaVS, for predicting docking poses and binding affinities. Our approach outperforms other state-of-the-art methods on a wide range of benchmarks, partially due to our ability to model receptor flexibility. We incorporate this into a new open-source artificial intelligence accelerated virtual screening platform for drug discovery. Using this platform, we screen multi-billion compound libraries against two unrelated targets, a ubiquitin ligase target KLHDC2 and the human voltage-gated sodium channel Na V 1.7. For both targets, we discover hit compounds, including seven hits (14% hit rate) to KLHDC2 and four hits (44% hit rate) to Na V 1.7, all with single digit micromolar binding affinities. Screening in both cases is completed in less than seven days. Finally, a high resolution X-ray crystallographic structure validates the predicted docking pose for the KLHDC2 ligand complex, demonstrating the effectiveness of our method in lead discovery.

Science & Technology - Other Topics↗

A New Drug Discovery Platform: Application to DNA Polymerase Eta and Apurinic/Apyrimidinic Endonuclease 1

The ability to quickly discover reliable hits from screening and rapidly convert them into lead compounds, which can be verified in functional assays, is central to drug discovery. The expedited validation of novel targets and the identification of modulators to advance to preclinical studies can significantly increase drug development success. Our SaXPyTM (“SAR by X-ray Poses Quickly”) platform, which is applicable to any X-ray crystallography-enabled drug target, couples the established methods of protein X-ray crystallography and fragment-based drug discovery (FBDD) with advanced computational and medicinal chemistry to deliver small molecule modulators or targeted protein degradation ligands in a short timeframe. Our approach, especially for elusive or “undruggable” targets, allows for (i) hit generation; (ii) the mapping of protein–ligand interactions; (iii) the assessment of target ligandability; (iv) the discovery of novel and potential allosteric binding sites; and (v) hit-to-lead execution. These advances inform chemical tractability and downstream biology and generate novel intellectual property. We describe here the application of SaXPy in the discovery and development of DNA damage response inhibitors against DNA polymerase eta (Pol η or POLH) and apurinic/apyrimidinic endonuclease 1 (APE1 or APEX1). Notably, our SaXPy platform allowed us to solve the first crystal structures of these proteins bound to small molecules and to discover novel binding sites for each target.

59 BASIC BIOLOGICAL SCIENCES↗

Annual Report for Structure-Aware Unsupervised, Transformational Machine Learning for Drug Discovery

The major goal of this project is to develop machine learning (ML) methods to enable improved predictive power on real drug discovery for novel targets. More specifically, we plan to demonstrate the capability and effectiveness of ML tools utilizing unlabeled large-volume protein-ligand datasets. We also plan to demonstrate the capability and effectiveness of the developed methods by testing on a realistic drug discovery task to identify pan-coronavirus protease inhibitors such as SARS-CoV-2. While the overall goals and milestones remain consistent with the original proposal, certain technical details have been modified, which we will describe in this report.

97 MATHEMATICS AND COMPUTING↗

Structure-Aware Unsupervised, Transformational Machine Learning for Drug Discovery (DTRA Basic Research Final Report)

The major goal of this project is to develop machine learning (ML) methods to enable improved predictive power on real drug discovery for novel targets. More specifically, we planned to demonstrate the capability and effectiveness of ML tools utilizing unlabeled large-volume protein-ligand datasets. We investigated multiple pre-training approaches for 3D protein-ligand structure-based foundation models, without relying on experimental binding data. We also addressed scenarios in which crystal structures are unavailable or binding data are limited. We also planned to develop a complete pipeline to screen novel compounds as well as to demonstrate the capability and effectiveness of the developed methods by testing on a realistic drug discovery task such as SARS-CoV-2. While the major goals and milestones remain consistent with the original proposal, certain technical details have been adjusted, based on the experimental results and related outcomes.

97 MATHEMATICS AND COMPUTING↗

Drug discovery efforts at George Mason University

With over 39,000 students, and research expenditures in excess of $200 million, George Mason University (GMU) is the largest R1 (Carnegie Classification of very high research activity) university in Virginia. Mason scientists have been involved in the discovery and development of novel diagnostics and therapeutics in areas as diverse as infectious diseases and cancer. Below are highlights of the efforts being led by Mason researchers in the drug discovery arena. To enable targeted cellular delivery, and non-biomedical applications, Veneziano and colleagues have developed a synthesis strategy that enables the design of self-assembling DNA nanoparticles (DNA origami) with prescribed shape and size in the 10 to 100 nm range. The nanoparticles can be loaded with molecules of interest such as drugs, proteins and peptides, and are a promising new addition to the drug delivery platforms currently in use. The investigators also recently used the DNA origami nanoparticles to fine tune the spatial presentation of immunogens to study the impact on B cell activation. These studies are an important step towards the rational design of vaccines for a variety of infectious agents. To elucidate the parameters for optimizing the delivery efficiency of lipid nanoparticles (LNPs), Buschmann, Paige and colleagues have devised methods for predicting and experimentally validating the pKa of LNPs based on the structure of the ionizable lipids used to formulate the LNPs. These studies may pave the way for the development of new LNP delivery vehicles that have reduced systemic distribution and improved endosomal release of their cargo post administration. To better understand protein-protein interactions and identify potential drug targets that disrupt such interactions, Luchini and colleagues have developed a methodology that identifies contact points between proteins using small molecule dyes. The dye molecules noncovalently bind to the accessible surfaces of a protein complex with very high affinity, but are excluded from contact regions. When the complex is denatured and digested with trypsin, the exposed regions covered by the dye do not get cleaved by the enzyme, whereas the contact points are digested. The resulting fragments can then be identified using mass spectrometry. The data generated can serve as the basis for designing small molecules and peptides that can disrupt the formation of protein complexes involved in disease processes. For example, using peptides based on the interleukin 1 receptor accessory protein (IL-1RAcP), Luchini, Liotta, Paige and colleagues disrupted the formation of IL-1/IL-R/IL-1RAcP complex and demonstrated that the inhibition of complex formation reduced the inflammatory response to IL-1B. Working on the discovery of novel antimicrobial agents, Bishop, van Hoek and colleagues have discovered a number of antimicrobial peptides from reptiles and other species. DRGN-1, is a synthetic peptide based on a histone H1-derived peptide that they had identified from Komodo Dragon plasma. DRGN-1 was shown to disrupt bacterial biofilms and promote wound healing in an animal model. The peptide, along with others, is being developed and tested in preclinical studies. Other research by van Hoek and colleagues focuses on in silico antimicrobial peptide discovery, screening of small molecules for antibacterial properties, as well as assessment of diffusible signal factors (DFS) as future therapeutics. The above examples provide insight into the cutting-edge studies undertaken by GMU scientists to develop novel methodologies and platform technologies important to drug discovery.

59 BASIC BIOLOGICAL SCIENCES↗

American Heart Association Protein Atlas Portal for Accelerated Drug Discovery

The American Heart Association's Protein Atlas Data Portal is a cutting-edge, open-source atlas of the structural determinants of the binding of drugs to (all of) their target proteins as an unbiased means to accelerate drug discovery. The portal will provide researchers with access to comprehensive data that can help increase the selection of effective drugs and cut the time-to-market of new drugs by half.

drug targets↗

Green genes from blue greens: challenges and solutions to unlocking the potential of cyanobacteria in drug discovery

Cyanobacteria are prolific producers of biologically active compounds that are important in influencing ecology, behavior of interacting organisms, and as leads in drug discovery efforts. Here we discuss the challenges faced by all natural product researchers, especially those that focus on cyanobacteria, and then describe progress that has been made in these areas. We also propose some solutions, paths forward, and thoughts for consideration on these challenges.

Philmus, Benjamin↗

A deep generative model for deciphering cellular dynamics and in silico drug discovery in complex diseases

Human diseases are characterized by intricate cellular dynamics. Single-cell transcriptomics provides critical insights, yet a persistent gap remains in computational tools for detailed disease progression analysis and targeted in silico drug interventions. Here we introduce UNAGI, a deep generative neural network tailored to analyse time-series single-cell transcriptomic data. This tool captures the complex cellular dynamics underlying disease progression, enhancing drug perturbation modelling and screening. When applied to a dataset from patients with idiopathic pulmonary fibrosis, UNAGI learns disease-informed cell embeddings that sharpen our understanding of disease progression, leading to the identification of potential therapeutic drug candidates. Validation using proteomics reveals the accuracy of UNAGI’s cellular dynamics analysis, and the use of the fibrotic cocktail-treated human precision-cut lung slices confirms UNAGI’s predictions that nifedipine, an antihypertensive drug, may have anti-fibrotic effects on human tissues. UNAGI’s versatility extends to other diseases, including COVID, demonstrating adaptability and confirming its broader applicability in decoding complex cellular dynamics beyond idiopathic pulmonary fibrosis, amplifying its use in the quest for therapeutic solutions across diverse pathological landscapes.

Neural Network↗

Advances in Computational Approaches for Estimating Passive Permeability in Drug Discovery

Passive permeation of cellular membranes is a key feature of many therapeutics. The relevance of passive permeability spans all biological systems as they all employ biomembranes for compartmentalization. A variety of computational techniques are currently utilized and under active development to facilitate the characterization of passive permeability. These methods include lipophilicity relations, molecular dynamics simulations, and machine learning, which vary in accuracy, complexity, and computational cost. This review briefly introduces the underlying theories, such as the prominent inhomogeneous solubility diffusion model, and covers a number of recent applications. Various machine-learning applications, which have demonstrated good potential for high-volume, data-driven permeability predictions, are also discussed. Due to the confluence of novel computational methods and next-generation exascale computers, we anticipate an exciting future for computationally driven permeability predictions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Flow matching meets biology and life science: a survey

Over the past decade, advances in generative modeling, such as generative adversarial networks, masked autoencoders, and diffusion models, have significantly transformed biological research and discovery, enabling breakthroughs in molecule design, protein generation, catalysis discovery, drug discovery, and beyond. At the same time, biological applications have served as valuable testbeds for evaluating the capabilities of generative models. Recently, flow matching has emerged as a powerful and efficient alternative to diffusion-based generative modeling, with growing interest in its application to problems in biology and life sciences. This paper presents the first comprehensive survey of recent developments in flow matching and its applications in biological domains. We begin by systematically reviewing the foundations and variants of flow matching, and then categorize its applications into three major areas: biological sequence modeling, molecule generation and design, and peptide and protein generation. For each, we provide an in-depth review of recent progress. We also summarize commonly used datasets and software tools, and conclude with a discussion of potential future directions.

59 BASIC BIOLOGICAL SCIENCES↗

A Deep Multimodal Representation Learning Framework for Accurate Molecular Properties Prediction

Drug discovery is a complex and challenging process, requiring the optimization of candidate compounds to identify those with the potential to become safe and effective drugs. Predicting molecular properties is an indispensable step in the drug discovery pipeline. Traditionally, this process is costly and time-intensive, involving multiple rounds of experiments and clinical trials, rendering it impractical for every candidate compound. Deep learning techniques have emerged as a promising approach to drug discovery to reduce the cost and time required to identify novel drugs. However, prevalent research in deep learning models focused on predicting molecular properties has primarily fixated on single-modal models, which utilize a single modality of data, neglecting the potential benefits of combining different data modalities. To overcome this limitation, we introduce MRL-Mol: a deep \textbf{M}ultimodal \textbf{R}epresentation \textbf{L}earning framework for accurate \textbf{Mol}ecular properties prediction. MRL-Mol harnesses three data modalities: sequence, graph, and image, augmenting the depth of comprehension. Leveraging a large-scale unlabeled dataset~($\sim$1M unique molecules), we pretrain MRL-Mol to extract inter- and intra-modal information. Our study demonstrates the superior performance of MRL-Mol in predicting molecular properties across six benchmark datasets, including both classification and regression tasks. Notably, MRL-Mol outperforms other state-of-the-art molecular properties prediction models. These findings suggest that by combining information from multiple data modalities, MRL-Mol can comprehend molecules better than single-modal deep learning models and identify molecular properties with better accuracy.

Yang, Yuxin↗

Drug-induced kidney injury: challenges and opportunities

Abstract Drug-induced kidney injury (DIKI) is a frequently reported adverse event, associated with acute kidney injury, chronic kidney disease, and end-stage renal failure. Prospective cohort studies on acute injuries suggest a frequency of around 14%–26% in adult populations and a significant concern in pediatrics with a frequency of 16% being attributed to a drug. In drug discovery and development, renal injury accounts for 8 and 9% of preclinical and clinical failures, respectively, impacting multiple therapeutic areas. Currently, the standard biomarkers for identifying DIKI are serum creatinine and blood urea nitrogen. However, both markers lack the sensitivity and specificity to detect nephrotoxicity prior to a significant loss of renal function. Consequently, there is a pressing need for the development of alternative methods to reliably predict drug-induced kidney injury (DIKI) in early drug discovery. In this article, we discuss various aspects of DIKI and how it is assessed in preclinical models and in the clinical setting, including the challenges posed by translating animal data to humans. We then examine the urinary biomarkers accepted by both the US Food and Drug Administration (FDA) and the European Medicines Agency for monitoring DIKI in preclinical studies and on a case-by-case basis in clinical trials. We also review new approach methodologies (NAMs) and how they may assist in developing novel biomarkers for DIKI that can be used earlier in drug discovery and development.

Connor, Skylar (ORCID:0000000233479180)↗

Comparative Assessment of Pose Prediction Accuracy in RNA–Ligand Docking

Structure-based virtual high-throughput screening is used in early-stage drug discovery. Over the years, docking protocols and scoring functions for protein–ligand complexes have evolved to improve the accuracy in the computation of binding strengths and poses. In the past decade, RNA has also emerged as a target class for new small-molecule drugs. However, most ligand docking programs have been validated and tested for proteins and not RNA. Here, we test the docking power (pose prediction accuracy) of three state-of-the-art docking protocols on 173 RNA–small molecule crystal structures. The programs are AutoDock4 (AD4) and AutoDock Vina (Vina), which were designed for protein targets, and rDock, which was designed for both protein and nucleic acid targets. AD4 performed relatively poorly. For RNA targets for which a crystal structure of a bound ligand used to limit the docking search space is available and for which the goal is to identify new molecules for the same pocket, rDock performs slightly better than Vina, with success rates of 48% and 63%, respectively. However, in the more common type of early-stage drug discovery setting, in which no structure of a ligand–target complex is known and for which a larger search space is defined, rDock performed similarly to Vina, with a low success rate of ~27%. Further, Vina was found to have bias for ligands with certain physicochemical properties, whereas rDock performs similarly for all ligand properties. Thus, for projects where no ligand–protein structure already exists, Vina and rDock are both applicable. However, the relatively poor performance of all methods relative to protein–target docking illustrates a need for further methods refinement.

59 BASIC BIOLOGICAL SCIENCES↗

Labels as a feature: Network homophily for systematically annotating human GPCR drug-target interactions

Machine learning has revolutionized drug discovery by enabling the exploration of vast, uncharted chemical spaces essential for discovering novel patentable drugs. Despite the critical role of human G protein-coupled receptors in FDA-approved drugs, exhaustive in-distribution drug-target interaction testing across all pairs of human G protein-coupled receptors and known drugs is rare due to significant economic and technical challenges. This often leaves off-target effects unexplored, which poses a considerable risk to drug safety. In contrast to the traditional focus on out-of-distribution exploration (drug discovery), we introduce a neighborhood-to-prediction model termed Chemical Space Neural Networks that leverages network homophily and training-free graph neural networks with labels as features. We show that Chemical Space Neural Networks’ ability to make accurate predictions strongly correlates with network homophily. Thus, labels as features strongly increase a machine learning model’s capacity to enhance in-distribution prediction accuracy, which we show by integrating labeled data during inference. We validate these advancements in a high-throughput yeast biosensing system (3773 drug-target interactions, 539 compounds, 7 human G protein-coupled receptors) to discover novel drug-target interactions for FDA-approved drugs and to expand the general understanding of how to build reliable predictors to guide experimental verification.

Hansson, Frederik G↗

Ligand-Based Compound Activity Prediction via Few-Shot Learning

Predicting the activities of new compounds against biophysical or phenotypic assays based on the known activities of one or a few existing compounds is a common goal in early stage drug discovery. This problem can be cast as a “few-shot learning” challenge, and prior studies have developed few-shot learning methods to classify compounds as active versus inactive. However, the ability to go beyond classification and rank compounds by expected affinity is more valuable. We describe Few-Shot Compound Activity Prediction (FS-CAP), a novel neural architecture trained on a large bioactivity data set to predict compound activities against an assay outside the training set, based on only the activities of a few known compounds against the same assay. Our model aggregates encodings generated from the known compounds and their activities to capture assay information and uses a separate encoder for the new compound whose activity is to be predicted. The new method provides encouraging results relative to traditional chemical-similarity-based techniques as well as other state-of-the-art few-shot learning methods in tests on a variety of ligand-based drug discovery settings and data sets.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗