Search NASASearch

SEARCH · Search NASA

Results for “Drug discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

An artificial intelligence accelerated virtual screening platform for drug discovery

Abstract Structure-based virtual screening is a key tool in early drug discovery, with growing interest in the screening of multi-billion chemical compound libraries. However, the success of virtual screening crucially depends on the accuracy of the binding pose and binding affinity predicted by computational docking. Here we develop a highly accurate structure-based virtual screen method, RosettaVS, for predicting docking poses and binding affinities. Our approach outperforms other state-of-the-art methods on a wide range of benchmarks, partially due to our ability to model receptor flexibility. We incorporate this into a new open-source artificial intelligence accelerated virtual screening platform for drug discovery. Using this platform, we screen multi-billion compound libraries against two unrelated targets, a ubiquitin ligase target KLHDC2 and the human voltage-gated sodium channel Na V 1.7. For both targets, we discover hit compounds, including seven hits (14% hit rate) to KLHDC2 and four hits (44% hit rate) to Na V 1.7, all with single digit micromolar binding affinities. Screening in both cases is completed in less than seven days. Finally, a high resolution X-ray crystallographic structure validates the predicted docking pose for the KLHDC2 ligand complex, demonstrating the effectiveness of our method in lead discovery.

Science & Technology - Other Topics

Annual Report for Structure-Aware Unsupervised, Transformational Machine Learning for Drug Discovery

The major goal of this project is to develop machine learning (ML) methods to enable improved predictive power on real drug discovery for novel targets. More specifically, we plan to demonstrate the capability and effectiveness of ML tools utilizing unlabeled large-volume protein-ligand datasets. We also plan to demonstrate the capability and effectiveness of the developed methods by testing on a realistic drug discovery task to identify pan-coronavirus protease inhibitors such as SARS-CoV-2. While the overall goals and milestones remain consistent with the original proposal, certain technical details have been modified, which we will describe in this report.

97 MATHEMATICS AND COMPUTING

Structure-Aware Unsupervised, Transformational Machine Learning for Drug Discovery (DTRA Basic Research Final Report)

The major goal of this project is to develop machine learning (ML) methods to enable improved predictive power on real drug discovery for novel targets. More specifically, we planned to demonstrate the capability and effectiveness of ML tools utilizing unlabeled large-volume protein-ligand datasets. We investigated multiple pre-training approaches for 3D protein-ligand structure-based foundation models, without relying on experimental binding data. We also addressed scenarios in which crystal structures are unavailable or binding data are limited. We also planned to develop a complete pipeline to screen novel compounds as well as to demonstrate the capability and effectiveness of the developed methods by testing on a realistic drug discovery task such as SARS-CoV-2. While the major goals and milestones remain consistent with the original proposal, certain technical details have been adjusted, based on the experimental results and related outcomes.

97 MATHEMATICS AND COMPUTING

Green genes from blue greens: challenges and solutions to unlocking the potential of cyanobacteria in drug discovery

Cyanobacteria are prolific producers of biologically active compounds that are important in influencing ecology, behavior of interacting organisms, and as leads in drug discovery efforts. Here we discuss the challenges faced by all natural product researchers, especially those that focus on cyanobacteria, and then describe progress that has been made in these areas. We also propose some solutions, paths forward, and thoughts for consideration on these challenges.

Philmus, Benjamin

A deep generative model for deciphering cellular dynamics and in silico drug discovery in complex diseases

Human diseases are characterized by intricate cellular dynamics. Single-cell transcriptomics provides critical insights, yet a persistent gap remains in computational tools for detailed disease progression analysis and targeted in silico drug interventions. Here we introduce UNAGI, a deep generative neural network tailored to analyse time-series single-cell transcriptomic data. This tool captures the complex cellular dynamics underlying disease progression, enhancing drug perturbation modelling and screening. When applied to a dataset from patients with idiopathic pulmonary fibrosis, UNAGI learns disease-informed cell embeddings that sharpen our understanding of disease progression, leading to the identification of potential therapeutic drug candidates. Validation using proteomics reveals the accuracy of UNAGI’s cellular dynamics analysis, and the use of the fibrotic cocktail-treated human precision-cut lung slices confirms UNAGI’s predictions that nifedipine, an antihypertensive drug, may have anti-fibrotic effects on human tissues. UNAGI’s versatility extends to other diseases, including COVID, demonstrating adaptability and confirming its broader applicability in decoding complex cellular dynamics beyond idiopathic pulmonary fibrosis, amplifying its use in the quest for therapeutic solutions across diverse pathological landscapes.

Neural Network

Flow matching meets biology and life science: a survey

Over the past decade, advances in generative modeling, such as generative adversarial networks, masked autoencoders, and diffusion models, have significantly transformed biological research and discovery, enabling breakthroughs in molecule design, protein generation, catalysis discovery, drug discovery, and beyond. At the same time, biological applications have served as valuable testbeds for evaluating the capabilities of generative models. Recently, flow matching has emerged as a powerful and efficient alternative to diffusion-based generative modeling, with growing interest in its application to problems in biology and life sciences. This paper presents the first comprehensive survey of recent developments in flow matching and its applications in biological domains. We begin by systematically reviewing the foundations and variants of flow matching, and then categorize its applications into three major areas: biological sequence modeling, molecule generation and design, and peptide and protein generation. For each, we provide an in-depth review of recent progress. We also summarize commonly used datasets and software tools, and conclude with a discussion of potential future directions.

59 BASIC BIOLOGICAL SCIENCES

A Deep Multimodal Representation Learning Framework for Accurate Molecular Properties Prediction

Drug discovery is a complex and challenging process, requiring the optimization of candidate compounds to identify those with the potential to become safe and effective drugs. Predicting molecular properties is an indispensable step in the drug discovery pipeline. Traditionally, this process is costly and time-intensive, involving multiple rounds of experiments and clinical trials, rendering it impractical for every candidate compound. Deep learning techniques have emerged as a promising approach to drug discovery to reduce the cost and time required to identify novel drugs. However, prevalent research in deep learning models focused on predicting molecular properties has primarily fixated on single-modal models, which utilize a single modality of data, neglecting the potential benefits of combining different data modalities. To overcome this limitation, we introduce MRL-Mol: a deep \textbf{M}ultimodal \textbf{R}epresentation \textbf{L}earning framework for accurate \textbf{Mol}ecular properties prediction. MRL-Mol harnesses three data modalities: sequence, graph, and image, augmenting the depth of comprehension. Leveraging a large-scale unlabeled dataset~($\sim$1M unique molecules), we pretrain MRL-Mol to extract inter- and intra-modal information. Our study demonstrates the superior performance of MRL-Mol in predicting molecular properties across six benchmark datasets, including both classification and regression tasks. Notably, MRL-Mol outperforms other state-of-the-art molecular properties prediction models. These findings suggest that by combining information from multiple data modalities, MRL-Mol can comprehend molecules better than single-modal deep learning models and identify molecular properties with better accuracy.

Yang, Yuxin

Drug-induced kidney injury: challenges and opportunities

Abstract Drug-induced kidney injury (DIKI) is a frequently reported adverse event, associated with acute kidney injury, chronic kidney disease, and end-stage renal failure. Prospective cohort studies on acute injuries suggest a frequency of around 14%–26% in adult populations and a significant concern in pediatrics with a frequency of 16% being attributed to a drug. In drug discovery and development, renal injury accounts for 8 and 9% of preclinical and clinical failures, respectively, impacting multiple therapeutic areas. Currently, the standard biomarkers for identifying DIKI are serum creatinine and blood urea nitrogen. However, both markers lack the sensitivity and specificity to detect nephrotoxicity prior to a significant loss of renal function. Consequently, there is a pressing need for the development of alternative methods to reliably predict drug-induced kidney injury (DIKI) in early drug discovery. In this article, we discuss various aspects of DIKI and how it is assessed in preclinical models and in the clinical setting, including the challenges posed by translating animal data to humans. We then examine the urinary biomarkers accepted by both the US Food and Drug Administration (FDA) and the European Medicines Agency for monitoring DIKI in preclinical studies and on a case-by-case basis in clinical trials. We also review new approach methodologies (NAMs) and how they may assist in developing novel biomarkers for DIKI that can be used earlier in drug discovery and development.

Connor, Skylar (ORCID:0000000233479180)

Labels as a feature: Network homophily for systematically annotating human GPCR drug-target interactions

Machine learning has revolutionized drug discovery by enabling the exploration of vast, uncharted chemical spaces essential for discovering novel patentable drugs. Despite the critical role of human G protein-coupled receptors in FDA-approved drugs, exhaustive in-distribution drug-target interaction testing across all pairs of human G protein-coupled receptors and known drugs is rare due to significant economic and technical challenges. This often leaves off-target effects unexplored, which poses a considerable risk to drug safety. In contrast to the traditional focus on out-of-distribution exploration (drug discovery), we introduce a neighborhood-to-prediction model termed Chemical Space Neural Networks that leverages network homophily and training-free graph neural networks with labels as features. We show that Chemical Space Neural Networks’ ability to make accurate predictions strongly correlates with network homophily. Thus, labels as features strongly increase a machine learning model’s capacity to enhance in-distribution prediction accuracy, which we show by integrating labeled data during inference. We validate these advancements in a high-throughput yeast biosensing system (3773 drug-target interactions, 539 compounds, 7 human G protein-coupled receptors) to discover novel drug-target interactions for FDA-approved drugs and to expand the general understanding of how to build reliable predictors to guide experimental verification.

Hansson, Frederik G

Ligand-Based Compound Activity Prediction via Few-Shot Learning

Predicting the activities of new compounds against biophysical or phenotypic assays based on the known activities of one or a few existing compounds is a common goal in early stage drug discovery. This problem can be cast as a “few-shot learning” challenge, and prior studies have developed few-shot learning methods to classify compounds as active versus inactive. However, the ability to go beyond classification and rank compounds by expected affinity is more valuable. We describe Few-Shot Compound Activity Prediction (FS-CAP), a novel neural architecture trained on a large bioactivity data set to predict compound activities against an assay outside the training set, based on only the activities of a few known compounds against the same assay. Our model aggregates encodings generated from the known compounds and their activities to capture assay information and uses a separate encoder for the new compound whose activity is to be predicted. The new method provides encouraging results relative to traditional chemical-similarity-based techniques as well as other state-of-the-art few-shot learning methods in tests on a variety of ligand-based drug discovery settings and data sets.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

ACES-GNN: can graph neural network learn to explain activity cliffs?

Graph Neural Networks (GNNs) have revolutionized molecular property prediction by leveraging graph-based representations, yet their opaque decision-making processes hinder broader adoption in drug discovery. This study introduces the Activity-Cliff-Explanation-Supervised GNN (ACES-GNN) framework, designed to simultaneously improve predictive accuracy and interpretability by integrating explanation supervision for activity cliffs (ACs) into GNN training. ACs, defined by structurally similar molecules with significant potency differences, pose challenges for traditional models due to their reliance on shared structural features. By aligning model attributions with chemist-friendly interpretations, the ACES-GNN framework bridges the gap between prediction and explanation. Validated across 30 pharmacological targets, ACES-GNN consistently enhances both predictive accuracy and attribution quality for ACs compared to unsupervised GNNs. Our results demonstrate a positive correlation between improved predictions and accurate explanations, offering a robust and adaptable framework to better understand and interpret ACs. This work underscores the potential of explanation-guided learning to advance interpretable artificial intelligence in molecular modeling and drug discovery.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Membrane protein reconstitution : New possibilities for structural biology, biophysical methods, and antibody/drug discovery

Nearly one-third of all proteins in eukaryotes are membrane proteins. Moreover, roughly 60% of Food and Drug Adminstration (FDA)-approved small-molecule drugs act on membrane proteins, which includes G protein–coupled receptors (GPCRs), ion channels, and transporters. Here, the vast majority of these membrane proteins are cell-surface accessible and thus amenable to drug discovery. At the same time, they are considerably more challenging to reconstitute and prepare for structure initiatives, antibody discovery, and drug screening. This series of reviews introduces the reader to current reconstitution systems, biophysical characterization of the membrane proteins and associated lipids, and common applications involving nuclear magnetic resonance (NMR), mass spectrometry (MS), and cryo-electron microscopy (cryo-EM).

36 MATERIALS SCIENCE

Protein Structure Inspired Discovery of a Novel Inducer of Anoikis in Human Melanoma

Drug discovery historically starts with an established function, either that of compounds or proteins. This can hamper discovery of novel therapeutics. As structure determines function, we hypothesized that unique 3D protein structures constitute primary data that can inform novel discovery. Using a computationally intensive physics-based analytical platform operating at supercomputing speeds, we probed a high-resolution protein X-ray crystallographic library developed by us. For each of the eight identified novel 3D structures, we analyzed binding of sixty million compounds. Top-ranking compounds were acquired and screened for efficacy against breast, prostate, colon, or lung cancer, and for toxicity on normal human bone marrow stem cells, both using eight-day colony formation assays. Effective and non-toxic compounds segregated to two pockets. One compound, Dxr2-017, exhibited selective anti-melanoma activity in the NCI-60 cell line screen. In eight-day assays, Dxr2-017 had an IC50 of 12 nM against melanoma cells, while concentrations over 2100-fold higher had minimal stem cell toxicity. Dxr2-017 induced anoikis, a unique form of programmed cell death in need of targeted therapeutics. Our findings demonstrate proof-of-concept that protein structures represent high-value primary data to support the discovery of novel acting therapeutics. This approach is widely applicable.

Oncology

Impact of structural biology and the protein data bank on us fda new drug approvals of low molecular weight antineoplastic agents 2019–2023

Abstract Open access to three-dimensional atomic-level biostructure information from the Protein Data Bank (PDB) facilitated discovery/development of 100% of the 34 new low molecular weight, protein-targeted, antineoplastic agents approved by the US FDA 2019–2023. Analyses of PDB holdings, the scientific literature, and related documents for each drug-target combination revealed that the impact of structural biologists and public-domain 3D biostructure data was broad and substantial, ranging from understanding target biology (100% of all drug targets), to identifying a given target as likely druggable (100% of all targets), to structure-guided drug discovery (>80% of all new small-molecule drugs, made up of 50% confirmed and >30% probable cases). In addition to aggregate impact assessments, illustrative case studies are presented for six first-in-class small-molecule anti-cancer drugs, including a selective inhibitor of nuclear export targeting Exportin 1 (selinexor, Xpovio), an ATP-competitive CSF-1R receptor tyrosine kinase inhibitor (pexidartinib,Turalia), a non-ATP-competitive inhibitor of the BCR-Abl fusion protein targeting the myristoyl binding pocket within the kinase catalytic domain of Abl (asciminib, Scemblix), a covalently-acting G12C KRAS inhibitor (sotorasib, Lumakras or Lumykras), an EZH2 methyltransferase inhibitor (tazemostat, Tazverik), and an agent targeting the basic-Helix-Loop-Helix transcription factor HIF-2α (belzutifan, Welireg).

60 APPLIED LIFE SCIENCES

An expedited screening platform for the discovery of anti-ageing compounds in vitro and in vivo

Background: Restraining or slowing ageing hallmarks at the cellular level have been proposed as a route to increased organismal lifespan and healthspan. Consequently, there is great interest in anti-ageing drug discovery. However, this currently requires laborious and lengthy longevity analysis. Here, we present a novel screening readout for the expedited discovery of compounds that restrain ageing of cell populations in vitro and enable extension of in vivo lifespan. Methods: Using Illumina methylation arrays, we monitored DNA methylation changes accompanying long-term passaging of adult primary human cells in culture. This enabled us to develop, test, and validate the CellPopAge Clock, an epigenetic clock with underlying algorithm, unique among existing epigenetic clocks for its design to detect anti-ageing compounds in vitro. Additionally, we measured markers of senescence and performed longevity experiments in vivo in Drosophila, to further validate our approach to discover novel anti-ageing compounds. Finally, we bench mark our epigenetic clock with other available epigenetic clocks to consolidate its usefulness and specialisation for primary cells in culture. Results: We developed a novel epigenetic clock, the CellPopAge Clock, to accurately monitor the age of a population of adult human primary cells. We find that the CellPopAge Clock can detect decelerated passage-based ageing of human primary cells treated with rapamycin or trametinib, well-established longevity drugs. We then utilise the CellPopAge Clock as a screening tool for the identification of compounds which decelerate ageing of cell populations, uncovering novel anti-ageing drugs, torin2 and dactolisib (BEZ-235). We demonstrate that delayed epigenetic ageing in human primary cells treated with anti-ageing compounds is accompanied by a reduction in senescence and ageing biomarkers. Finally, we extend our screening platform in vivo by taking advantage of a specially formulated holidic medium for increased drug bioavailability in Drosophila. We show that the novel anti-ageing drugs, torin2 and dactolisib (BEZ-235), increase longevity in vivo. Conclusions: Our method expands the scope of CpG methylation profiling to accurately and rapidly detecting anti-ageing potential of drugs using human cells in vitro, and in vivo, providing a novel accelerated discovery platform to test sought after anti-ageing compounds and geroprotectors.

60 APPLIED LIFE SCIENCES

A goldilocks computational protocol for inhibitor discovery targeting DNA damage responses including replication-repair functions

While many researchers can design knockdown and knockout methodologies to remove a gene product, this is mainly untrue for new chemical inhibitor designs that empower multifunctional DNA Damage Response (DDR) networks. Here, we present a robust Goldilocks (GL) computational discovery protocol to efficiently innovate inhibitor tools and preclinical drug candidates for cellular and structural biologists without requiring extensive virtual screen (VS) and chemical synthesis expertise. By computationally targeting DDR replication and repair proteins, we exemplify the identification of DDR target sites and compounds to probe cancer biology. Our GL pipeline integrates experimental and predicted structures to efficiently discover leads, allowing early-structure and early-testing (ESET) experiments by many laboratories. By employing an efficient VS protocol to examine protein-protein interfaces (PPIs) and allosteric interactions, we identify ligand binding sites beyond active sites, leveraging in silico advances for molecular docking and modeling to screen PPIs and multiple targets. A diverse 3,174 compound ESET library combines Diamond Light Source DSI-poised, Protein Data Bank fragments, and FDA-approved drugs to span relevant chemotypes and facilitate downstream hit evaluation efficiency for academic laboratories. Two VS per library and multiple ranked ligand binding poses enable target testing for several DDR targets. This GL library and protocol can thus strategically probe multiple DDR network targets and identify readily available compounds for early structural and activity testing to overcome bottlenecks that can limit timely breakthrough drug discoveries. By testing accessible compounds to dissect multi-functional DDRs and suggesting inhibitor mechanisms from initial docking, the GL approach may enable more groups to help accelerate discovery, suggest new sites and compounds for challenging targets including emerging biothreats and advance cancer biology for future precision medicine clinical trials.

59 BASIC BIOLOGICAL SCIENCES

Knowledge graph-aided Bayesian active learning for top- K genetic interaction discovery

In silico methods for predicting the effects of multi-gene perturbations hold great promise for advancing functional genomics, computational drug discovery, and disease modeling. However, the development of these predictive algorithms for mammalian systems has been hampered by limited datasets and high experimental costs. In this study, we present a Bayesian active learning framework designed to discover pairwise host gene knockdowns that effectively inhibit viral proliferation in an in vitro HIV-1 infection model. Our method leverages a biological knowledge graph as side information and employs a computationally efficient batch diversification approach. We evaluated this framework using a dataset of viral load measurements obtained from multi-day dual-gene depletion experiments, encompassing all possible pairwise knockdowns of over 350 host genes associated with HIV infection. We demonstrate that our framework rapidly identifies the most effective gene knockdown pairs for reducing viral load. Furthermore, we show that incorporating side information enhances performance during the early stages of active learning (low data regime), while our batch diversification strategy significantly boosts performance in later stages (high data regime). This framework is general and can be adapted to explore gene interactions in other contexts, such as synthetic lethality prediction and mapping epistatic effects across quantitative trait loci.

Computational biology and bioinformatics