Search NASASearch

SEARCH · Search NASA

Results for “Drug discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Structure-Aware Unsupervised, Transformational Machine Learning for Drug Discovery (DTRA Basic Research Final Report)

The major goal of this project is to develop machine learning (ML) methods to enable improved predictive power on real drug discovery for novel targets. More specifically, we planned to demonstrate the capability and effectiveness of ML tools utilizing unlabeled large-volume protein-ligand datasets. We investigated multiple pre-training approaches for 3D protein-ligand structure-based foundation models, without relying on experimental binding data. We also addressed scenarios in which crystal structures are unavailable or binding data are limited. We also planned to develop a complete pipeline to screen novel compounds as well as to demonstrate the capability and effectiveness of the developed methods by testing on a realistic drug discovery task such as SARS-CoV-2. While the major goals and milestones remain consistent with the original proposal, certain technical details have been adjusted, based on the experimental results and related outcomes.

97 MATHEMATICS AND COMPUTING

Green genes from blue greens: challenges and solutions to unlocking the potential of cyanobacteria in drug discovery

Cyanobacteria are prolific producers of biologically active compounds that are important in influencing ecology, behavior of interacting organisms, and as leads in drug discovery efforts. Here we discuss the challenges faced by all natural product researchers, especially those that focus on cyanobacteria, and then describe progress that has been made in these areas. We also propose some solutions, paths forward, and thoughts for consideration on these challenges.

Philmus, Benjamin

A deep generative model for deciphering cellular dynamics and in silico drug discovery in complex diseases

Human diseases are characterized by intricate cellular dynamics. Single-cell transcriptomics provides critical insights, yet a persistent gap remains in computational tools for detailed disease progression analysis and targeted in silico drug interventions. Here we introduce UNAGI, a deep generative neural network tailored to analyse time-series single-cell transcriptomic data. This tool captures the complex cellular dynamics underlying disease progression, enhancing drug perturbation modelling and screening. When applied to a dataset from patients with idiopathic pulmonary fibrosis, UNAGI learns disease-informed cell embeddings that sharpen our understanding of disease progression, leading to the identification of potential therapeutic drug candidates. Validation using proteomics reveals the accuracy of UNAGI’s cellular dynamics analysis, and the use of the fibrotic cocktail-treated human precision-cut lung slices confirms UNAGI’s predictions that nifedipine, an antihypertensive drug, may have anti-fibrotic effects on human tissues. UNAGI’s versatility extends to other diseases, including COVID, demonstrating adaptability and confirming its broader applicability in decoding complex cellular dynamics beyond idiopathic pulmonary fibrosis, amplifying its use in the quest for therapeutic solutions across diverse pathological landscapes.

Neural Network

Flow matching meets biology and life science: a survey

Over the past decade, advances in generative modeling, such as generative adversarial networks, masked autoencoders, and diffusion models, have significantly transformed biological research and discovery, enabling breakthroughs in molecule design, protein generation, catalysis discovery, drug discovery, and beyond. At the same time, biological applications have served as valuable testbeds for evaluating the capabilities of generative models. Recently, flow matching has emerged as a powerful and efficient alternative to diffusion-based generative modeling, with growing interest in its application to problems in biology and life sciences. This paper presents the first comprehensive survey of recent developments in flow matching and its applications in biological domains. We begin by systematically reviewing the foundations and variants of flow matching, and then categorize its applications into three major areas: biological sequence modeling, molecule generation and design, and peptide and protein generation. For each, we provide an in-depth review of recent progress. We also summarize commonly used datasets and software tools, and conclude with a discussion of potential future directions.

59 BASIC BIOLOGICAL SCIENCES

Labels as a feature: Network homophily for systematically annotating human GPCR drug-target interactions

Machine learning has revolutionized drug discovery by enabling the exploration of vast, uncharted chemical spaces essential for discovering novel patentable drugs. Despite the critical role of human G protein-coupled receptors in FDA-approved drugs, exhaustive in-distribution drug-target interaction testing across all pairs of human G protein-coupled receptors and known drugs is rare due to significant economic and technical challenges. This often leaves off-target effects unexplored, which poses a considerable risk to drug safety. In contrast to the traditional focus on out-of-distribution exploration (drug discovery), we introduce a neighborhood-to-prediction model termed Chemical Space Neural Networks that leverages network homophily and training-free graph neural networks with labels as features. We show that Chemical Space Neural Networks’ ability to make accurate predictions strongly correlates with network homophily. Thus, labels as features strongly increase a machine learning model’s capacity to enhance in-distribution prediction accuracy, which we show by integrating labeled data during inference. We validate these advancements in a high-throughput yeast biosensing system (3773 drug-target interactions, 539 compounds, 7 human G protein-coupled receptors) to discover novel drug-target interactions for FDA-approved drugs and to expand the general understanding of how to build reliable predictors to guide experimental verification.

Hansson, Frederik G

ACES-GNN: can graph neural network learn to explain activity cliffs?

Graph Neural Networks (GNNs) have revolutionized molecular property prediction by leveraging graph-based representations, yet their opaque decision-making processes hinder broader adoption in drug discovery. This study introduces the Activity-Cliff-Explanation-Supervised GNN (ACES-GNN) framework, designed to simultaneously improve predictive accuracy and interpretability by integrating explanation supervision for activity cliffs (ACs) into GNN training. ACs, defined by structurally similar molecules with significant potency differences, pose challenges for traditional models due to their reliance on shared structural features. By aligning model attributions with chemist-friendly interpretations, the ACES-GNN framework bridges the gap between prediction and explanation. Validated across 30 pharmacological targets, ACES-GNN consistently enhances both predictive accuracy and attribution quality for ACs compared to unsupervised GNNs. Our results demonstrate a positive correlation between improved predictions and accurate explanations, offering a robust and adaptable framework to better understand and interpret ACs. This work underscores the potential of explanation-guided learning to advance interpretable artificial intelligence in molecular modeling and drug discovery.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Membrane protein reconstitution : New possibilities for structural biology, biophysical methods, and antibody/drug discovery

Nearly one-third of all proteins in eukaryotes are membrane proteins. Moreover, roughly 60% of Food and Drug Adminstration (FDA)-approved small-molecule drugs act on membrane proteins, which includes G protein–coupled receptors (GPCRs), ion channels, and transporters. Here, the vast majority of these membrane proteins are cell-surface accessible and thus amenable to drug discovery. At the same time, they are considerably more challenging to reconstitute and prepare for structure initiatives, antibody discovery, and drug screening. This series of reviews introduces the reader to current reconstitution systems, biophysical characterization of the membrane proteins and associated lipids, and common applications involving nuclear magnetic resonance (NMR), mass spectrometry (MS), and cryo-electron microscopy (cryo-EM).

36 MATERIALS SCIENCE

A goldilocks computational protocol for inhibitor discovery targeting DNA damage responses including replication-repair functions

While many researchers can design knockdown and knockout methodologies to remove a gene product, this is mainly untrue for new chemical inhibitor designs that empower multifunctional DNA Damage Response (DDR) networks. Here, we present a robust Goldilocks (GL) computational discovery protocol to efficiently innovate inhibitor tools and preclinical drug candidates for cellular and structural biologists without requiring extensive virtual screen (VS) and chemical synthesis expertise. By computationally targeting DDR replication and repair proteins, we exemplify the identification of DDR target sites and compounds to probe cancer biology. Our GL pipeline integrates experimental and predicted structures to efficiently discover leads, allowing early-structure and early-testing (ESET) experiments by many laboratories. By employing an efficient VS protocol to examine protein-protein interfaces (PPIs) and allosteric interactions, we identify ligand binding sites beyond active sites, leveraging in silico advances for molecular docking and modeling to screen PPIs and multiple targets. A diverse 3,174 compound ESET library combines Diamond Light Source DSI-poised, Protein Data Bank fragments, and FDA-approved drugs to span relevant chemotypes and facilitate downstream hit evaluation efficiency for academic laboratories. Two VS per library and multiple ranked ligand binding poses enable target testing for several DDR targets. This GL library and protocol can thus strategically probe multiple DDR network targets and identify readily available compounds for early structural and activity testing to overcome bottlenecks that can limit timely breakthrough drug discoveries. By testing accessible compounds to dissect multi-functional DDRs and suggesting inhibitor mechanisms from initial docking, the GL approach may enable more groups to help accelerate discovery, suggest new sites and compounds for challenging targets including emerging biothreats and advance cancer biology for future precision medicine clinical trials.

59 BASIC BIOLOGICAL SCIENCES

Knowledge graph-aided Bayesian active learning for top- K genetic interaction discovery

In silico methods for predicting the effects of multi-gene perturbations hold great promise for advancing functional genomics, computational drug discovery, and disease modeling. However, the development of these predictive algorithms for mammalian systems has been hampered by limited datasets and high experimental costs. In this study, we present a Bayesian active learning framework designed to discover pairwise host gene knockdowns that effectively inhibit viral proliferation in an in vitro HIV-1 infection model. Our method leverages a biological knowledge graph as side information and employs a computationally efficient batch diversification approach. We evaluated this framework using a dataset of viral load measurements obtained from multi-day dual-gene depletion experiments, encompassing all possible pairwise knockdowns of over 350 host genes associated with HIV infection. We demonstrate that our framework rapidly identifies the most effective gene knockdown pairs for reducing viral load. Furthermore, we show that incorporating side information enhances performance during the early stages of active learning (low data regime), while our batch diversification strategy significantly boosts performance in later stages (high data regime). This framework is general and can be adapted to explore gene interactions in other contexts, such as synthetic lethality prediction and mapping epistatic effects across quantitative trait loci.

Computational biology and bioinformatics

Architector 2.0: Expanded Capabilities for Metal Complex Engineering

Automated three-dimensional molecular construction from two-dimensional graph representations is critical to high-throughput discovery eIorts. Software capabilities in this area have accelerated research across fields ranging from protein design and drug discovery to transition metal catalyst development. When Architector was first introduced, it uniquely enabled high-throughput, chemically relevant three-dimensional construction of f-element complexes. Since its introduction, Architector has been applied in large-scale computational campaigns, targeted studies in critical mineral extraction, and artificial intelligence-driven discovery eIorts.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Mutation of active site glutamate in serine hydroxymethyltransferase allows trapping a reactive intermediate: a combined neutron and X-ray crystallography study

Serine hydroxymethyltransferase (SHMT) is a pyridoxal-5′-phosphate (PLP) dependent enzyme that catalyzes a chemical transformation essential for the one-carbon (1C) metabolism. SHMT reversibly converts L-Ser into Gly and transfers a 1C unit to tetrahydrofolate (THF) to give 5,10-methylene-THF (5,10-MTHF). 5,10-MTHF, a 1C-unit donor, plays a crucial role in the downstream biomolecular syntheses required for the cell homeostasis and proliferation. SHMT is a prominent target for the drug discovery to battle bacterial and parasitic infections, and to treat various types of cancer. SHMT-catalyzed chemistry is governed by the general acid-base catalysis. Knowledge of the catalytic mechanism can aid drug design but can only be achieved when the atomic details of each reaction step are mapped, including accurate determination of hydrogen atom positions. Here we utilized the inactive E53Q mutant of Thermus thermophilus ( Tth ) SHMT to directly determine protonation states with room-temperature neutron crystallography and to capture a reactive intermediate containing the PLP-L-Ser external aldimine and THF in the enzyme active site. We observed protonation of the Schiff base nitrogen (N SB ) in the PLP internal aldimine but no change in the protonation states of other ionizable PLP groups and active site residues compared to wild-type Tth SHMT. X-ray structural analysis of the ternary intermediate complex E53Q-Ser-THF that eluded previous structural characterization shows the strategic positioning of the E53Q side chain in close proximity to the external aldimine and THF and reinforces the proposed role for E53 as the driver of proton transfer events along the reaction pathway.

Drago, Victoria N. [Oak Ridge National Laboratory

Discovery of highly potent and ALK2/ALK1 selective kinase inhibitors using DNA-encoded chemistry technology

Activin receptor type 1 (ACVR1; ALK2) and activin receptor like type 1 (ACVRL1; ALK1) are transforming growth factor beta family receptors that integrate extracellular signals of bone morphogenic proteins (BMPs) and activins into Mothers Against Decapentaplegic homolog 1/5 (SMAD1/SMAD5) signaling complexes. Several activating mutations in ALK2 are implicated in fibrodysplasia ossificans progressiva (FOP), diffuse intrinsic pontine gliomas, and ependymomas. The ALK2 R206H mutation is also present in a subset of endometrial tumors, melanomas, non-small lung cancers, and colorectal cancers, and ALK2 expression is elevated in pancreatic cancer. Using DNA-encoded chemistry technology, we screened 3.94 billion unique compounds from our diverse DNA-encoded chemical libraries (DECLs) against the kinase domain of ALK2. Off-DNA synthesis of DECL hits and biochemical validation revealed nanomolar potent ALK2 inhibitors. Further structure-activity relationship studies yielded center for drug discovery (CDD)-2789, a potent [NanoBRET (NB) cell IC50: 0.54 μM] and metabolically stable analog with good pharmacological profile. Crystal structures of ALK2 bound with CDD-2281, CDD-2282, or CDD-2789 show that these inhibitors bind the active site through Van der Waals interactions and solvent-mediated hydrogen bonds. CDD-2789 exhibits high selectivity toward ALK2/ALK1 in KINOMEscan analysis and NB K192 assay. In cell-based studies, ALK2 inhibitors effectively attenuated activin A and BMP-induced Phosphorylated SMAD1/5 activation in fibroblasts from individuals with FOP in a dose-dependent manner. Thus, CDD-2789 is a valuable tool compound for further investigation of the biological functions of ALK2 and ALK1 and the therapeutic potential of specific inhibition of ALK2.

Jimmidi, Ravikumar

Optimization of Species-Selective Reversible Proteasome Inhibitors for the Treatment of Malaria

Abstract Malaria remains a critical global health challenge, with increasing resistance to frontline therapies necessitating novel drug targets. The proteasome has emerged as a promising target for antimalarial drug discovery. This study describes efforts to optimize a series of species-selective reversible inhibitors targeting the Plasmodium falciparum 20S proteasome. Starting from the carboxypiperidine scaffold identified through a high-throughput viability screen, we conducted iterative structure–activity relationship studies, leading to the development of highly potent and selective inhibitors with good oral bioavailability. Lead compounds demonstrated nanomolar potency against P. falciparum blood-stage parasites and selective inhibition of the parasite proteasome over the human counterpart. Cryo-EM structural studies confirmed binding at the β5 subunit, while in vivo pharmacokinetic studies identified promising candidates for further development. These findings support proteasome inhibition as a viable strategy for novel antimalarial drug development.

Gahalawat, Suraksha [UT Southwestern Medical Cente

The 'Biologically-Inspired Computing' Column

The field of Biology changed dramatically in 1953, with the determination by Francis Crick and James Dewey Watson of the double helix structure of DNA. This discovery changed Biology for ever, allowing the sequencing of the human genome, and the emergence of a "new Biology" focused on DNA, genes, proteins, data, and search. Computational Biology and Bioinformatics heavily rely on computing to facilitate research into life and development. Simultaneously, an understanding of the biology of living organisms indicates a parallel with computing systems: molecules in living cells interact, grow, and transform according to the "program" dictated by DNA. Moreover, paradigms of Computing are emerging based on modelling and developing computer-based systems exploiting ideas that are observed in nature. This includes building into computer systems self-management and self-governance mechanisms that are inspired by the human body's autonomic nervous system, modelling evolutionary systems analogous to colonies of ants or other insects, and developing highly-efficient and highly-complex distributed systems from large numbers of (often quite simple) largely homogeneous components to reflect the behaviour of flocks of birds, swarms of bees, herds of animals, or schools of fish. This new field of "Biologically-Inspired Computing", often known in other incarnations by other names, such as: Autonomic Computing, Pervasive Computing, Organic Computing, Biomimetics, and Artificial Life, amongst others, is poised at the intersection of Computer Science, Engineering, Mathematics, and the Life Sciences. Successes have been reported in the fields of drug discovery, data communications, computer animation, control and command, exploration systems for space, undersea, and harsh environments, to name but a few, and augur much promise for future progress.

Hinchey, Mike

Automated Label‐Free Assay for Viral Detection and Inhibitor Screening via Biomembrane‐Functionalized Microelectrode Arrays

Most virus infection assays have indirect readout such as virus number following entry (e.g., PCR, cell lysis). While effective, these technologies are labor‐intensive, require specialized environments (e.g., sterile or RNA‐free), and detect later‐stage viral events like lysis or cell death, lacking sensitivity to early fusion events. To address these limitations, we present biologically relevant 2D membrane materials, host‐cell‐derived supported lipid bilayers (hcd‐SLBs), integrated with organic microelectrode arrays (OMEAs) for detection of severe acute respiratory syndrome coronavirus 2 (SARS‐CoV‐2) fusion. By overexpressing angiotensin‐converting enzyme 2 (ACE2) receptors on the native membranes, the platform functions as a viral sensor capable of detecting virus pseudo particles (VPPs) through the late pathway. Additionally, hcd‐SLBs extracted from human lung epithelium expressing native ACE2 detect fusion events through the early pathway. The platform's utility as a drug‐screening tool is demonstrated by testing antibodies targeting either the ACE2 on the host membrane or the viral spike (S) proteins. To enhance the throughput, microfluidics are integrated for automation and OMEAs are incorporated within each channel, miniaturizing the testing units. This system supports high‐throughput data generation, automation, and scalability, providing an efficient platform for viral fusion detection that advances the study of pathogen‐host interactions and accelerates antiviral drug discovery.

Biology

Polysulfamates as “Macroisosteres” of Polyurethanes with Improved Degradability

Abstract Addressing the environmental persistence of plastics requires the development of next‐generation polymers that combine high performance with enhanced degradability. Progress toward this grand challenge has been impeded, in part, by the absence of a general blueprint for the macromolecular design of such materials. Herein, we introduce a “macroisostere” design strategy, where the carbonyl group (–CO–) in polyurethanes (PUs) is replaced with a sulfonyl group (–SO 2 –), resulting in a virtually unknown family of polymers called polysulfamates. This approach, inspired by the use of bioisosteres in drug discovery, aims to preserve key interchain interactions that contribute to thermomechanical performance while enhancing the hydrolytic lability of the polymer backbone. The optimization of a Sulfur(VI) Fluoride Exchange (SuFEx) polymerization allowed the synthesis of ten polysulfamates structurally analogous to common PUs. Comparative analysis of one PU and its polysulfamate analog showed that this isosteric substitution increases thermal stability, slightly lowers the glass transition temperature, and retains similar hardness and reduced Young's modulus. Notably, the S(VI)‐based polysulfamate demonstrated significantly enhanced hydrolytic degradability. These results highlight the potential of the “macroisostere” approach as a generalizable strategy for designing high‐performance, degradable alternatives to traditional plastics.

Chemistry

A Step-by-Step Protocol from METASPACE to Biological Interpretation

Mass spectrometry imaging (MSI) represents an exceptional tool for exploring complex biological systems spatially at the molecular level. However, due to its multidimensional nature and large-scale data output, it presents considerable challenges when it comes to extracting meaningful biological insights. Recent advancements, such as the METASPACE platform, have enabled researchers to efficiently process, annotate, and interpret MSI datasets by leveraging machine learning and cloud-based infrastructure. In this tutorial, we present a detailed and user-friendly R-pipeline designed to help METASPACE users navigate untargeted metabolomic annotations and transform them into practical insights about their biological systems. By combining METASPACE annotations with rapid R-based screening, this workflow not only streamlined the analytical process but also enhanced the understanding of spatial molecular distribution, especially for complex systems. Here, this easy-to-follow approach has the potential for applications in diagnostics, drug discovery, environmental and ecological processes, and more. We envision this pipeline to be particularly useful for newcomers to the field of MSI and

Moreno Pedraza, Abigail

Interpretable, extensible linear and symbolic regression models for charge density prediction using a hierarchy of many-body correlation descriptors

Here, density functional theory (DFT) is routinely used to make electronic structure predictions for high-throughput screening of materials and molecules for technologically relevant areas, like the identification of better catalysts, electronic materials, and drug discovery. However, the DFT formalism is limited by (a) its poor (quadratic-to-quartic) scaling, and (b) the need to perform repeated eigenvalue computations of the electronic Hamiltonian as part of its self-consistent field (SCF) iteration procedure to obtain the converged ground state electron density, ρ (r). Approaches that directly predict ρ (r) of a structure with high accuracy can accelerate conventional SCF calculations and can also be used in linearly scaling methods such as orbital-free DFT. To this end, we present a procedure to predict the ground state electron density of molecular and periodic three-dimensional systems directly from the atomic structure with a particular emphasis on physical interpretability. In our framework, ρ (r) is modeled using many-body correlation descriptors that accurately capture the effects of local atomic arrangements in the neighborhood of a grid point. Our use of a linear regression scheme to fit to charge density data enables transparent analysis of the relative contributions of various types of local atomic correlations. By systematically including increasingly complex correlations, our model is shown to accurately predict ρ (r) for a variety of chemically and electronically diverse systems — amorphous Ge, Al(001) slab, crystalline Ga 2 O 3 , molecular benzene, and polyethylene. We then demonstrate a symbolic regression-based protocol to construct easily computable, interpretable features from lower-order correlations that significantly improves our electron density predictions with effectively no increase in the computational cost.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC