Search NASA⌕ Search

SEARCH · Search NASA

Results for “virtual screening”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

An artificial intelligence accelerated virtual screening platform for drug discovery

Abstract Structure-based virtual screening is a key tool in early drug discovery, with growing interest in the screening of multi-billion chemical compound libraries. However, the success of virtual screening crucially depends on the accuracy of the binding pose and binding affinity predicted by computational docking. Here we develop a highly accurate structure-based virtual screen method, RosettaVS, for predicting docking poses and binding affinities. Our approach outperforms other state-of-the-art methods on a wide range of benchmarks, partially due to our ability to model receptor flexibility. We incorporate this into a new open-source artificial intelligence accelerated virtual screening platform for drug discovery. Using this platform, we screen multi-billion compound libraries against two unrelated targets, a ubiquitin ligase target KLHDC2 and the human voltage-gated sodium channel Na V 1.7. For both targets, we discover hit compounds, including seven hits (14% hit rate) to KLHDC2 and four hits (44% hit rate) to Na V 1.7, all with single digit micromolar binding affinities. Screening in both cases is completed in less than seven days. Finally, a high resolution X-ray crystallographic structure validates the predicted docking pose for the KLHDC2 ligand complex, demonstrating the effectiveness of our method in lead discovery.

Science & Technology - Other Topics↗

Optimal decision-making in high-throughput virtual screening pipelines

Screening large pools of molecular candidates to identify those with specific design criteria or targeted properties is demanding in various science and engineering domains. While a high-throughput virtual screening (HTVS) pipeline can provide efficient means to achieving this goal, its design and operation often rely on experts' intuition, potentially resulting in suboptimal performance. In this paper, we fill this critical gap by presenting a systematic framework that can maximize the return on computational investment (ROCI) of such HTVS campaigns. Based on various scenarios, we empirically validate the proposed framework and demonstrate its potential to accelerate scientific discoveries through optimal computational campaigns, especially in the context of virtual screening.

97 MATHEMATICS AND COMPUTING↗

Discovery of Novel Rhizoctonia solani DHFR Inhibitors as Fungicides Using Virtual Screening

Dihydrofolate reductase (DHFR) is an essential enzyme in the folate pathway and has been recognized as a well-known target for antibacterial and antifungal drugs. We discovered eight compounds from the ZINC database using virtual screening to inhibit Rhizoctonia solani (R. solani), a fungal pathogen in crops. These compounds were evaluated with in vitro assays for enzymatic and antifungal activity. Among these, compound Hit8 is the most active R. solani DHFR inhibitor, with the IC 50 of 10.2 μM. The selectivity of inhibition is 22.3 against human DHFR with the IC 50 of 227.7 μM. Moreover, Hit8 has higher antifungal activity against R. solani (EC 50 of 38.2 mg L –1 ) compared with validamycin A (EC 50 of 67.6 mg L –1 ), a well-documented fungicide. These results suggest that Hit8 may be a potential fungicide. Finally, our study exemplifies a computer-aided method to discover novel inhibitors that could target plant pathogenic fungi.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Identification of inhibitors against SARS-CoV-2 variants of concern using virtual screening and metadynamics-based enhanced sampling

Among the variants of SARS-CoV-2, some are more infectious than the Wild-type. Interestingly, these mutations enable the virus to evade the therapeutic efforts. Hence, there is a need for candidate drug molecules that can potently bind with all the variants. Here we have adopted a strategy combining virtual screening, molecular docking followed by rigorous sampling by metadynamics simulations to find candidate molecules. From our results we found four highly potent drug candidates that can bind to the Spike-RBD of all the variants of the virus. Additionally, we also found that certain signature residues on the RBM region commonly bind to each of these inhibitors. Thus, our study not only gives information on the chemical compounds, but also residues on the proteins which could be targeted for future drug and vaccine development studies.

60 APPLIED LIFE SCIENCES↗

Do Molecular Fingerprints Identify Diverse Active Drugs in Large-Scale Virtual Screening? (No)

Computational approaches for small-molecule drug discovery now regularly scale to the consideration of libraries containing billions of candidate small molecules. One promising approach to increased the speed of evaluating billion-molecule libraries is to develop succinct representations of each molecule that enable the rapid identification of molecules with similar properties. Molecular fingerprints are thought to provide a mechanism for producing such representations. Here, we explore the utility of commonly used fingerprints in the context of predicting similar molecular activity. We show that fingerprint similarity provides little discriminative power between active and inactive molecules for a target protein based on a known active—while they may sometimes provide some enrichment for active molecules in a drug screen, a screened data set will still be dominated by inactive molecules. We also demonstrate that high-similarity actives appear to share a scaffold with the query active, meaning that they could more easily be identified by structural enumeration. Furthermore, even when limited to only active molecules, fingerprint similarity values do not correlate with compound potency. In sum, these results highlight the need for a new wave of molecular representations that will improve the capacity to detect biologically active molecules based on their similarity to other such molecules.

59 BASIC BIOLOGICAL SCIENCES↗

Advancing energy storage through solubility prediction: leveraging the potential of deep learning

Solubility prediction plays a crucial role in energy storage applications, such as redox flow batteries, because it directly affects the efficiency and reliability. Researchers have developed various methods that utilize quantum calculations and descriptors to predict the aqueous solubilities of organic molecules. Notably, machine learning models based on descriptors have shown promise for solubility prediction. As deep learning tools, graph neural networks (GNNs) have emerged to capture complex structure–property relationships for material property prediction. Specifically, MolGAT, a type of GNN model, was designed to incorporate n-dimensional edge attributes, enabling the modeling of intricacies in molecular graphs and enhancing the prediction capabilities. In a previous study, MolGAT successfully screened 23 467 promising redox-active molecules from a database of over 500 000 compounds, based on redox potential predictions. This study focused on applying the MolGAT model to predict the aqueous solubility (log S) of a broad range of organic compounds, including those previously screened for redox activity. The model was trained on a diverse sample of 8494 organic molecules from AqSolDB and benchmarked against literature data, demonstrating superior accuracy compared with other state of the art graph-based and descriptor-based models. Subsequently, the trained MolGAT model was employed to screen redox-active organic compounds identified in the first phase of high-throughput virtual screening, targeting favorable solubility in energy storage applications. The second round of screening, which considered solubility, yielded 12 332 promising redox-active and soluble organic molecules suitable for use in aqueous redox flow batteries. Thus, the two-phase high-throughput virtual screening approach utilizing MolGAT, specifically trained for redox potential and solubility, is an effective strategy for selecting suitable intrinsically soluble redox-active molecules from extensive databases, potentially advancing energy storage through reliable material development. This indicates that the model is reliable for predicting the solubility of various molecules and provides valuable insights for energy storage, pharmaceutical, environmental, and chemical applications.

25 ENERGY STORAGE↗

Multivariate chemogenomic screening prioritizes new macrofilaricidal leads

Development of direct acting macrofilaricides for the treatment of human filariases is hampered by limitations in screening throughput imposed by the parasite life cycle. In vitro adult screens typically assess single phenotypes without prior enrichment for chemicals with antifilarial potential. We developed a multivariate screen that identified dozens of compounds with submicromolar macrofilaricidal activity, achieving a hit rate of >50% by leveraging abundantly accessible microfilariae. Adult assays were multiplexed to thoroughly characterize compound activity across relevant parasite fitness traits, including neuromuscular control, fecundity, metabolism, and viability. Seventeen compounds from a diverse chemogenomic library elicited strong effects on at least one adult trait, with differential potency against microfilariae and adults. Our screen identified five compounds with high potency against adults but low potency or slow-acting microfilaricidal effects, at least one of which acts through a novel mechanism. We show that the use of microfilariae in a primary screen outperforms model nematode developmental assays and virtual screening of protein structures inferred with deep learning. These data provide new leads for drug development, and the high-content and multiplex assays set a new foundation for antifilarial discovery.

59 BASIC BIOLOGICAL SCIENCES↗

Exploration of structure-activity relationships for the SARS-CoV-2 macrodomain from shape-based fragment linking and active learning

The macrodomain of severe acute respiratory syndrome coronavirus 2 nonstructural protein 3 is required for viral pathogenesis and is an emerging antiviral target. We previously performed an x-ray crystallography–based fragment screen and found submicromolar inhibitors by fragment linking. However, these compounds had poor membrane permeability and liabilities that complicated optimization. Here, we developed a shape-based virtual screening pipeline—FrankenROCS. We screened the Enamine high-throughput collection of 2.1 million compounds, selecting 39 compounds for testing, with the most potent binding with a 130 μM median inhibitory concentration (IC 50 ). We then paired FrankenROCS with an active learning algorithm (Thompson sampling) to efficiently search the Enamine REAL database of 22 billion molecules, testing 32 compounds with the most potent binding with a 220 μM IC 50 . Further optimization led to analogs with IC 50 values better than 10 μM. This lead series has improved membrane permeability and is poised for optimization. FrankenROCS is a scalable method for fragment linking to exploit synthesis-on-demand libraries.

Science & Technology - Other Topics↗

Expediting field-effect transistor chemical sensor design with neuromorphic spiking graph neural networks

Improving the sensitive and selective detection of analytes in a variety of applications requires accelerating the rational design of field-effect transistor (FET) chemical sensors. Achieving high-performance detection relies on identifying optimal probe materials that can effectively interact with target analytes, a process traditionally driven by chemical intuition and time-consuming trial-and-error methods. To address the difficulties in probe screening for FET sensor development, this work presents a methodology that combines neuromorphic machine learning (ML) architectures, specifically a hybrid spiking graph neural network (SGNN), with an enriched dataset of physicochemical properties through semi-automated data extraction using large language models. Achieving a classification accuracy of 0.89 in predicting sensor sensitivity categories, the SGNN model outperformed traditional ML techniques by leveraging its ability to capture both global physicochemical properties and sparse topological features through a hybrid modeling framework. Next-generation sensor design was informed by the actionable insights into the connections between material properties and sensing performance offered by the SGNN framework. Through virtual screening for the detection of per- and polyfluoroalkyl substances (PFAS) as a use case, the effectiveness of the SGNN model was further validated. Density functional theory simulations confirmed graphene as a promising active material for PFAS detection as suggested by the SGNN framework. By bridging gaps in predictive modeling and data availability, this integrated approach provides a strong foundation for accelerating advancements in FET sensor design and innovation.

Ferreira, Rodrigo Pires [Univ. of Chicago, IL (Uni↗

Spin-Controllable Dynamics in Defect-Engineered Carbon Nanotubes as Single Photon Emitters: Data-Driven Modeling and Computations

Quantum technologies, such as quantum computing and sensing, require efficient single-photon emission (SPE) sources that operate at room temperature in telecom wavelengths. While several materials can serve as SPE sources, no single platform meets all the criteria for efficiency, ambient operation, and scalability. Single-walled carbon nanotubes (SWCNTs) with covalently attached molecules offer a promising solution. Their SPE can be easily tuned via modifications of the SWCNT's diameter, chirality, and bonded molecules, enabling emission across near-IR to telecom wavelengths at ambient conditions. However, to fully realize the potential of SWCNTs and unlock their quantum capabilities, a deeper understanding of how structural defects from molecular adducts affect their emission and competing photoexcited processes is essential. To address this gap in our knowledge, this project combined quantum chemistry calculations with data-driven methods of cheminformatics (QSAR) and machine learning (ML). The developed computational approaches have provided several design strategies for covalent functionalization of SWCNTs to improve their optical response. The collaboration with Los Alamos National Lab (LANL) enabled direct comparison of computational and experimental data, facilitating method validation. This partnership was enhanced through access to LANL's Center for Integrated Nanotechnologies (CINT) utilizing User Facility Program and summer internships, which provided three NDSU graduate students with hands-on experience at LANL. The outcomes of this project included (1) Advancing the current stage of computational methods in accurate modeling of non-adiabatic spin-dependent photoexcited dynamics and its applicability to nanosystems consisting of thousands of atoms, realized as open-access codes linked to existing DFT-based software; (2) Establishing the relationship between the structure of adducts and SWCNTs and intrinsic excitonic and spin properties of defect states for guiding novel synthetic strategies and experimental probes of chemically functionalized SWCNTs as near-IR emitting materials; (3) Generating virtual libraries of hypothetical functionalized SWCNTs for virtual screening of their chemical structures and optical properties, leveraging new functionalities of SWCNTs; (4) Offering a unique experience for NDSU graduate students that prepared them for future scientific careers related to materials modeling and big data processing. These results were summarized in 12 published journal papers and 3 recently submitted papers. One of a key finding is that the position of defect sites on the SWCNT surface primarily drives the emission redshift (up to 100 meV), while the polarity of the defect-inducing molecules has a much smaller effect (~10 meV). However, the electron-donating or withdrawing properties of a molecule influence selecting reactivity of defect sites. These insights important for optimizing synthetic protocols for desired emissions in SWCNTs. We also revealed that the interaction between two defects at various positions on the SWCNT enhances the redshift and optical activity of states, favoring strong near-IR emission. This suggests that manipulations in defect concentrations is a promising strategy for controlling efficient emission. Mostly important, the defect position was found controllable by the spin states of photoexcited intermediates: Excited aromatic molecules form ortho defects with SWCNTs at their singlet states in the presence of oxygen, while oxygen-free conditions favor para defects via the triplet-state mechanism. Additionally, a heat-activated [2+2] cycloaddition reaction facilitates divalent defect formation with fewer bonding positions that narrows emission bands. These groundbreaking findings have been experimentally validated and significantly advance our understanding of defect chemistry in SWCNTs. Using a novel encoding technique and 3D-MoRSE descriptors, we developed highly accurate ML/QSAR models to predict both the 3D structure and optical properties of SWCNTs with chemical defects. This model enabled the creation of a virtual library of 125,556 structures, providing new insights into the relationship between SWCNT-defect structure and emission.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Integrated data-driven and experimental approaches to accelerate lead optimization targeting SARS-CoV- 2 main protease

Identification of potential therapeutic candidates can be expedited by integrating computational modeling with domain aware machine learning (ML) approaches followed by experimental validation. Generative deep learning models have been recently developed that can generate thousands of new candidates, but their physiochemical properties are typically not optimized. Using our deep learning models and a scaffold as a starting point, we generated tens of thousands of compounds for SARS-CoV-2 M pro that preserve the core scaffold. Here we utilized and implemented several computational tools such as structural alert and toxicity analysis, high throughput virtual screening, ML-based 3D quantitative structure–activity relationships, multi-parameter optimization, and graph neural networks on libraries of generated candidates to predict biological activity and binding affinity a priori. From these collective computational results, eight promising candidates were identified and tested experimentally using Native Mass Spectrometry (MS) and FRET-based functional assays. Two compounds, with quinazoline-2-thiol and acetylpiperidine core moiety showed IC 50 values in the low micromolar range: 2.95±0.0017 µM and 3.41±0.0015 µM, respectively. The molecular dynamics simulations further highlight that binding of these compounds results in allosteric modulations in the chain B and the interface domains of the M pro . The key fragments from these top hits can be used as input for closed loop lead optimization in the integrated pipeline.

60 APPLIED LIFE SCIENCES↗

Predicting receptor-ligand pairing preferences in plant-microbe interfaces via molecular dynamics and machine learning

Microbiome assembly, structure, and dynamics significantly influence plant health. Secreted microbial signaling molecules initiate and mediate symbiosis by binding to structurally compatible plant receptors. For example, lipo-chitooligosaccharides (LCOs), produced by nitrogen-fixing rhizobial bacteria and various fungi, are recognized by plant lysin motif receptor-like kinases (LysM-RLKs), which activate the common symbiotic pathway. Accurately predicting these molecular interactions could reveal complementary signatures underlying the initial stages of endosymbiosis. Despite the breakthrough in protein-ligand structure prediction with deep learning-based tools, such as AlphaFold3, the large size and highly flexible nature of signaling compounds like LCOs present major challenges for detailed structural characterization and binding-affinity prediction. Typical structure-/physics-based methods of ligand virtual screening are designed for small, drug-like molecules, often rely on high-resolution, experimentally determined structures of the protein receptors, and rarely achieve sufficient sampling to obtain converged thermodynamic quantities with large ligands. In this study, we developed a hybrid molecular dynamics/machine learning (MD/ML) approach capable of predicting binding affinity rankings with high accuracy in systems involving large, flexible ligands, despite limited experimental structural information. Using coarse initial structural models, the predictions using the MD/ML workflow achieved strong alignment with experimental trends, particularly in the top-affinity tier for four legume LysM-RLKs (LYR3) binding to LCOs and a chitooligosaccharide. Furthermore, the MD-based conformation selection protocol provided critical structural insights into substrate specificity and binding mechanisms. This study demonstrates a powerful method to screen for challenging cognate ligand-receptors and advance our understanding of the molecular basis of microbial colonization in plants.

Lipo-chitooligosaccharides↗

Insights into the substrate specificity, structure, and dynamics of plant histidinol-phosphate aminotransferase (HISN6)

Histidinol-phosphate aminotransferase is the sixth protein (hence HISN6) in the histidine biosynthetic pathway in plants. HISN6 is a pyridoxal 5'-phosphate (PLP)-dependent enzyme that catalyzes the reversible conversion of imidazole acetol phosphate into L-histidinol phosphate (HOLP). Here, we show that plant HISN6 enzymes are closely related to the orthologs from Chloroflexota. The studied example, HISN6 from Medicago truncatula (MtHISN6), exhibits a surprisingly high affinity for HOLP, which is much higher than reported for bacterial homologs. Moreover, unlike the latter, MtHISN6 does not transaminate phenylalanine. High-resolution crystal structures of MtHISN6 in the open and closed states, as well as the complex with HOLP and the apo structure without PLP, bring new insights into the enzyme dynamics, pointing at a particular role of a string-like fragment that oscillates near the active site and participates in the HOLP binding. When MtHISN6 is compared to bacterial orthologs with known structures, significant differences arise in or near the string region. The high affinity of MtHISN6 appears linked to the particularly tight active site cavity. Finally, a virtual screening against a library of over 1.3 mln compounds revealed three sites in the MtHISN6 structure with the potential to bind small molecules. Such compounds could be developed into herbicides inhibiting plant HISN6 enzymes absent in animals, which makes them a potential target for weed control agents.

59 BASIC BIOLOGICAL SCIENCES↗

Data-Driven Discovery of Linear Molecular Probes with Optimal Selective Affinity for PFAS in Water

Approaches to tackle the wide and growing variety of highly persistent per- and polyfluoroalkyl substances (PFAS) are of pressing global need because of their detrimental human health effects, such as cancer, birth defects, and hormone imbalance. Sensitive, selective, and easy-to-use real-time sensors to monitor and detect PFAS and sorbents to extract them are critical to meeting government-mandated environmental concentrations. In this work, we combine all-atom molecular dynamics simulations, enhanced sampling, deep representational learning, and Bayesian optimization to perform high-throughput virtual screening for highly sensitive and selective molecular probes. Our molecular design space consists of 3850 linear hydrocarbon chains with varying degrees of halogenation with and without amine- and phosphine-based headgroups. By employing a data-driven search process, we efficiently explore the molecular design space to optimize the sensitivity to perfluorooctanesulfonic acid (PFOS) as a prototypical PFAS analyte and selectivity relative to a sodium dodecyl sulfate (SDS) interferent. We calculate 504 Gibbs free energies of probe-analyte and probe-interferent interactions and identify probes with PFOS association free energies of up to (-ΔG PFOS ) = 9.8 ± 0.2 kJ/mol and selectivities relative to SDS of (-ΔΔG PFOS–SDS ) = 3.1 ± 1.5 kJ/mol. A C 11 Br 23 P(CH 3 ) 2 probe containing 11 backbone brominated carbons and a tertiary phosphine headgroup possesses the most sensitive binding constant to PFOS within the defined search space of K b PFOS = 177.4 ± 12.7, and a semibrominated probe C 5 H 11 C 7 Br 14 N(CH 3 ) 2 containing 12 backbone carbons and a tertiary amine headgroup possesses the highest selectivity relative to SDS of K b PFOS /K b SDS = 4.6 ± 1.7. A retrospective analysis of our data to extract interpretable design rules reveals that the sensitivity of linear hydrogenated probes increases by approximately 1 kJ/mol per C–C bond. The addition or removal of halogen atoms and amine or phosphine headgroups produces nonmonotonic changes in both sensitivity and selectivity with changes to the sensitivity of up to 2.5 kJ/mol. Finally, this work places empirical limitations on the performance of a wide range of linear probes for PFOS detection and offers a generic strategy for high-throughput computational screening to promote selective and sensitive binding.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

CoarsenConf: Equivariant Coarsening with Aggregated Attention for Molecular Conformer Generation

Molecular conformer generation (MCG) is an important task in cheminformatics and drug discovery. The ability to efficiently generate low-energy 3D structures can avoid expensive quantum mechanical simulations, leading to accelerated virtual screenings and enhanced structural exploration. Several generative models have been developed for MCG, but many struggle to consistently produce high-quality conformers for meaningful downstream applications. To address these issues, we introduce CoarsenConf, which coarse-grains molecular graphs based on torsional angles and integrates them into an SE(3)-equivariant hierarchical variational autoencoder. Through equivariant coarse-graining, we aggregate the fine-grained atomic coordinates of subgraphs connected via rotatable bonds, creating a variable-length coarse-grained latent representation. Our model uses a novel aggregated attention mechanism to restore fine-grained coordinates from the coarse-grained latent representation, enabling efficient generation of accurate conformers. Furthermore, we evaluate the chemical and biochemical quality of our generated conformers on multiple downstream applications, including property prediction and large-scale oracle-based protein docking. Overall, CoarsenConf generates more accurate conformer ensembles compared to prior generative models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Discovering High Entropy Alloy Electrocatalysts in Vast Composition Spaces with Multiobjective Optimization

High entropy alloys (HEAs) are a highly promising class of materials for electrocatalysis as their unique active site distributions break the scaling relations that limit the activity of conventional transition metal catalysts. Existing Bayesian optimization (BO)-based virtual screening approaches focus on catalytic activity as the sole objective and correspondingly tend to identify promising materials that are unlikely to be entropically stabilized. Here, we overcome this limitation with a multiobjective BO framework for HEAs that simultaneously targets activity, cost-effectiveness, and entropic stabilization. With diversity-guided batch selection further boosting its data efficiency, the framework readily identifies numerous promising candidates for the oxygen reduction reaction that strike the balance between all three objectives in hitherto unchartered HEA design spaces comprising up to 10 elements.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Structure and Synthesizability of Iron–Sulfur Metal–Organic Frameworks

Sulfur-based metal–organic frameworks (MOFs) and coordination polymers (CPs) are an emerging class of hybrid materials that have received growing attention due to their magnetic, conductive, and catalytic properties with potential applications in electrocatalysis and energy storage. In this work, we report a high-throughput virtual screening protocol to predict the synthesizability of candidate metal–sulfur MOFs/CPs by computing the thermodynamically stable structures resulting from a particular combination of metal cluster, linker, cation, and synthetic conditions. Free energies are computed by using all-atom classical mechanical thermodynamic integration. Low-free-energy structures are refined using ab initio density functional theory, and pair distribution functions and powder X-ray diffraction patterns are calculated to complement and guide experimental structure determination. We validate the computational approach by retrospective predictions of the stable structure produced by experimental syntheses, and a subsequent screen predicts Fe 4 S 4 -BDT–TPP as a new thermodynamically stable one-dimensional (1D) CP comprising a redox-active Fe 4 S 4 cluster, a 1,4-benzenedithiolate (BDT) linker, and a tetraphenylphosphonium (TPP) countercation. Furthermore, this material is experimentally synthesized, and the 1D chain structure of the crystal is confirmed using microcrystal electron diffraction. The computational screening pipeline is generically transferable to neutral and ionic MOFs/CPs comprising arbitrary metal clusters, linkers, cations, and synthetic conditions, and we make it freely available as an open source tool to guide and accelerate the discovery and engineering of novel porous materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗