Search NASA⌕ Search

SEARCH · Search NASA

Results for “high-throughput virtual screening”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Optimal decision-making in high-throughput virtual screening pipelines

Screening large pools of molecular candidates to identify those with specific design criteria or targeted properties is demanding in various science and engineering domains. While a high-throughput virtual screening (HTVS) pipeline can provide efficient means to achieving this goal, its design and operation often rely on experts' intuition, potentially resulting in suboptimal performance. In this paper, we fill this critical gap by presenting a systematic framework that can maximize the return on computational investment (ROCI) of such HTVS campaigns. Based on various scenarios, we empirically validate the proposed framework and demonstrate its potential to accelerate scientific discoveries through optimal computational campaigns, especially in the context of virtual screening.

97 MATHEMATICS AND COMPUTING↗

Optimal high-throughput virtual screening pipeline for efficient selection of redox-active organic materials

As global interest in renewable energy continues to increase, there has been a pressing need for developing novel energy storage devices based on organic electrode materials that can overcome the shortcomings of the current lithium-ion batteries. One critical challenge for this quest is to find materials whose redox potential (RP) meets specific design targets. In this study, we propose a computational framework for addressing this challenge through the effective design and optimal operation of a high-throughput virtual screening (HTVS) pipeline that enables rapid screening of organic materials that satisfy the desired criteria. Starting from a high-fidelity model for estimating the RP of a given material, we show how a set of surrogate models with different accuracy and complexity may be designed to construct a highly accurate and efficient HTVS pipeline. We demonstrate that the proposed HTVS pipeline construction and operation strategies substantially enhance the overall screening throughput.

36 MATERIALS SCIENCE↗

Advancing energy storage through solubility prediction: leveraging the potential of deep learning

Solubility prediction plays a crucial role in energy storage applications, such as redox flow batteries, because it directly affects the efficiency and reliability. Researchers have developed various methods that utilize quantum calculations and descriptors to predict the aqueous solubilities of organic molecules. Notably, machine learning models based on descriptors have shown promise for solubility prediction. As deep learning tools, graph neural networks (GNNs) have emerged to capture complex structure–property relationships for material property prediction. Specifically, MolGAT, a type of GNN model, was designed to incorporate n-dimensional edge attributes, enabling the modeling of intricacies in molecular graphs and enhancing the prediction capabilities. In a previous study, MolGAT successfully screened 23 467 promising redox-active molecules from a database of over 500 000 compounds, based on redox potential predictions. This study focused on applying the MolGAT model to predict the aqueous solubility (log S) of a broad range of organic compounds, including those previously screened for redox activity. The model was trained on a diverse sample of 8494 organic molecules from AqSolDB and benchmarked against literature data, demonstrating superior accuracy compared with other state of the art graph-based and descriptor-based models. Subsequently, the trained MolGAT model was employed to screen redox-active organic compounds identified in the first phase of high-throughput virtual screening, targeting favorable solubility in energy storage applications. The second round of screening, which considered solubility, yielded 12 332 promising redox-active and soluble organic molecules suitable for use in aqueous redox flow batteries. Thus, the two-phase high-throughput virtual screening approach utilizing MolGAT, specifically trained for redox potential and solubility, is an effective strategy for selecting suitable intrinsically soluble redox-active molecules from extensive databases, potentially advancing energy storage through reliable material development. This indicates that the model is reliable for predicting the solubility of various molecules and provides valuable insights for energy storage, pharmaceutical, environmental, and chemical applications.

25 ENERGY STORAGE↗

Pose Classification Using Three-Dimensional Atomic Structure-Based Neural Networks Applied to Ion Channel–Ligand Docking

The identification of promising lead compounds showing pharmacological activities toward a biological target is essential in early stage drug discovery. With the recent increase in available small-molecule databases, virtual high-throughput screening using physics-based molecular docking has emerged as an essential tool in assisting fast and cost-efficient lead discovery and optimization. However, the best scored docking poses are often suboptimal, resulting in incorrect screening and chemical property calculation. We address the pose classification problem by leveraging data-driven machine learning approaches to identify correct docking poses from AutoDock Vina and Glide screens. To enable effective classification of docking poses, we present two convolutional neural network approaches: a three-dimensional convolutional neural network (3D-CNN) and an attention-based point cloud network (PCN) trained on the PDBbind refined set. We demonstrate the effectiveness of our proposed classifiers on multiple evaluation data sets including the standard PDBbind CASF-2016 benchmark data set and various compound libraries with structurally different protein targets including an ion channel data set extracted from Protein Data Bank (PDB) and an in-house KCa3.1 inhibitor data set. Our experiments show that excluding false positive docking poses using the proposed classifiers improves virtual high-throughput screening to identify novel molecules against each target protein compared to the initial screen based on the docking scores.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data-Driven Discovery of Linear Molecular Probes with Optimal Selective Affinity for PFAS in Water

Approaches to tackle the wide and growing variety of highly persistent per- and polyfluoroalkyl substances (PFAS) are of pressing global need because of their detrimental human health effects, such as cancer, birth defects, and hormone imbalance. Sensitive, selective, and easy-to-use real-time sensors to monitor and detect PFAS and sorbents to extract them are critical to meeting government-mandated environmental concentrations. In this work, we combine all-atom molecular dynamics simulations, enhanced sampling, deep representational learning, and Bayesian optimization to perform high-throughput virtual screening for highly sensitive and selective molecular probes. Our molecular design space consists of 3850 linear hydrocarbon chains with varying degrees of halogenation with and without amine- and phosphine-based headgroups. By employing a data-driven search process, we efficiently explore the molecular design space to optimize the sensitivity to perfluorooctanesulfonic acid (PFOS) as a prototypical PFAS analyte and selectivity relative to a sodium dodecyl sulfate (SDS) interferent. We calculate 504 Gibbs free energies of probe-analyte and probe-interferent interactions and identify probes with PFOS association free energies of up to (-ΔG PFOS ) = 9.8 ± 0.2 kJ/mol and selectivities relative to SDS of (-ΔΔG PFOS–SDS ) = 3.1 ± 1.5 kJ/mol. A C 11 Br 23 P(CH 3 ) 2 probe containing 11 backbone brominated carbons and a tertiary phosphine headgroup possesses the most sensitive binding constant to PFOS within the defined search space of K b PFOS = 177.4 ± 12.7, and a semibrominated probe C 5 H 11 C 7 Br 14 N(CH 3 ) 2 containing 12 backbone carbons and a tertiary amine headgroup possesses the highest selectivity relative to SDS of K b PFOS /K b SDS = 4.6 ± 1.7. A retrospective analysis of our data to extract interpretable design rules reveals that the sensitivity of linear hydrogenated probes increases by approximately 1 kJ/mol per C–C bond. The addition or removal of halogen atoms and amine or phosphine headgroups produces nonmonotonic changes in both sensitivity and selectivity with changes to the sensitivity of up to 2.5 kJ/mol. Finally, this work places empirical limitations on the performance of a wide range of linear probes for PFOS detection and offers a generic strategy for high-throughput computational screening to promote selective and sensitive binding.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Comparative Assessment of Pose Prediction Accuracy in RNA–Ligand Docking

Structure-based virtual high-throughput screening is used in early-stage drug discovery. Over the years, docking protocols and scoring functions for protein–ligand complexes have evolved to improve the accuracy in the computation of binding strengths and poses. In the past decade, RNA has also emerged as a target class for new small-molecule drugs. However, most ligand docking programs have been validated and tested for proteins and not RNA. Here, we test the docking power (pose prediction accuracy) of three state-of-the-art docking protocols on 173 RNA–small molecule crystal structures. The programs are AutoDock4 (AD4) and AutoDock Vina (Vina), which were designed for protein targets, and rDock, which was designed for both protein and nucleic acid targets. AD4 performed relatively poorly. For RNA targets for which a crystal structure of a bound ligand used to limit the docking search space is available and for which the goal is to identify new molecules for the same pocket, rDock performs slightly better than Vina, with success rates of 48% and 63%, respectively. However, in the more common type of early-stage drug discovery setting, in which no structure of a ligand–target complex is known and for which a larger search space is defined, rDock performed similarly to Vina, with a low success rate of ~27%. Further, Vina was found to have bias for ligands with certain physicochemical properties, whereas rDock performs similarly for all ligand properties. Thus, for projects where no ligand–protein structure already exists, Vina and rDock are both applicable. However, the relatively poor performance of all methods relative to protein–target docking illustrates a need for further methods refinement.

59 BASIC BIOLOGICAL SCIENCES↗

Structure and Synthesizability of Iron–Sulfur Metal–Organic Frameworks

Sulfur-based metal–organic frameworks (MOFs) and coordination polymers (CPs) are an emerging class of hybrid materials that have received growing attention due to their magnetic, conductive, and catalytic properties with potential applications in electrocatalysis and energy storage. In this work, we report a high-throughput virtual screening protocol to predict the synthesizability of candidate metal–sulfur MOFs/CPs by computing the thermodynamically stable structures resulting from a particular combination of metal cluster, linker, cation, and synthetic conditions. Free energies are computed by using all-atom classical mechanical thermodynamic integration. Low-free-energy structures are refined using ab initio density functional theory, and pair distribution functions and powder X-ray diffraction patterns are calculated to complement and guide experimental structure determination. We validate the computational approach by retrospective predictions of the stable structure produced by experimental syntheses, and a subsequent screen predicts Fe 4 S 4 -BDT–TPP as a new thermodynamically stable one-dimensional (1D) CP comprising a redox-active Fe 4 S 4 cluster, a 1,4-benzenedithiolate (BDT) linker, and a tetraphenylphosphonium (TPP) countercation. Furthermore, this material is experimentally synthesized, and the 1D chain structure of the crystal is confirmed using microcrystal electron diffraction. The computational screening pipeline is generically transferable to neutral and ionic MOFs/CPs comprising arbitrary metal clusters, linkers, cations, and synthetic conditions, and we make it freely available as an open source tool to guide and accelerate the discovery and engineering of novel porous materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Active learning of polarizable nanoparticle phase diagrams for the guided design of triggerable self-assembling superlattices

Polarizable nanoparticles are of interest in materials science because of their rich and complex phase behavior that can be used to engineer nanostructured materials with long-range crystalline order. To understand and rationally navigate the design space of polarizable nanoparticles for self-assembling highly ordered superlattices, we developed a coarse-grained computational model to describe the nanoparticle-nanoparticle interactions in implicit solvent and employ the computationally efficient image method to model many-body polarization interactions. We conducted high-throughput virtual screening over a five-dimensional particle design space spanned by temperature, particle size, particle charge, particle dielectric, and solvent dielectric using enhanced sampling molecular dynamics calculations within an active learning framework to efficiently map out the regions of thermodynamic stability of the self-assembled aggregates. We validate our predictions in comparisons against small angle x-ray scattering measurements of gold nanoparticles surface functionalized with metal chalcogenide ligands. Lastly, we use our validated phase maps to computationally design switchable nanostructured materials capable of triggered assembly and disassembly as a function of temperature and solvent dielectric with potential applications as sensors, smart windows, optoelectronic devices, and in medical diagnostics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Detection of multi-reference character imbalances enables a transfer learning approach for virtual high throughput screening with coupled cluster accuracy at DFT cost

Appropriately identifying and treating molecules and materials with significant multi-reference (MR) character is crucial for achieving high data fidelity in virtual high-throughput screening (VHTS). Despite development of numerous MR diagnostics, the extent to which a single value of such a diagnostic indicates the MR effect on a chemical property prediction is not well established. We evaluate MR diagnostics for over 10 000 transition-metal complexes (TMCs) and compare to those for organic molecules. We observe that only some MR diagnostics are transferable from one chemical space to another. By studying the influence of MR character on chemical properties (i.e., MR effect) that involve multiple potential energy surfaces (i.e., adiabatic spin splitting, ΔE H–L , and ionization potential, IP), we show that differences in MR character are more important than the cumulative degree of MR character in predicting the magnitude of an MR effect. Motivated by this observation, we build transfer learning models to predict CCSD(T)-level adiabatic ΔE H–L and IP from lower levels of theory. By combining these models with uncertainty quantification and multi-level modeling, we introduce a multi-pronged strategy that accelerates data acquisition by at least a factor of three while achieving coupled cluster accuracy (i.e., to within 1 kcal mol –1 MAE) for robust VHTS.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Many-body expansion based machine learning models for octahedral transition metal complexes

Abstract Graph-based machine learning (ML) models for material properties show great potential to accelerate virtual high-throughput screening of large chemical spaces. However, in their simplest forms, graph-based models do not include any 3D information and are unable to distinguish stereoisomers such as those arising from different orderings of ligands around a metal center in coordination complexes. In this work we present a modification to revised autocorrelation descriptors, a molecular graph featurization method, for predicting spin state dependent properties of octahedral transition metal complexes (TMCs). Inspired by analytical semi-empirical models for TMCs, the new modeling strategy is based on the many-body expansion (MBE) and allows one to tune the captured stereoisomer information by changing the truncation order of the MBE. We present the necessary modifications to include this approach in two commonly used ML methods, kernel ridge regression and feed-forward neural networks. On a test set composed of all possible isomers of binary TMCs, the best MBE models achieve mean absolute errors (MAEs) of 2.75 kcal mol −1 on spin-splitting energies and 0.26 eV on frontier orbital energy gaps, a 30%–40% reduction in error compared to models based on our previous approach. We also observe improved generalization to previously unseen ligands where the best-performing models exhibit MAEs of 4.00 kcal mol −1 (i.e. a 0.73 kcal mol −1 reduction) on the spin-splitting energies and 0.53 eV (i.e. a 0.10 eV reduction) on the frontier orbital energy gaps. Because the new approach incorporates insights from electronic structure theory, such as ligand additivity relationships, these models exhibit systematic generalization from homoleptic to heteroleptic complexes, allowing for efficient screening of TMC search spaces.

Meyer, Ralf (ORCID:0000000322360261)↗

Discovery of an autoinhibited conformation in mesotrypsin reveals a strategy for selective serine protease inhibition

Selective inhibition of the more than 100 S1 family serine proteases is a long-standing challenge due to their active site similarity. Mesotrypsin, implicated in cancer progression, exemplifies these difficulties; no current inhibitors achieve selectivity over other human trypsins. We found an unexpected autoinhibited conformation of mesotrypsin via x-ray crystallography, revealing a cryptic pocket adjacent to the active site. Using high-throughput virtual screening targeting this cryptic pocket, we identified a conformationally selective small-molecule inhibitor that stabilizes the inactive state of mesotrypsin. This inhibitor demonstrates selectivity for mesotrypsin over other trypsins. Our findings challenge the accepted view of digestive trypsins as constitutively active enzymes lacking potential for allosteric regulation. Furthermore, analyses of other structures suggest that dynamic sampling of closed states with analogous allosteric cryptic pockets appears widespread among S1 serine proteases. These observations point to a potentially generalizable strategy to achieve selective inhibition, offering broad implications for drug development targeting serine proteases in cancer and other diseases.

Coban, Matt↗

Deep Generative Models for Materials Discovery and Machine Learning-Accelerated Innovation

Machine learning and artificial intelligence (AI/ML) methods are beginning to have significant impact in chemistry and condensed matter physics. For example, deep learning methods have demonstrated new capabilities for high-throughput virtual screening, and global optimization approaches for inverse design of materials. Recently, a relatively new branch of AI/ML, deep generative models (GMs), provide additional promise as they encode material structure and/or properties into a latent space, and through exploration and manipulation of the latent space can generate new materials. These approaches learn representations of a material structure and its corresponding chemistry or physics to accelerate materials discovery, which differs from traditional AI/ML methods that use statistical and combinatorial screening of existing materials via distinct structure-property relationships. However, application of GMs to inorganic materials has been notably harder than organic molecules because inorganic structure is often more complex to encode. In this work we review recent innovations that have enabled GMs to accelerate inorganic materials discovery. We focus on different representations of material structure, their impact on inverse design strategies using variational autoencoders or generative adversarial networks, and highlight the potential of these approaches for discovering materials with targeted properties needed for technological innovation.

36 MATERIALS SCIENCE↗

Protein-ligand binding affinity prediction using multi-instance learning with docking structures

Recent advances in 3D structure-based deep learning approaches demonstrate improved accuracy in predicting protein-ligand binding affinity in drug discovery. These methods complement physics-based computational modeling such as molecular docking for virtual high-throughput screening. Despite recent advances and improved predictive performance, most methods in this category primarily rely on utilizing co-crystal complex structures and experimentally measured binding affinities as both input and output data for model training. Nevertheless, co-crystal complex structures are not readily available and the inaccurate predicted structures from molecular docking can degrade the accuracy of the machine learning methods. We introduce a novel structure-based inference method utilizing multiple molecular docking poses for each complex entity. Our proposed method employs multi-instance learning with an attention network to predict binding affinity from a collection of docking poses. We validate our method using multiple datasets, including PDBbind and compounds targeting the main protease of SARS-CoV-2. The results demonstrate that our method leveraging docking poses is competitive with other state-of-the-art inference models that depend on co-crystal structures. This method offers binding affinity prediction without requiring co-crystal structures, thereby increasing its applicability to protein targets lacking such data.

97 MATHEMATICS AND COMPUTING↗

Exploration of structure-activity relationships for the SARS-CoV-2 macrodomain from shape-based fragment linking and active learning

The macrodomain of severe acute respiratory syndrome coronavirus 2 nonstructural protein 3 is required for viral pathogenesis and is an emerging antiviral target. We previously performed an x-ray crystallography–based fragment screen and found submicromolar inhibitors by fragment linking. However, these compounds had poor membrane permeability and liabilities that complicated optimization. Here, we developed a shape-based virtual screening pipeline—FrankenROCS. We screened the Enamine high-throughput collection of 2.1 million compounds, selecting 39 compounds for testing, with the most potent binding with a 130 μM median inhibitory concentration (IC 50 ). We then paired FrankenROCS with an active learning algorithm (Thompson sampling) to efficiently search the Enamine REAL database of 22 billion molecules, testing 32 compounds with the most potent binding with a 220 μM IC 50 . Further optimization led to analogs with IC 50 values better than 10 μM. This lead series has improved membrane permeability and is poised for optimization. FrankenROCS is a scalable method for fragment linking to exploit synthesis-on-demand libraries.

Science & Technology - Other Topics↗

Drugsniffer: An Open Source Workflow for Virtually Screening Billions of Molecules for Binding Affinity to Protein Targets

The SARS-CoV2 pandemic has highlighted the importance of efficient and effective methods for identification of therapeutic drugs, and in particular has laid bare the need for methods that allow exploration of the full diversity of synthesizable small molecules. While classical high-throughput screening methods may consider up to millions of molecules, virtual screening methods hold the promise of enabling appraisal of billions of candidate molecules, thus expanding the search space while concurrently reducing costs and speeding discovery. Here, we describe a new screening pipeline, called drugsniffer, that is capable of rapidly exploring drug candidates from a library of billions of molecules, and is designed to support distributed computation on cluster and cloud resources. As an example of performance, our pipeline required ~40,000 total compute hours to screen for potential drugs targeting three SARS-CoV2 proteins among a library of ~3.7 billion candidate molecules.

59 BASIC BIOLOGICAL SCIENCES↗

High-throughput spin-bath characterization of spin defects in semiconductors

Detailed knowledge of the local environments of spin defects in semiconductors, such as nitrogenvacancy (NV) centers in diamond or divacancies in silicon carbide, is crucial for optimizing control and entanglement protocols in quantum sensing and information applications. However, at present a direct experimental characterization of individual defect environments is not scalable, as conventional spin-bath measurements are time consuming and difficult to automate. Achieving high-throughput characterization requires short experiments to probe the spin bath. However, with fewer and noisier measurements, the inverse problem of recovering spin-bath properties from measured data becomes ill posed, with multiple spin baths having a high likelihood of yielding the same data. In this work, we present a set of computational tools to resolve the ill-posed inverse problem of recovering the atomic positions and hyperfine couplings of random nuclei surrounding spin defects from sparse, noisy experimental coherence data, which can be obtained in hours. Here, we use a trans-dimensional Bayesian approach that incorporates ab initio data to yield full posterior distributions over nuclear spin environments, enabling robust recovery from limited data. We also provide practical tools and guidelines to determine the limits of detectability for hyperfine couplings under specific dynamical decoupling sequences and sampling conditions. In addition, we demonstrate how the tools developed here, in combination with ab initio simulations of spin baths, can guide the design of efficient experimental protocols for application-specific high-throughput screening. To showcase the utility of our approach, we apply it to design fast dynamical decoupling experiments to characterize the spin baths often individual NV centers in diamond. While the primary focus is on accelerating spin-bath characterization of spin defects, this Bayesian approach also lays the foundation for digital-twin studies of spin defects, where a virtual model of the spin-defect system evolves in real time with ongoing experimental measurements. Together, the set of tools we designed and applied paves the way for scalable deployment of spin defects in semiconductors for quantum sensing and information applications.

Bayesian methods↗

NGPINT V3: a containerized orchestration Python software for discovery of next-generation protein–protein interactions

Abstract Summary Batch yeast two-hybrid (Y2H) assays, leveraged with next-generation sequencing, have afforded successful innovations for the analysis of protein–protein interactions. NGPINT is a Conda-based software designed to process the millions of raw sequencing reads resulting from Y2H–next-generation interaction screens. Over time, increasing compatibility and dependency issues have prevented clean NGPINT installation and operation. A system-wide update was essential to continue effective use with its companion software, Y2H-SCORES. We present NGPINT V3, a containerized implementation built with both Singularity and Docker, allowing accessibility across virtually any operating system and computing environment. Availability and implementation This update includes streamlined dependencies and container images hosted on Sylabs (https://cloud.sylabs.io/library/schuyler/ngpint/ngpint) and Dockerhub (https://hub.docker.com/r/schuylerds/ngpint), facilitating easier adoption and integration into high-throughput and cloud-computing workflows. Full instructions and software can be also found in the GitHub repository https://github.com/Wiselab2/NGPINT_V3 and Zenodo https://doi.org/10.5281/zenodo.15256036.

Biochemistry & Molecular Biology↗

Machine learning prediction on the fractional free volume of polymer membranes

Fractional free volume (FFV) characterizes the microstructural level features of polymers and affects their properties including thermal, mechanical, and separation performance. Experimental measurements and theoretical analyses have been used to quantify the FFV of polymers, but challenges remain because of their limitations. Experimental measurements are laborious and based on semi empirical equations, while Bondi’s group contribution theory involves ambiguities like the determination of van der Waals volume and the choice of factor values in the theoretical equation. To efficiently evaluate the FFV of polymers, this study utilizes high-throughput molecular dynamics (MD) simulations to build a large dataset regarding polymer’s FFV. Based on this large dataset, we further build machine learning (ML) models to establish the composition-structure relation. Inspired by group contribution theory which correlates polymer’s functional groups to FFV, our ML models correlate polymer’s substructures or physico-chemical indexes to FFV. Here, our study first benchmarks the MD simulation protocol to obtain reliable FFV of polymers and then carries out high-throughput MD simulations for more than 6,500 homopolymers and 1,400 polyamides. Such a large and diverse dataset makes the well-trained ML models more generalizable, compared with the group contribution theory. The efficiency of a feed forward neural network model is further demonstrated by applying it to a hypothetical polyimide dataset of more than 8 million chemical structures. The predicted FFVs of hypothetical polyimides are further validated by MD simulations. The obtained FFVs of the 8 million polymers, plus their previously reported gas separation performances, demonstrate the promising capability of ML virtual screening for the discovery of polymer membranes with exceptional permeability/selectivity.

36 MATERIALS SCIENCE↗