Search NASA⌕ Search

SEARCH · Search NASA

Results for “Molecular fingerprints”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Do Molecular Fingerprints Identify Diverse Active Drugs in Large-Scale Virtual Screening? (No)

Computational approaches for small-molecule drug discovery now regularly scale to the consideration of libraries containing billions of candidate small molecules. One promising approach to increased the speed of evaluating billion-molecule libraries is to develop succinct representations of each molecule that enable the rapid identification of molecules with similar properties. Molecular fingerprints are thought to provide a mechanism for producing such representations. Here, we explore the utility of commonly used fingerprints in the context of predicting similar molecular activity. We show that fingerprint similarity provides little discriminative power between active and inactive molecules for a target protein based on a known active—while they may sometimes provide some enrichment for active molecules in a drug screen, a screened data set will still be dominated by inactive molecules. We also demonstrate that high-similarity actives appear to share a scaffold with the query active, meaning that they could more easily be identified by structural enumeration. Furthermore, even when limited to only active molecules, fingerprint similarity values do not correlate with compound potency. In sum, these results highlight the need for a new wave of molecular representations that will improve the capacity to detect biologically active molecules based on their similarity to other such molecules.

59 BASIC BIOLOGICAL SCIENCES↗

Molecular fingerprint and machine learning to accelerate design of high-performance homochiral metal–organic frameworks

In this report computational screening was employed to calculate the enantioseparation capabilities of 45 functionalized homochiral metal–organic frameworks (FHMOFs), and machine learning (ML) and molecular fingerprint (MF) techniques were used to find new FHMOFs with high performance. With increasing temperature, the enantioselectivities for (R,S)-1,3-dimethyl-1,2-propadiene are improved. The “glove effect” in the chiral pockets was proposed to explain the correlations between the steric effect of functional groups and performance of FHMOFs. Moreover, the neighborhood component analysis and RDKit/MACCS MFs show the highest predictive effect on enantioselectivities among the four ML classification algorithms with nine MFs that were tested. Based on the importance of MF, 85 new FHMOFs were designed, and a newly designed FHMOF, NO 2 -NHOH-FHMOF, with high similarity to the optimal MFs achieved improved chiral separation performance, with enantioselectivities of 85%. The design principles and new chiral pockets obtained by ML and MFs could facilitate the development of new materials for chiral separation.

molecular fingerprint↗

Tabletop soft x-ray absorption spectroscopy for molecular fingerprinting

For applications related to nuclear security, safeguards, and nonproliferation, it is often critical to know the molecular compositions of lanthanide- and actinide-containing samples. Spectroscopy is a widely used tool that looks at the interaction between light and matter: Different species absorb or emit light at unique wavelengths which act as signatures. However, there is a limited number of tools that can achieve high-sensitivity, accurate measurements of lanthanide and actinide molecular compositions. Candidate methods include mass spectrometry, which usually destroys at least part of the sample and requires complicated stoichiometry to guess the original sample’s molecular compositions; optical spectroscopies, which have great atomic but limited molecular sensitivities or other drawbacks which make sensing molecules difficult like limited light sources or strong absorption in the atmosphere; and nuclear spectroscopies (gamma, neutron) which also have limited sources and long (>minute) collection times. As such, the purpose of our research is to develop a new tool to better distinguish between subtle differences in molecules containing lanthanides and actinides. Soft x-ray spectroscopy is sensitive to molecular form and is minimally intrusive/nondestructive to the sample. However, soft x-ray light with sufficient brightness for spectroscopy is typically limited to user-facilities like synchrotrons or free electron lasers, where beamtimes are competitive, and work with radiological materials may be difficult or entirely prohibited. To overcome this issue, our Team has developed a custom tabletop laser-driven, soft x-ray light source which employs high harmonic generation (HHG). Soft x-ray spectroscopy can distinguish between subtly different molecules, in the spectral range which we need to study these heavy elements. A tabletop system provides an effective and affordable tool to find both the elemental and chemical specificity of lanthanide and samples. Creating a light source in the soft x-ray spectrum is difficult because these wavelengths in the range 5-20 nm (20-350 eV photon energies) only penetrate several 100s of nm in most solid materials and only reflect well in shallow, grazing incident angles. The results are applicable to nuclear forensics, because molecular fingerprinting of lanthanide and actinide samples can be used to back out the origin and processing method of nuclear materials (Skrodzki, et al.).

36 MATERIALS SCIENCE↗

J -Resolved Molecular Fingerprinting by Parahydrogen Hyperpolarized Low-Field NMR

A J-resolved spectroscopy that depends on homonuclear scalar coupling in the strong-coupling regime and heteronuclear coupling in the weak regime expands complex peak patterns to a second axis. Hyperpolarization by Signal Amplification by Reversible Exchange (SABRE) enables the spectroscopy at a low magnetic field of 0.82 mT. Overlapping peaks of molecules such as 3-fluoropyridine and 3,5-difluoropyridine are resolved. Density matrix simulations of the 1 H and 19 F spins indicate a strong dependence on the signs and values of the J-coupling constants, including the homonuclear couplings that are not directly observable. The best matching peak positions and intensities predict coupling constants, including couplings between chemically equivalent nuclear spins, ranging in magnitude from 0.4 to 9.0 Hz for the two molecules. Simulations of other spin systems show unique patterns for molecules containing 1 H and 19 F or 13 C. The dependence of the J-resolved peak patterns on all coupling constants in a spin system presents a new modality for portable and inexpensive identification of molecules.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Invariant Molecular Representations for Heterogeneous Catalysis

Catalyst screening is a critical step in the discovery and development of heterogeneous catalysts, which are vital for a wide range of chemical processes. In recent years, computational catalyst screening, primarily through density functional theory (DFT), has gained significant attention as a method for identifying promising catalysts. However, the computation of adsorption energies for all likely chemical intermediates present in complex surface chemistries is computationally intensive and costly due to the expensive nature of these calculations and the intrinsic idiosyncrasies of the methods or data sets used. This study introduces a novel machine learning (ML) method to learn adsorption energies from multiple DFT functionals by using invariant molecular representations (IMRs). To do this, we first extract molecular fingerprints for the reaction intermediates and later use a Siamese-neural-network-based training strategy to learn invariant molecular representations or the IMR across all available functionals. Our Siamese network-based representations demonstrate superior performance in predicting adsorption energies compared with other molecular representations. Notably, when considering mean absolute values of adsorption energies as 0.43 eV (PBE-D3), 0.46 eV (BEEF-vdW), 0.81 eV (RPBE), and 0.37 eV (scan+rVV10), our IMR method has achieved the lowest mean absolute errors (MAEs) of 0.18 0.10, 0.16, and 0.18 eV, respectively. These results emphasize the superior predictive capacity of our Siamese network-based representations. The empirical findings in this study illuminate the efficacy, robustness, and dependability of our proposed ML paradigm in predicting adsorption energies, specifically for propane dehydrogenation on a platinum catalyst surface.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Fingerprinting Interactions between Proteins and Ligands for Facilitating Machine Learning in Drug Discovery

Molecular recognition is fundamental in biology, underpinning intricate processes through specific protein–ligand interactions. This understanding is pivotal in drug discovery, yet traditional experimental methods face limitations in exploring the vast chemical space. Computational approaches, notably quantitative structure–activity/property relationship analysis, have gained prominence. Molecular fingerprints encode molecular structures and serve as property profiles, which are essential in drug discovery. While two-dimensional (2D) fingerprints are commonly used, three-dimensional (3D) structural interaction fingerprints offer enhanced structural features specific to target proteins. Machine learning models trained on interaction fingerprints enable precise binding prediction. Recent focus has shifted to structure-based predictive modeling, with machine-learning scoring functions excelling due to feature engineering guided by key interactions. Notably, 3D interaction fingerprints are gaining ground due to their robustness. Various structural interaction fingerprints have been developed and used in drug discovery, each with unique capabilities. This review recapitulates the developed structural interaction fingerprints and provides two case studies to illustrate the power of interaction fingerprint-driven machine learning. The first elucidates structure–activity relationships in β2 adrenoceptor ligands, demonstrating the ability to differentiate agonists and antagonists. The second employs a retrosynthesis-based pre-trained molecular representation to predict protein–ligand dissociation rates, offering insights into binding kinetics. Despite remarkable progress, challenges persist in interpreting complex machine learning models built on 3D fingerprints, emphasizing the need for strategies to make predictions interpretable. Binding site plasticity and induced fit effects pose additional complexities. Interaction fingerprints are promising but require continued research to harness their full potential.

3D structural interaction fingerprints↗

Phonon engineering of boron nitride via isotopic enrichment

Phonon polaritons (PhPs) enable a variety of applications, yet it requires PhPs supported in the desired frequency range. The two BN allotropes (cubic and hexagonal, cBN and hBN) are of particular interest, as their optic phonons fall within the so-called molecular-fingerprint region (~ 1000–1610 cm -1 ). However, there remains a spectral gap between PhPs covered by these two, limiting applications. Thus, we isotopically engineered hBN and cBN and examined the optic phonons. Furthermore, for hBN, enhancement of the optic phonon lifetimes and shifted frequencies are observed. However, lifetimes are observed to decrease with the enrichment of cBN by 10 B, 11 B, and 15 N. We propose that the reduced lifetimes are not due to intrinsic loss, but rather increased defect concentrations resulting from the modified growth, supported by first-principles calculations. Thus, reducing the extrinsic defects in isotopically engineered cBN may present a path toward overcoming these restrictions for applications in the molecular-fingerprint region.

36 MATERIALS SCIENCE↗

A Deep Neural Network for Accurate and Robust Prediction of the Glass Transition Temperature of Polyhydroxyalkanoate Homo- and Copolymers

The purpose of this study was to develop a data-driven machine learning model to predict the performance properties of polyhydroxyalkanoates (PHAs), a group of biosourced polyesters featuring excellent performance, to guide future design and synthesis experiments. A deep neural network (DNN) machine learning model was built for predicting the glass transition temperature, Tg, of PHA homo- and copolymers. Molecular fingerprints were used to capture the structural and atomic information of PHA monomers. The other input variables included the molecular weight, the polydispersity index, and the percentage of each monomer in the homo- and copolymers. The results indicate that the DNN model achieves high accuracy in estimation of the glass transition temperature of PHAs. In addition, the symmetry of the DNN model is ensured by incorporating symmetry data in the training process. The DNN model achieved better performance than the support vector machine (SVD), a nonlinear ML model and least absolute shrinkage and selection operator (LASSO), a sparse linear regression model. The relative importance of factors affecting the DNN model prediction were analyzed. Sensitivity of the DNN model, including strategies to deal with missing data, were also investigated. Compared with commonly used machine learning models incorporating quantitative structure–property (QSPR) relationships, it does not require an explicit descriptor selection step but shows a comparable performance. The machine learning model framework can be readily extended to predict other properties.

quantitative structure–property relationship (QSPR↗

AIMSim : An accessible cheminformatics platform for similarity operations on chemicals datasets

The recent advances in deep learning, generative modeling, and statistical learning have ushered in a renewed interest in traditional cheminformatics tools and methods. Quantifying molecular similarity is essential in molecular generative modeling, exploratory molecular synthesis campaigns, and drug-discovery applications to assess how new molecules differ from existing ones. Further, most tools target advanced users and lack general implementations accessible to the larger community. In this work, we introduce Artificial Intelligence Molecular Similarity (AIMSim), an accessible cheminformatics platform for performing similarity operations on collections of molecules called molecular datasets. AIMSim provides a unified platform to perform similarity-based tasks on molecular datasets, such as diversity quantification, outlier and novelty analysis, clustering, dimensionality reduction, and inter-molecular comparisons. AIMSim implements all major binary similarity metrics and molecular fingerprints and is provided as a Python package that includes support for command-line use as well as a Graphical User Interface for code-free utilization with fully interactive plots.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Dynamically tunable membrane metasurfaces for infrared spectroscopy and strong light-matter interactions

Mid-infrared spectroscopy enables biochemical sensing by identifying vibrational molecular fingerprints, but it faces limitations in instrumentation portability and analytical sensitivity. Optical metasurfaces with strong mid-infrared photonic resonances provide an attractive solution towards on-chip spectrometry and sensitive molecular detection, yet their static nature hinders their anticipated impact. Here, we introduce and demonstrate dynamically tunable silicon membrane metasurfaces exhibiting high-Q transmissive resonances in the fingerprint region. By harnessing silicon’s thermo-optical properties, we achieve continuous modulation of coupling-induced transparency (CIT) modes that emerge upon the interference of quasi-bound states in the continuum (q-BICs) and surface lattice modes (SLMs). We measure a spectral tuning rate of 0.06 cm −1 K −1 by continuously sweeping the sharp CIT resonances over a 23.5 cm −1 spectral range across a temperature range of 300–700 K. In the current proof‑of‑concept implementation, the dynamic transmission control enables non-contact chemical analysis of polymer films by detecting characteristic absorption bands of polystyrene (1450 and 1492 cm −1 ) and poly(methyl methacrylate) (1730 cm −1 ) without requiring conventional spectrometers. When analyte molecules fill the metasurface-generated photonic cavities, we demonstrate vibrational strong coupling between the poly(methyl methacrylate)’s carbonyl band and the CIT mode, manifested in a Rabi splitting of ~43 cm −1 . Our results establish a new photonic platform that unites spectral precision, strong field enhancement, and reconfigurability, offering diverse potential for compact mid-infrared spectroscopy, molecular sensing, and programmable polaritonic photonics.

74 ATOMIC AND MOLECULAR PHYSICS↗

Graph neural networks for CO 2 solubility predictions in Deep Eutectic Solvents

Deep Eutectic Solvents (DESs) are a promising class of solvents for CO 2 capture. DESs are complex mixtures that can be designed to optimize CO solubility and overall capture process efficiency. However, the vast design landscape of DES mixtures makes experimental investigation prohibitive; as such, there is a need for computational models that can quickly and efficiently navigate the design space and inform data collection efforts. In this work, we propose Graph Neural Network (GNN) models for predicting CO 2 solubility for DESs; the GNN leverages a mixture graph representation that captures the molecular structure of the DES components as well as their intermolecular interactions. Here, we compare the GNN framework against alternative architectures (neural networks, graph convolution networks, and random forests) and data representations (molecular fingerprints, sigma profiles, and graphs). We show that the proposed approach offers superior predictive performance; specifically, we show that solubility can be predicted reliably directly from molecular structure (without the need of using sigma profiles as proposed in previous studies). This result is important, as obtaining sigma profiles requires expensive density functional theory computations. We also explored the ability of GNNs to predict solubility for new DES mixtures and operating conditions. We found that the model extrapolates across temperature reliably. However, we also found deficiencies in the ability of the model to predict solubility for DES mixtures, pressures, and molar ratio not included in the training sets; we show that this is due to an inherent lack of chemical diversity in datasets available in the literature. The proposed computational capabilities can thus help navigate the design space of DES and inform data collection efforts. Our models, data, and benchmarks are shared as Python code implemented in Jupyter notebooks.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Response of soybean rhizosphere communities to human hygiene water addition as determined by community level physiological profiling (CLPP) and terminal restriction fragment length polymorphism (TRFLP) analysis

In this report, we describe an experiment conducted at Kennedy Space Center in the biomass production chamber (BPC) using soybean plants for purification and processing of human hygiene water. Specifically, we tested whether it was possible to detect changes in the root-associated bacterial assemblage of the plants and ultimately to identify the specific microorganism(s) which differed when plants were exposed to hygiene water and other hydroponic media. Plants were grown in hydroponics media corresponding to four different treatments: control (Hoagland's solution), artificial gray water (Hoagland's+surfactant), filtered gray water collected from human subjects on site, and unfiltered gray water. Differences in rhizosphere microbial populations in all experimental treatments were observed when compared to the control treatment using both community level physiological profiles (BIOLOG) and molecular fingerprinting of 16S rRNA genes by terminal restriction fragment length polymorphism analysis (TRFLP). Furthermore, screening of a clonal library of 16S rRNA genes by TRFLP yielded nearly full length SSU genes associated with the various treatments. Most 16S rRNA genes were affiliated with the Klebsiella, Pseudomonas, Variovorax, Burkholderia, Bordetella and Isosphaera groups. This molecular approach demonstrated the ability to rapidly detect and identify microorganisms unique to experimental treatments and provides a means to fingerprint microbial communities in the biosystems being developed at NASA for optimizing advanced life support operations.

NASA Discipline Life Support Systems↗

Technical Assistance for RingIR Aerosol Penetration Study with RingIR

Sandia provided technical assistance to RingIR to test and evaluate of the RingIR molecular detector. The detector can identify gas phase species using molecular fingerprinting and has potential application for SARS-CoV-2 detection in near real time. As part of the development process Sandia will utilize the biological aerosol test bed deployed at the Aerosol Complex to evaluate the penetration of MS2 bacteriophage aerosol through the RingIR system. The objective of this project is to provide experimentally derived measurements of the RingIR molecular detector penetration efficiency, including external exhaust filter for mitigation of exhaust aerosol and operation using MS2 bacteriophage as a biological surrogate to the SARS-CoV-2 virus.

59 BASIC BIOLOGICAL SCIENCES↗

Mid Infra-Red Laser Sensor for Continuous Sulfur Trioxide Monitoring to Improve Coal-Fired Power Plant Performance during Flexible Operations

During the course of this project, we performed exhaustive research and development of SO3/H2SO4 sensing technology for coal-fired power plant applications (Figure 1). The development culminated in a successful field campaign of a prototype continuous real-time H2SO4 monitor at a coal-fired power plant (TRL 6) accomplishing the primary goal of the project. The developed sensors utilize tunable laser absorption spectroscopy (TLAS) operating in the mid-infrared (Mid-IR) wavelength region, which is the so-called “molecular fingerprint” region. Systems operating in the Mid-IR have orders of magnitude more sensitivity than systems operating at shorter wavelengths, such as near-infrared (NIR). However, NIR systems are more widespread due to more mature supporting technology (e.g., fiber optics, optical components, etc.). In this project, we not only produced a specific Mid-IR sensor, we also advanced Mid-IR sensor technology in general through the development and demonstration of such supporting technology. In this project, we also developed proprietary broad tuning lasers enabling the ability to effectively measure SO3, H2SO4, H2O, and SO2. Different molecular species have unique spectral signatures that can be probed with lasers operating at different wavelengths. Standard TLAS uses relatively narrow wavelength tuning distributed feedback (DFB) lasers, which can typically only target a single species with narrow features, and are not appropriate for species with broad features, such as SO3 or H2SO4. In contrast, by developing unique, broad-tuning laser technology, we were able to measure these species, as well as SO2 and H2O simultaneously. Furthermore, to enable real-time analysis at a power plant, we modified a commercially available heated gas cell to operate in the Mid-IR wavelength range and fiber coupled the lasers to enable remote delivery of the beams. To generate reference data (library spectra), our collaborators at the University of California Irvine (UCI) developed a catalytic SO3 generation facility. It is worth mentioning that representative H2SO4 and SO3 Mid-IR spectra are not a part of any publicly available database and the data generated under this project is a valuable resource in and of itself. In addition, based on the UCI study we determined that detection of SO3 is complicated by the very strong SO2 absorption. For that reason, we concentrated on H2SO4 detection. Since SO3 and H2SO4 exist in a flue gas in a state of equilibrium, which depends on temperature and humidity, by measuring water concentration and controlling the temperature of the gas cell, we developed an approach to determine SO3 concentration from the H2SO4 measurement. During the development phase of the project, we performed three testing campaigns at our collaborator’s FERCo flue gas facility with conditions representative of the coal-fired power plant (~ 40ppm SO3, 1700ppm to 2800 ppm SO2, 10% water) with the exception of particulate matter. After three test campaigns at FERCo we performed field testing at Harrison Power Station. The final system was mounted on a duct and measured H2SO4, SO2 and water. The tests were highly successful with a demonstrated real-time H2SO4 precision of 1 ppm with a 1 second update. Our collaborators at EPRI conducted an industry survey and determined that there is a very high interest for the SO3/H2SO4 monitoring in the power generation industry as well as in heavy industries in general. Furthermore, work performed by OptoKnowledge beyond the scope of this project under a synergistic DOE SBIR determined another approach to SO3 detection. We applied for Phase II on this SBIR for development a of versatile SO3/H2SO4 sensor but were not selected. We are currently looking for another opportunity to leverage all the technological advancements produced by this project including but not limited to the flue gas facility at UCI, the hardware and software developed, and relationships with FERCo, EPRI CEMTEK, and Harrison Station.

01 COAL, LIGNITE, AND PEAT↗

Hit Expansion of a Noncovalent SARS-CoV-2 Main Protease Inhibitor

Inhibition of the SARS-CoV-2 main protease (M pro ) is a major focus of drug discovery efforts against COVID-19. Here we report a hit expansion of non-covalent inhibitors of M pro . Starting from a recently discovered scaffold (The COVID Moonshot Consortium. Open Science Discovery of Oral Non-Covalent SARS-CoV-2 Main Protease Inhibitor Therapeutics. bioRxiv 2020.10.29.339317) represented by an isoquinoline series, we searched a database of over a billion compounds using a cheminformatics molecular fingerprinting approach. We identified and tested 48 compounds in enzyme inhibition assays, of which 21 exhibited inhibitory activity above 50% at 20 μM. Among these, four compounds with IC 50 values around 1 μM were found. Interestingly, despite the large search space, the isoquinolone motif was conserved in each of these four strongest binders. Room-temperature X-ray structures of co-crystallized protein–inhibitor complexes were determined up to 1.9 Å resolution for two of these compounds as well as one of the stronger inhibitors in the original isoquinoline series, revealing essential interactions with the binding site and water molecules. Molecular dynamics simulations and quantum chemical calculations further elucidate the binding interactions as well as electrostatic effects on ligand binding. The results help explain the strength of this new non-covalent scaffold for M pro inhibition and inform lead optimization efforts for this series, while demonstrating the effectiveness of a high-throughput computational approach to expanding a pharmacophore library.

60 APPLIED LIFE SCIENCES↗

Iron-Tolerant Cyanobacteria: Ecophysiology and Fingerprinting

Although the iron-dependent physiology of marine and freshwater cyanobacterial strains has been the focus of extensive study, very few studies dedicated to the physiology and diversity of cyanobacteria inhabiting iron-depositing hot springs have been conducted. One of the few studies that have been conducted [B. Pierson, 1999] found that cyanobacterial members of iron depositing bacterial mat communities might increase the rate of iron oxidation in situ and that ferrous iron concentrations up to 1 mM significantly stimulated light dependent consumption of bicarbonate, suggesting a specific role for elevated iron in photosynthesis of cyanobacteria inhabiting iron-depositing hot springs. Our recent studies pertaining to the diversity and physiology of cyanobacteria populating iron-depositing hot springs in Great Yellowstone area (Western USA) indicated a number of different isolates exhibiting elevated tolerance to Fe(3+) (up to 1 mM). Moreover, stimulation of growth was observed with increased Fe(3+) (0.02-0.4 mM). Molecular fingerprinting of unialgal isolates revealed a new cyanobacterial genus and species Chroogloeocystis siderophila, an unicellular cyanobacterium with significant EPS sheath harboring colloidal Fe(3+) from iron enriched media. Our preliminary data suggest that some filamentous species of iron-tolerant cyanobacteria are capable of exocytosis of iron precipitated in cytoplasm. Prior to 2.4 Ga global oceans were likely significantly enriched in soluble iron [Lindsay et al, 2003], conditions which are not conducive to growth of most contemporary oxygenic cyanobacteria. Thus, iron-tolerant CB may have played important physiological and evolutionary roles in Earths history.

Brown, I. I.↗

Optimizing FPGA-based Accelerator Design for Large-Scale Molecular Similarity Search (Special Session Paper)

Molecular similarity search has been widely used in drug discovery to rapidly identify structurally similar compounds from large molecular databases. With the increasing size of chemical libraries, there is growing interest in the efficient ac- celeration of large-scale similarity search. Existing works mainly focus on CPU and GPU to accelerate the computation of Tatimoto coefficient in measuring the pairwise similarity between different molecular fingerprints. In this paper, we propose and optimize an FPGA-based accelerator design on exhaustive and approximate search algorithms. On exhaustive search using BitBound & fold- ing, we analyze the similarity cutoff and folding level relationship with search speedup and accuracy, and propose a scalable on- the-fly query engine on FPGAs to reduce the resource utilization and pipeline interval. We achieve a 450 million compounds-per- second processing throughput for a single query engine. On approximate search using hierarchical navigable small world (HNSW), a popular algorithm with high recall and query speed, we propose an FPGA-based graph traversal engine to utilize high throughput register array based priority queue and fine- grained distance calculation engine to increase the processing capability. Experimental results show that the proposed FPGA- based HNSW implementation achieves a 35× speedup than existing works on CPU. To the best of our knowledge, our FPGA- based implementation is the first attempt to accelerate molecular similarity search on FPGA and has the highest performance among existing approaches.

Peng, Hongwu↗