Search NASA⌕ Search

SEARCH · Search NASA

Results for “reference-free”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Statistically-driven Experimental Design to Improve Reference-free Quantification of Small Molecules by Liquid Chromatography-Mass Spectrometry

Non-targeted analysis of small molecules and metabolites in unknown, complex samples using liquid chromatography-tandem mass spectrometry remains challenging. One of the main bottlenecks is the extensive unannotated regions of metabolomics mass spectrometry data, resulting in knowledge gaps. Small molecule annotation in mass spectrometry data has conventionally relied on reference standards and libraries for compound identification and confirmation, which can constrain compound identification to those molecules already known, thus limiting the ability to discover new knowledge and new markers. Retention time prediction can facilitate and expedite unknown compound identification in non-targeted analysis of complex metabolomics samples. Additionally, accurate retention time predictions can also inform sample mixture design for LC-MS/MS analyses. However, current machine learning-based methods for retention time prediction are typically developed for specific chromatographic platforms and are not generalizable across scales. And while technologies and methods to improve reference-free metabolite identification for more comprehensive annotation of unknowns has received much attention, development of the same for quantitation without reference standards has been much more limited, despite its importance in toxicological, environmental, food safety, forensics, and clinical applications. We believe that a reference-free quantitation strategy that exploits mass spectrometry data already collected for reference-free identification can provide much more insight on unknowns, and move the metabolomics field for more complete unknowns characterization. As such, we pursue two efforts to improve upon current state-of-the-art methods in non-targeted analysis: (1) machine learning-based retention time prediction and (2) statistical design of experiments framework for reference-free quantitation. In this work, we develop and demonstrate (1) a generalizable retention time prediction capability across chromatographic conditions and scales, and (2) a statistical design-based framework for response factor contribution elucidation and reference-free quantitation. Evaluation of our retention time prediction model, PrediToR, showed approximately 24% improvement over current models, and we observed approximately 10X improvement in concentration estimation accuracy from our statistical design-based response factor model over a primarily ionization efficiency-based model. We expect that future efforts to improve upon these new capabilities will further advance non-targeted analysis of small molecules towards truly reference-free metabolomics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Evaluation of a Reference-Free Collision Cross Section Calibration Strategy for Proteomics Using SLIM-Based High-Resolution Ion Mobility Spectrometry–Mass Spectrometry

Ion mobility spectrometry (IMS) is a gas-phase analytical technique that separates ions with different sizes and shapes and is compatible with mass spectrometry (MS) to provide an additional separation dimension. The rapid nature of the IMS separation combined with the high sensitivity of MS-based detection and the ability to derive structural information on analytes in the form of the property collision cross section (CCS) makes IMS particularly well-suited for characterizing complex samples in -omics applications. In such applications, the quality of CCS from IMS measurements is critical to confident annotation of the detected components in the complex -omics samples. However, most IMS instrumentation in mainstream use requires calibration to calculate CCS from measured arrival times, with the most notable exception being drift tube IMS measurements using multifield methods. The strategy for calibrating CCS values, particularly selection of appropriate calibrants, has important implications for CCS accuracy, reproducibility, and transferability between laboratories. The conventional approach to CCS calibration involves explicitly defining calibrants ahead of data acquisition and crucially relies upon availability of reference CCS values. In this work, we present a novel reference-free approach to CCS calibration which leverages trends among putatively identified features and computational CCS prediction to conduct calibrations post-data acquisition and without relying on explicitly defined calibrants. We demonstrated the utility of this reference-free CCS calibration strategy for proteomics application using high-resolution structures for lossless ion manipulations (SLIM)-based IMS-MS. In conclusion, we first validated the accuracy of CCS values using a set of synthetic peptides and then demonstrated using a complex peptide sample from cell lysate.

59 BASIC BIOLOGICAL SCIENCES↗

Similarity Downselection: Finding the n Most Dissimilar Molecular Conformers for Reference-Free Metabolomics

Computational methods for creating in silico libraries of molecular descriptors (e.g., collision cross sections) are becoming increasingly prevalent due to the limited number of authentic reference materials available for traditional library building. These so-called “reference-free metabolomics” methods require sampling sets of molecular conformers in order to produce high accuracy property predictions. Due to the computational cost of the subsequent calculations for each conformer, there is a need to sample the most relevant subset and avoid repeating calculations on conformers that are nearly identical. The goal of this study is to introduce a heuristic method of finding the most dissimilar conformers from a larger population in order to help speed up reference-free calculation methods and maintain a high property prediction accuracy. Finding the set of the n items most dissimilar from each other out of a larger population becomes increasingly difficult and computationally expensive as either n or the population size grows large. Because there exists a pairwise relationship between each item and all other items in the population, finding the set of the n most dissimilar items is different than simply sorting an array of numbers. For instance, if you have a set of the most dissimilar n = 4 items, one or more of the items from n = 4 might not be in the set n = 5. An exact solution would have to search all possible combinations of size n in the population exhaustively. We present an open-source software called similarity downselection (SDS), written in Python and freely available on GitHub. SDS implements a heuristic algorithm for quickly finding the approximate set(s) of the n most dissimilar items. We benchmark SDS against a Monte Carlo method, which attempts to find the exact solution through repeated random sampling. We show that for SDS to find the set of n most dissimilar conformers, our method is not only orders of magnitude faster, but it is also more accurate than running Monte Carlo for 1,000,000 iterations, each searching for set sizes n = 3–7 out of a population of 50,000. We also benchmark SDS against the exact solution for example small populations, showing that SDS produces a solution close to the exact solution in these instances. Using theoretical approaches, we also demonstrate the constraints of the greedy algorithm and its efficacy as a ratio to the exact solution.

97 MATHEMATICS AND COMPUTING↗

Assessing the Impact of Measurement Precision on Metabolite Identification Probability in Multidimensional Mass Spectrometry-Based, Reference-Free Metabolomics

Identification of compounds with minimal ambiguity remains a central challenge in mass spectrometry-based metabolomics. Conventional compound identification relies on comparing analytical signatures (e.g., mass-to-charge ratio, collision cross section, tandem mass spectra) against reference data obtained from measurements of authentic chemical standards. The breadth of annotatable compounds using this approach is necessarily limited by availability of authentic standards, analytical throughput, and resolving power of the separations that underly the measurements. The maturation of computational methods, both theory-driven and artificial intelligence/machine learning-based, for prediction of various molecular properties relevant to multidimensional mass spectrometry measurements has opened the door to a new “reference-free” paradigm of compound annotation. Through augmenting existing reference data for molecular properties with computational predictions, the universe of identifiable chemical species can be expanded significantly beyond its current limits. An unexplored aspect of this novel approach is understanding how to gauge confidence in resulting annotations, especially as the compound search space is expanded. Intuitively, the confidence of a compound annotation is related to the inherent discriminatory power of the molecular properties used for identification, as well as the precision with which the properties are measured or predicted. In this work, we characterize this relationship between measurement precision and identification probability in a systematic and quantitative fashion for a defined region of chemical space that includes organic small molecule metabolites. Importantly, this work establishes a framework for conducting metabolite identification probability analysis that enables others to quantify this relationship for their own compounds and properties of interest.

Metabolite Identification↗

Reference-free direct digital lock-in method and apparatus

A reference-free direct digital lock-in system (RDDL 10) has a first input coupled to a periodic electrical signal and an output for outputting an indication of a magnitude of a desired periodic signal component. The RDDL also has a second input for receiving a signal (9) that specifies a reference period value, and operates to autonomously generate a lock-in reference signal having a specified period and a phase that is adjusted to maximize a magnitude of the outputted desired periodic signal component. In an embodiment of a measurement system that includes the RDDL 10 an optical source provides a chopped light beam having wavelengths within a predetermined range of wavelengths, and the periodic electrical signal is generated by at least one photodetector that is illuminated by the chopped light beam. In this embodiment the measurement system characterizes, for at least one wavelength of light that is generated by the optical source, a spectral response of the at least one photodetector. The RDDL can operate in nonreal-time upon previously generated and stored digital equivalent values of the periodic electrical signal or signals.

Henry, James E.↗

Molecular Vision - Multimodal, multitask retrieval of molecular structure from measured signatures for reference-free compound identification

We are currently at risk of generating false conclusions based on limited methods to identify small molecules in biological systems and in chemical forensics. By definition, the chemical structures of novel small molecules have not been determined, let alone measured or synthesized. Currently, unambiguous structure determination of small molecules is constrained by the time and effort needed to isolate compounds and perform de novo structure elucidation using laboratory-based methods, significantly extending the time to inform mitigation strategies. To address this gap, we have developed a deep learning approach to directly map molecular structure to experimental signatures. We aim to unify measurement technologies employed in untargeted small molecule identification studies—such as infrared (IR) spectrometry, tandem mass spectrometry (MS/MS), ion mobility spectrometry-derived collision cross section (CCS)—through use of a multimodal, multitask deep learning architecture. Where existing methods require direct generation of information-rich spectra and/or properties, an inherently difficult task, we will simplify molecular signature-based identification by posing the problem as a recognition or retrieval task. The model is thus presented with relevant endpoints – structure and one or more molecular signatures – and need only determine whether they are semantically related. Thus, our approach offers the following advantages over existing techniques: (i) circumvents difficulties associated with direct generation of molecular signatures from structure and structure from signatures; (ii) incorporates multiple molecular signatures simultaneously, as available, to support identification; and (iii) enables rapid computation of structural embeddings toward broad coverage of known chemical space. Taken together, the approach removes the need to explicitly obtain or compute reference spectra, representing a powerful method for compound identification that requires only experimentally observed signatures.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Reference-free structural variant detection in microbiomes via long-read co-assembly graphs

Motivation: The study of bacterial genome dynamics is vital for understanding the mechanisms underlying microbial adaptation, growth, and their impact on host phenotype. Structural variants (SVs), genomic alterations of 50 base pairs or more, play a pivotal role in driving evolutionary processes and maintaining genomic heterogeneity within bacterial populations. While SV detection in isolate genomes is relatively straightforward, metagenomes present broader challenges due to the absence of clear reference genomes and the presence of mixed strains. In response, our proposed method rhea, forgoes reference genomes and metagenome-assembled genomes (MAGs) by encompassing all metagenomic samples in a series (time or other metric) into a single co-assembly graph. The log fold change in graph coverage between successive samples is then calculated to call SVs that are thriving or declining. Results: We show rhea to outperform existing methods for SV and horizontal gene transfer (HGT) detection in two simulated mock metagenomes, particularly as the simulated reads diverge from reference genomes and an increase in strain diversity is incorporated. We additionally demonstrate use cases for rhea on series metagenomic data of environmental and fermented food microbiomes to detect specific sequence alterations between successive time and temperature samples, suggesting host advantage. Our approach leverages previous work in assembly graph structural and coverage patterns to provide versatility in studying SVs across diverse and poorly characterized microbial communities for more comprehensive insights into microbial gene flux.

59 BASIC BIOLOGICAL SCIENCES↗

Reference-Free, Projection Background-Oriented Schlieren

A projection background-oriented schlieren (P-BOS) system is developed and demonstrated. Instead of a background that has a speckle pattern printed, painted, or otherwise deposited onto its surface, and can thus not be altered, the pattern here is projected onto the background. This allows for changes to the speckle pattern without replacement of the background material, which can be time-consuming and expensive. Reference images are acquired simultaneously with flow images. A pre-test transformation between the reference and flow images allows for on-the-fly changes to the speckle pattern during a test without requiring a stoppage of the flow, allowing for optimization of the BOS signal. Using a programmable LCD screen as the speckled optic allows for remote control of the speckle pattern. Because the background is not speckled, the system can be easily transformed to acquired shadowgraph images by removing the speckled optic and closing the aperture of the light source, which is useful for achieving measurements with higher spatial resolution. The system can be nearly as compact as a conventional BOS system, and can be assembled with polarized optics to reduce or eliminate window reflections and glare, making this system particularly well-suited for wind tunnel testing.

Joshua M Weisberger↗

Elucidating the Gas-Phase Behavior of Nitazene Analog Protomers Using Structures for Lossless Ion Manipulations Ion Mobility-Orbitrap Mass Spectrometry

2-benzylbenzimidazoles, or “nitazenes”, are a class of novel synthetic opioids (NSOs) that are increasingly being detected alongside fentanyl analogs and other opioids in drug overdose cases. Nitazenes can be 20x more potent than fentanyl but are not routinely tested for during postmortem or clinical toxicology drug screens; thus, their prevalence in drug overdose cases may be under-reported. Traditional analytical workflows utilizing liquid chromatography-tandem mass spectrometry (LC-MS/MS) often require additional confirmation with authentic reference standards to identify a novel nitazene. However, additional analytical measurements with ion mobility spectrometry (IMS) may provide a path towards reference-free identification, which would greatly accelerate NSO identification rates in toxicology labs. Presented here are the first IMS and collision cross section (CCS) measurements on a set of fourteen nitazene analogs using a Structures for Lossless Ion Manipulations (SLIM)-Orbitrap MS. All nitazenes exhibited two high intensity baseline-separated IMS distributions, which fentanyls and other drug and drug-like compounds also exhibit. Incorporating water into the electrospray ionization (ESI) solution caused the intensities of the higher mobility IMS distributions to increase the intensities of the lower mobility IMS distributions to decrease. Nitazenes lacking a nitro group at the R1 position exhibited the greatest shifts in signal intensities due to water. Furthermore, IMS-MS/MS experiments showed that the higher mobility IMS distributions of all nitazenes produced fragment ions with m/z 72, 100, and other low intensity fragments while the lower mobility IMS distributions only produced fragment ions with m/z 72 and 100. The IMS, solvent, and fragmentation studies provide experimental evidence that nitazenes potentially exhibit three gas-phase protomers. In conclusion, the cyclic IMS capability of SLIM was also employed to partially resolve four sets of structurally similar nitazene isomers (e.g., protonitazene/isotonitazene, butonitazene/isobutonitazene/secbutonitazene), showcasing the potential of using high-resolution IMS separations in MS-based workflows for reference-free identification of emerging nitazenes and other NSOs.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Heracles: Predictive Tools for Opioid Crisis Intervention - m/q Initiative Project Report

The opioid crisis in the United States is being fueled primarily by fentanyl and its molecular analogs, which can be anywhere from 50 to 1,000 times more potent than morphine. Fentanyl itself is straightforward to synthesize; furthermore, the structure is such that fentanyl’s flexible, rotatable side chains are easy to modify to create new analogs. Reference-free computational techniques to predict and identify new fentanyls have the potential to provide a desperately needed preemptive advantage to regulatory stakeholders and toxicologists. The computational pipeline Heracles was developed with this preemptive advantage in mind. Heracles has two primary components: 1) the creation of an in silico library of putative fentanyl analogs, and 2) a downselection pipeline to prioritize generated fentanyl analogs predicted to be potent and easy to synthesize. Experimental observables were also predicted for prioritized analogs, with validation of the observables begun. Heracles has demonstrated potential to aid in the advancement of reference-free paradigms while providing new tools to first responders and other stakeholders attempting to mitigate the opioid crisis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Identification of Unique Fragmentation Patterns of Fentanyl Analog Protomers Using Structures for Lossless Ion Manipulations Ion Mobility-Orbitrap Mass Spectrometry

The opioid crisis in the United States is being fueled by the rapid emergence of new fentanyl analogs and precursors that can elude traditional library-based screening methods, which require data from known reference compounds. Since reference compounds are unavailable for new fentanyl analogs, we examined if fentanyls (fentanyl + fentanyl analogs) could be identified in a reference-free manner using a combination of electrospray ionization (ESI), high-resolution ion mobility (IM) spectrometry, high-resolution mass spectrometry (MS), and higher-energy collision-induced dissociation (MS/MS). We analyzed a mixture containing nine fentanyls and W-15 (a structurally similar molecule) and found that the protonated forms of all fentanyls uniquely exhibited two baseline separated IM distributions that produced different MS/MS patterns. Upon fragmentation, both IM distributions of all fentanyls produced two high intensity fragments resulting from amine site cleavages. The higher mobility distributions of all fentanyls also produced several low intensity fragments, but surprisingly, these same fragments exhibited much greater intensities in the lower mobility distributions. This observation demonstrates that many fragments of fentanyls predominantly originate from one of two different gas-phase structures (suggestive of protomers). Furthermore, increasing the water concentration in the ESI solution increased the intensity of the lower mobility distribution relative to the higher mobility distribution, which further supports that fentanyls exist as two gas-phase protomers. In conclusion, our new observations on the IM and MS/MS properties of fentanyls can be exploited to positively identify them as fentanyls without requiring reference libraries and will hopefully assist first responders and law enforcement in combating new and emerging fentanyls.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Dual-Gated Structures for Lossless Ion Manipulations-Ion Mobility Orbitrap Mass Spectrometry Platform for Combined Ultra-High-Resolution Molecular Analysis

High-resolution ion mobility spectrometry-mass spectrometry (HR-IMS-MS) instruments have enormously advanced the ability to characterize complex biological mixtures. Unfortunately, HR-IMS and HR-MS measurements are typically performed independently due to mismatches in analysis time scales. Here we overcome this limitation by using a dual-gated ion injection approach to couple an 11-meter path length structures for lossless ion manipulations (SLIM) module to a Q-Exactive Plus Orbitrap MS. The dual-gate setup was implemented by placing one ion gate before the SLIM module and a second ion gate after. The dual-gated ion injection approach allowed the new SLIM-Orbitrap platform to simultaneously perform an 11-meter SLIM separation, Orbitrap mass analysis using the highest selectable mass resolution setting (up to 140k), and high-energy collision induced dissociation (HCD) in ~25 minutes over an $m/z$ range of ~1500 amu. The SLIM-Orbitrap was initially characterized using a mixture of standard phosphazene cations and demonstrated an average SLIM CCS resolving power (Rp CCS ) of ~218 and SLIM peak capacity of ~156 while simultaneously obtaining high mass resolutions. SLIM-Orbitrap analysis with fragmentation was then performed on mixtures of standard peptides and two reverse peptides (SDGRG 1+ , GRGDS 1+ , Rp CCS = 305) to demonstrate the utility of combined HR-IMS-MS/MS measurements for peptide identification. Our new HR-IMS-MS/MS capability was further demonstrated by analyzing a complex lipid mixture and showcasing SLIM separations on isobaric lipids. In conclusion, this new SLIM-Orbitrap platform demonstrates a critical new capability for proteomics and lipidomics applications, and the high-resolution multimodal data obtainable with this system establishes the foundation for reference-free identification of unknown ion structures.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Unraveling the functional dark matter through global metagenomics

Metagenomes encode an enormous diversity of proteins, reflecting a multiplicity of functions and activities1,2. Exploration of this vast sequence space has been limited to a comparative analysis against reference microbial genomes and protein families derived from those genomes. Here, to examine the scale of yet untapped functional diversity beyond what is currently possible through the lens of reference genomes, we develop a computational approach to generate reference-free protein families from the sequence space in metagenomes. We analyse 26,931 metagenomes and identify 1.17 billion protein sequences longer than 35 amino acids with no similarity to any sequences from 102,491 reference genomes or the Pfam database3. Using massively parallel graph-based clustering, we group these proteins into 106,198 novel sequence clusters with more than 100 members, doubling the number of protein families obtained from the reference genomes clustered using the same approach. We annotate these families on the basis of their taxonomic, habitat, geographical and gene neighbourhood distributions and, where sufficient sequence diversity is available, predict protein three-dimensional models, revealing novel structures. Overall, our results uncover an enormously diverse functional space, highlighting the importance of further exploring the microbial functional dark matter.

54 ENVIRONMENTAL SCIENCES↗

MnEdgeNet for accurate decomposition of mixed oxidation states for Mn XAS and EELS L2,3 edges without reference and calibration

Accurate decomposition of the mixed Mn oxidation states is highly important for characterizing the electronic structures, charge transfer and redox centers for electronic, and electrocatalytic and energy storage materials that contain Mn. Electron energy loss spectroscopy (EELS) and soft X-ray absorption spectroscopy (XAS) measurements of the Mn L2,3 edges are widely used for this purpose. To date, although the measurements of the Mn L2,3 edges are straightforward given the sample is prepared properly, an accurate decomposition of the mix valence states of Mn remains non-trivial. For both EELS and XAS, 2+, 3+, and 4+ reference spectra need to be taken on the same instrument/beamline and preferably in the same experimental session because the instrumental resolution and the energy axis offset could vary from one session to another. To circumvent this hurdle, in this study, we adopted a deep learning approach and developed a calibration-free and reference-free method to decompose the oxidation state of Mn L2,3 edges for both EELS and XAS. A deep learning regression model is trained to accurately predict the composition of the mix valence state of Mn. To synthesize physics-informed and ground-truth labeled training datasets, we created a forward model that takes into account plural scattering, instrumentation broadening, noise, and energy axis offset. With that, we created a 1.2 million-spectrum database with 1-by-3 oxidation state composition ground truth vectors. The library includes a sufficient variety of data including both EELS and XAS spectra. By training on this large database, our convolutional neural network achieves 85% accuracy on the validation dataset. We tested the model and found it is robust against noise (down to PSNR of 10) and plural scattering (up to t/λ = 1). We further validated the model against spectral data that were not used in training. In particular, the model shows high accuracy and high sensitivity for the decomposition of Mn 3 O 4 , MnO, Mn 2 O 3 , and MnO 2 . The accurate decomposition of Mn 3 O 4 experimental data shows the model is quantitatively correct and can be deployed for real experimental data. Our model will not only be a valuable tool to researchers and material scientists but also can assist experienced electron microscopists and synchrotron scientists in the automated analysis of Mn L edge data.

25 ENERGY STORAGE↗

A machine learning decision criterion for reducing scan time for hyperspectral neutron computed tomography systems

We present the first machine learning-based autonomous hyperspectral neutron computed tomography experiment performed at the Spallation Neutron Source. Hyperspectral neutron computed tomography allows the characterization of samples by enabling the reconstruction of crystallographic information and elemental/isotopic composition of objects relevant to materials science. High quality reconstructions using traditional algorithms such as the filtered back projection require a high signal-to-noise ratio across a wide wavelength range combined with a large number of projections. This results in scan times of several days to acquire hundreds of hyperspectral projections, during which end users have minimal feedback. To address these challenges, a golden ratio scanning protocol combined with model-based image reconstruction algorithms have been proposed. This novel approach enables high quality real-time reconstructions from streaming experimental data, thus providing feedback to users, while requiring fewer yet a fixed number of projections compared to the filtered back projection method. In this paper, we propose a novel machine learning criterion that can terminate a streaming neutron tomography scan once sufficient information is obtained based on the current set of measurements. Our decision criterion uses a quality score which combines a reference-free image quality metric computed using a pre-trained deep neural network with a metric that measures differences between consecutive reconstructions. The results show that our method can reduce the measurement time by approximately a factor of five compared to a baseline method based on filtered back projection for the samples we studied while automatically terminating the scans.

97 MATHEMATICS AND COMPUTING↗