Search NASA⌕ Search

Engineering topics

Chu, Fanny

Publications and source records attributed to Chu, Fanny.

Statistically-driven Experimental Design to Improve Reference-free Quantification of Small Molecules by Liquid Chromatography-Mass Spectrometry

Non-targeted analysis of small molecules and metabolites in unknown, complex samples using liquid chromatography-tandem mass spectrometry remains challenging. One of the main bottlenecks is the extensive unannotated regions of metabolomics mass spectrometry data, resulting in knowledge gaps. Small molecule annotation in mass spectrometry data has conventionally relied on reference standards and libraries for compound identification and confirmation, which can constrain compound identification to those molecules already known, thus limiting the ability to discover new knowledge and new markers. Retention time prediction can facilitate and expedite unknown compound identification in non-targeted analysis of complex metabolomics samples. Additionally, accurate retention time predictions can also inform sample mixture design for LC-MS/MS analyses. However, current machine learning-based methods for retention time prediction are typically developed for specific chromatographic platforms and are not generalizable across scales. And while technologies and methods to improve reference-free metabolite identification for more comprehensive annotation of unknowns has received much attention, development of the same for quantitation without reference standards has been much more limited, despite its importance in toxicological, environmental, food safety, forensics, and clinical applications. We believe that a reference-free quantitation strategy that exploits mass spectrometry data already collected for reference-free identification can provide much more insight on unknowns, and move the metabolomics field for more complete unknowns characterization. As such, we pursue two efforts to improve upon current state-of-the-art methods in non-targeted analysis: (1) machine learning-based retention time prediction and (2) statistical design of experiments framework for reference-free quantitation. In this work, we develop and demonstrate (1) a generalizable retention time prediction capability across chromatographic conditions and scales, and (2) a statistical design-based framework for response factor contribution elucidation and reference-free quantitation. Evaluation of our retention time prediction model, PrediToR, showed approximately 24% improvement over current models, and we observed approximately 10X improvement in concentration estimation accuracy from our statistical design-based response factor model over a primarily ionization efficiency-based model. We expect that future efforts to improve upon these new capabilities will further advance non-targeted analysis of small molecules towards truly reference-free metabolomics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Proteomics Analysis of Human Contaminant Proteins

Complete characterization of unknowns via proteomics remains challenging. There exist regions of mass spectrometry-based proteomics data where empirical measurements are not attributed to peptides, and/or sequenced peptides from mass spectra are not attributed to any source. These uncharacterized regions are known as the “dark” proteome. Many proteomics tools rely on some a priori knowledge of sample composition; few tools allow for investigation of unknowns without relying on composition assumptions. Further, the potential low abundance of minor traces in these uncharacterized regions can make elucidation of the “dark” proteome challenging. Herein, we describe the development and evaluation of approaches to study the “dark” proteome and move towards an untargeted approach for more complete characterization, namely by studying minor human protein traces in non-human samples and combining that approach with non-human source organism identification without relying on assumptions. Human protein markers, in the form of genetically variant peptides, have been extensively examined in a variety of human matrices, including blood, plasma, and hair, but have yet to be investigated in non-human samples, such as cell cultures, as human contaminant traces. Genetically variant peptides are those that are found in proteins carrying single nucleotide polymorphisms. In this work, we aimed to (1) investigate the feasibility of detecting human contaminant genetically variant peptides (GVPs) in a diverse set of non-human organisms using public proteomics data and a computational pipeline, as well as to (2) develop a combined capability for untargeted source organism characterization and GVP detection. To our knowledge, this is the first report of applying these approaches towards a more complete proteomic characterization of unknowns. We successfully demonstrate the feasibility of broad human contaminant GVP detection in proteomics data, develop a better understanding of GVP detectability, characterize the sample-to-sample variability in GVP detection, and identify a core set of GVPs that can potentially be used as markers indicative of the human contaminant traces portion of the “dark” proteome. Further, we developed and evaluated a combined pipeline, MARLOWE-GVP, that enables both untargeted source organism characterization and GVP detection. We show high accuracy of correct source organism characterization and high degree of similarity of human contaminant GVP detection compared to the conventional approach. Success on both these efforts have allowed us to advance our understanding and characterization of the “dark” proteome.

59 BASIC BIOLOGICAL SCIENCES↗