Search NASA⌕ Search

Engineering topics

Chang, Christine H.

Publications and source records attributed to Chang, Christine H..

Reviews and syntheses: Opportunities for robust use of peak intensities from high-resolution mass spectrometry in organic matter studies

Abstract. Earth's biogeochemical cycles are intimately tied to the biotic and abiotic processing of organic matter (OM). Spatial and temporal variations in OM chemistry are often studied using direct infusion, high-resolution Fourier transform mass spectrometry (FTMS). An increasingly common approach is to use ecological metrics (e.g., within-sample diversity) to summarize high-dimensional FTMS data, notably Fourier transform ion cyclotron resonance mass spectrometry (FT-ICR MS). However, problems can arise when FTMS peak-intensity data are used in a way that is analogous to abundances in ecological analyses (e.g., species abundance distributions). Using peak-intensity data in this way requires the assumption that intensities act as direct proxies for concentrations. Here, we show that comparisons of the same peak across samples (within-peak) may carry information regarding variations in relative concentration, but comparing different peaks (between-peak) within or between samples does not. We further developed a simulation model to study the quantitative implications of using peak intensities to compute ecological metrics (e.g., intensity-weighted mean properties and diversity) that rely on information about both within-peak and between-peak shifts in relative abundance. We found that, despite analytical limitations in linking concentration to intensity, ecological metrics often perform well in terms of providing robust qualitative inferences and sometimes quantitatively accurate estimates of diversity and mean molecular characteristics. We conclude with recommendations for the robust use of peak intensities for natural organic matter studies. A primary recommendation is the use and extension of the simulation model to provide objective guidance on the degree to which conceptual and quantitative inferences can be made for a given analysis of a given dataset. Broad use of this approach can help ensure rigorous scientific outcomes from the use of FTMS peak intensities in environmental applications.

54 ENVIRONMENTAL SCIENCES↗

Molecular Vision - Multimodal, multitask retrieval of molecular structure from measured signatures for reference-free compound identification

We are currently at risk of generating false conclusions based on limited methods to identify small molecules in biological systems and in chemical forensics. By definition, the chemical structures of novel small molecules have not been determined, let alone measured or synthesized. Currently, unambiguous structure determination of small molecules is constrained by the time and effort needed to isolate compounds and perform de novo structure elucidation using laboratory-based methods, significantly extending the time to inform mitigation strategies. To address this gap, we have developed a deep learning approach to directly map molecular structure to experimental signatures. We aim to unify measurement technologies employed in untargeted small molecule identification studies—such as infrared (IR) spectrometry, tandem mass spectrometry (MS/MS), ion mobility spectrometry-derived collision cross section (CCS)—through use of a multimodal, multitask deep learning architecture. Where existing methods require direct generation of information-rich spectra and/or properties, an inherently difficult task, we will simplify molecular signature-based identification by posing the problem as a recognition or retrieval task. The model is thus presented with relevant endpoints – structure and one or more molecular signatures – and need only determine whether they are semantically related. Thus, our approach offers the following advantages over existing techniques: (i) circumvents difficulties associated with direct generation of molecular signatures from structure and structure from signatures; (ii) incorporates multiple molecular signatures simultaneously, as available, to support identification; and (iii) enables rapid computation of structural embeddings toward broad coverage of known chemical space. Taken together, the approach removes the need to explicitly obtain or compute reference spectra, representing a powerful method for compound identification that requires only experimentally observed signatures.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Snekmer: a scalable pipeline for protein sequence fingerprinting based on amino acid recoding

Abstract Motivation The vast expansion of sequence data generated from single organisms and microbiomes has precipitated the need for faster and more sensitive methods to assess evolutionary and functional relationships between proteins. Representing proteins as sets of short peptide sequences (kmers) has been used for rapid, accurate classification of proteins into functional categories; however, this approach employs an exact-match methodology and thus may be limited in terms of sensitivity and coverage. We have previously used similarity groupings, based on the chemical properties of amino acids, to form reduced character sets and recode proteins. This amino acid recoding (AAR) approach simplifies the construction of protein representations in the form of kmer vectors, which can link sequences with distant sequence similarity and provide accurate classification of problematic protein families. Results Here, we describe Snekmer, a software tool for recoding proteins into AAR kmer vectors and performing either (i) construction of supervised classification models trained on input protein families or (ii) clustering for de novo determination of protein families. We provide examples of the operation of the tool against a set of nitrogen cycling families originally collected using both standard hidden Markov models and a larger set of proteins from Uniprot and demonstrate that our method accurately differentiates these sequences in both operation modes. Availability and implementation Snekmer is written in Python using Snakemake. Code and data used in this article, along with tutorial notebooks, are available at http://github.com/PNNL-CompBio/Snekmer under an open-source BSD-3 license. Supplementary information Supplementary data are available at Bioinformatics Advances online.

59 BASIC BIOLOGICAL SCIENCES↗

Effects of warming on bacterial growth rates in a peat soil under ambient and elevated CO 2

Boreal peatlands are important global carbon reservoirs that are particularly vulnerable to predicted climate changes such as increasing CO 2 and temperature. Since microbial activities regulate the balance of carbon sequestered into soil organic matter or remineralized to CO 2 , characterizing their response to these environmental factors is critical to predicting how peatland ecosystems will affect climate-carbon cycle feedbacks. Here we examined in-situ taxon-specific variation in microbial growth under long-term elevated CO 2 and across a gradient of warming treatments in a northern Minnesota peat bog using quantitative stable isotope probing with 18 O-water. Across temperatures, bacterial taxa were grouped according to the excess atom fraction 18 O (EAF) of their genomes, a proxy for DNA replication and hence, growth. Taxon-specific growth across CO 2 and temperature treatments clustered into relatively few response patterns. While a large portion of taxa showed little to no growth under ambient CO 2 , many of the same taxa grew rapidly under elevated CO 2 . We found support for phylogenetic conservation of response patterns among Acidobacteria and Proteobacteria, the two most abundant phyla in our data. Our results suggest certain taxa may be primed for new climate conditions and have a greater influence on carbon cycling with implications for future climate mitigation strategies.

16S amplicon sequencing, Carbon Dioxide (CO2), pea↗

Mass Spectrometry Adduct Calculator

We describe the Mass Spectrometry Adduct Calculator (MSAC), an automated Python tool to calculate the adduct ion masses of a parent molecule. Here, adduct refers to a version of a parent molecule [M] that is charged due to addition or loss of atoms and electrons resulting in a charged ion, e.g. [M+H] + . MSAC includes a database of 2,341 potential ions and their mass-to-charge ratios (m/z) as extracted from the NIST/EPA/NIH Mass Spectral Library (NIST17), the Global Natural Products Social Molecular Networking Public Spectral Libraries (GNPS), and MassBank of North America (MoNA). The calculator relies on user-selected subsets of the combined database to calculate expected m/z for adducts of molecules supplied as formulas This tool is intended to help researchers create identification libraries to collect evidence for the presence of molecules in mass spectrometry data. While the included adduct database focuses on adducts typically detected during liquid chromatography-mass spectrometry analyses, users may supply their own lists of adducts and charge states for calculating expected m/z. We also analyzed statistics on adducts from spectra contained in the three selected mass spectral libraries. MSAC is freely available at https://github.com/pnnl/MSAC.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗