Search NASA⌕ Search

Engineering topics

Baker, Erin S.

Publications and source records attributed to Baker, Erin S..

PubChemLite Plus Collision Cross Section (CCS) Values for Enhanced Interpretation of Nontarget Environmental Data

Finding relevant chemicals in the vast (known) chemical space is a major challenge for environmental and exposomics studies leveraging nontarget high resolution mass spectrometry (NT-HRMS) methods. Chemical databases now contain hundreds of millions of chemicals, yet many are not relevant. This article details an extensive collaborative, open science effort to provide a dynamic collection of chemicals for environmental, metabolomics, and exposomics research, along with supporting information about their relevance to assist researchers in the interpretation of candidate hits. The PubChemLite for Exposomics collection is compiled from ten annotation categories within PubChem, enhanced with patent, literature and annotation counts, predicted partition coefficient (logP) values, as well as predicted collision cross section (CCS) values using CCSbase. Monthly versions are archived on Zenodo under a CC-BY license, supporting reproducible research, and a new interface has been developed, including historical trends of patent and literature data, for researchers to browse the collection. This article details how PubChemLite can support researchers in environmental and exposomics studies, describes efforts to increase the availability of experimental CCS values, and explores known limitations and potential for future developments. The data and code behind these efforts are openly available.

PubChem↗

A high-throughput workflow to analyze sequence-conformation relationships and explore hydrophobic patterning in disordered peptoids

Understanding how a macromolecule’s primary sequence governs its conformational landscape is crucial for elucidating its function, yet these design principles are still emerging for macromolecules with intrinsic disorder. Herein, we introduce a high-throughput workflow that implements a practical colorimetric conformational assay, introduces a semi-automated sequencing protocol using matrix-assisted laser desorption/ionization and tandem mass spectrometry (MALDI-MS/MS), and develops a generalizable sequence-structure algorithm. Using a model system of 20mer peptidomimetics containing polar glycine and hydrophobic N-butylglycine residues, we identified nine classifications of conformational disorder and isolated 122 unique sequences across varied compositions and conformations. Conformational distributions of three compositionally identical library sequences were corroborated through atomistic simulations and ion mobility spectrometry coupled with liquid chromatography. A data-driven strategy was developed using existing sequence variables and data-derived “motifs” to inform a machine-learning algorithm toward conformation prediction. Here, this multifaceted approach enhances our understanding of sequence-conformation relationships and offers a powerful tool for accelerating the discovery of materials with conformational control.

data-driven analysis↗

The lipidomics reporting checklist a framework for transparency of lipidomic experiments and repurposing resource data

The rapid increase in lipidomic studies has led to a collaborative effort within the community to establish standards and criteria for producing, documenting, and disseminating data. Creating a dynamic checklist that condenses key information about lipidomic experiments into common terminology will enhance the field's consistency, comparability, and repeatability. Here, we describe the structure and rationale of the established Lipidomics Minimal Reporting Checklist to increase transparency in lipidomics research.

59 BASIC BIOLOGICAL SCIENCES↗

Deciphering ApoE Genotype-Driven Proteomic and Lipidomic Alterations in Alzheimer’s Disease Across Distinct Brain Regions

Alzheimer’s disease (AD) is a neurodegenerative disease with a complex etiology influenced by confounding factors such as genetic polymorphisms, age, sex, and race. Traditionally, AD research has not prioritized these influences, resulting in dramatically skewed cohorts such as three times the number of Apolipoprotein E (APOE) e4-allele carriers in AD relative to healthy cohorts. Thus, the resulting molecular changes of AD have previously been complicated by the influence of apolipoprotein E disparities. To explore how apolipoprotein E polymorphism influences AD progression, 62 post-mortem patients consisting of 33 Alzheimer’s disease (AD) and 29 controls (Ctrl) were studied to balance the number of e4-allele carriers and facilitate a molecular comparison of the apolipoprotein E genotype. Lipid and protein perturbations were assessed across AD diagnosed brains compared to Ctrl brains, e4 allele carriers (APOE4+ for those carrying 1 or 2 e4s and APOE4- for non-e4 carriers), and differences in e3e3 and e3e4 Ctrl brains across two brain regions (frontal cortex (FCX) and cerebellum (CBM)). In conclusion, the region-specific influences of apolipoprotein E on AD mechanisms showcased mitochondrial dysfunction and cell proteostasis at the core of AD pathophysiology in the post-mortem brains, indicating these two processes may be influenced by genotypic differences and brain morphology.

Alzheimer’s disease↗

PeakDecoder enables machine learning-based metabolite annotation and accurate profiling in multidimensional mass spectrometry measurements

Multidimensional measurements using state-of-the-art separations and mass spectrometry provide advantages in untargeted metabolomics analyses for studying biological and environmental bio-chemical processes. However, the lack of rapid analytical methods and robust algorithms for these heterogeneous data has limited its application. Here, we develop and evaluate a sensitive and high-throughput analytical and computational workflow to enable accurate metabolite profiling. Our workflow combines liquid chromatography, ion mobility spectrometry and data-independent acquisition mass spectrometry with PeakDecoder, a machine learning-based algorithm that learns to distinguish true co-elution and co-mobility from raw data and calculates metabolite identification error rates. We apply PeakDecoder for metabolite profiling of various engineered strains of Aspergillus pseudoterreus, Aspergillus niger, Pseudomonas putida and Rhodosporidium toruloides. Results, validated manually and against selected reaction monitoring and gas-chromatography platforms, show that 2683 features could be confidently annotated and quantified across 116 microbial sample runs using a library built from 64 standards.

59 BASIC BIOLOGICAL SCIENCES↗