Search NASA⌕ Search

SEARCH · Search NASA

Results for “false discovery rate”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Target–decoy false discovery rate estimation using Crema

Assigning statistical confidence estimates to discoveries produced by a tandem mass spectrometry proteomics experiment is critical to enabling principled interpretation of the results and assessing the cost/benefit ratio of experimental follow-up. The most common technique for computing such estimates is to use target-decoy competition (TDC), in which observed spectra are searched against a database of real (target) peptides and a database of shuffled or reversed (decoy) peptides. TDC procedures for estimating the false discovery rate (FDR) at a given score threshold have been developed for application at the level of spectra, peptides, or proteins. Although these techniques are relatively straightforward to implement, it is common in the literature to skip over the implementation details or even to make mistakes in how the TDC procedures are applied in practice. Here we present Crema, an open-source Python tool that implements several TDC methods of spectrum-, peptide- and protein-level FDR estimation. Crema is compatible with a variety of existing database search tools and provides a straightforward way to obtain robust FDR estimates.

59 BASIC BIOLOGICAL SCIENCES↗

Enhanced climate reproducibility testing with false discovery rate correction

Simulating the Earth's climate is an important and complex problem, thus climate models are similarly complex, comprised of millions of lines of code. In order to appropriately utilize the latest computational and software infrastructure advancements in Earth system models running on modern hybrid computing architectures to improve their performance, precision, accuracy, or all three; it is important to ensure that model simulations are repeatable and robust. This introduces the need for establishing statistical or non-bit-for-bit reproducibility, since bit-for-bit reproducibility may not always be achievable. Here, we propose a short-simulation ensemble-based test for an atmosphere model to evaluate the null hypothesis that modified model results are statistically equivalent to that of the original model. We implement this test in version 2 of the US Department of Energy's Energy Exascale Earth System Model (E3SM). The test evaluates a standard set of output variables across the two simulation ensembles and uses a false discovery rate correction to account for multiple testing. The false positive rates of the test are examined using re-sampling techniques on large simulation ensembles and are found to be lower than the currently implemented bootstrapping-based testing approach in E3SM. We also evaluate the statistical power of the test using perturbed simulation ensemble suites, each with a progressively larger magnitude of change to a tuning parameter. The new test is generally found to exhibit more statistical power than the current approach, being able to detect smaller changes in parameter values with higher confidence.

Kelleher, Michael E. [Oak Ridge National Laborator↗

Increased inflammation as well as decreased endoplasmic reticulum stress and translation differentiate pancreatic islets from donors with pre-symptomatic stage 1 type 1 diabetes and non-diabetic donors

Aims/hypothesis Progression to type 1 diabetes is associated with genetic factors, the presence of autoantibodies and a decline in beta cell insulin secretion in response to glucose. Very little is known regarding the molecular changes that occur in human insulin-secreting beta cells prior to the onset of type 1 diabetes. Herein, we applied an unbiased proteomics approach to identify changes in proteins and potential mechanisms of islet dysfunction in islet-autoantibody-positive organ donors with pre-symptomatic stage 1 type 1 diabetes (HbA1c ≤42 mmol/mol [6.0%]). We aimed to identify pathways in islets that are indicative of beta cell dysfunction. Methods Multiple islet sections were collected through laser microdissection of frozen pancreatic tissues from organ donors positive for single or multiple islet autoantibodies (AAb + , n=5), and age (±2 years)- and sex-matched non-diabetic (ND) control donors (n=5) obtained from the Network for Pancreatic Organ donors with Diabetes (nPOD). Islet sections were subjected to MS-based proteomics and analysed with label-free quantification followed by pathway and functional annotations. Results Analyses resulted in ~4500 proteins identified with low false discovery rate (<1%), with 2165 proteins reliably quantified in every islet sample. We observed large inter-donor variations that presented a challenge for statistical analysis of proteome changes between donor groups. We therefore focused on only the donors with stage 1 type 1 diabetes who were positive for multiple autoantibodies (mAAb + , n=3) and genetic risk compared with their matched ND controls (n=3) for the final statistical analysis. Approximately 10% of the proteins (n=202) were significantly different (unadjusted p<0.025, q<0.15) for mAAb + vs ND donor islets. The significant alterations clustered around major functions for upregulation in the immune response and glycolysis, and downregulation in endoplasmic reticulum (ER) stress response as well as protein translation and synthesis. The observed proteome changes were further supported by several independent published datasets, including a proteomics dataset from in vitro proinflammatory cytokine-treated human islets and single-cell RNA-seq datasets from AAb + individuals. Conclusions/interpretation In situ human islet proteome alterations in stage 1 type 1 diabetes centred around several major functional categories, including an expected increase in immune response genes (elevated antigen presentation/HLA), with decreases in protein synthesis and ER stress response, as well as compensatory metabolic response. The dataset serves as a proteomics resource for future studies on beta cell changes during type 1 diabetes progression and pathogenesis. Data availability The LC-MS raw datasets that support the findings of this study have been deposited in the online repository: MassIVE (https://massive.ucsd.edu/ProteoSAFe/static/massive.jsp) with accession no. MSV000090212.

Autoantibody-positive↗

Autoencoder-Based Anomaly Detection System for Online Data Quality Monitoring of the CMS Electromagnetic Calorimeter

The CMS detector is a general-purpose apparatus that detects high-energy collisions produced at the LHC. Online data quality monitoring of the CMS electromagnetic calorimeter is a vital operational tool that allows detector experts to quickly identify, localize, and diagnose a broad range of detector issues that could affect the quality of physics data. A real-time autoencoder-based anomaly detection system using semi-supervised machine learning is presented enabling the detection of anomalies in the CMS electromagnetic calorimeter data. A novel method is introduced which maximizes the anomaly detection performance by exploiting the time-dependent evolution of anomalies as well as spatial variations in the detector response. The autoencoder-based system is able to efficiently detect anomalies, while maintaining a very low false discovery rate. The performance of the system is validated with anomalies found in 2018 and 2022 LHC collision data. In addition, the first results from deploying the autoencoder-based system in the CMS online data quality monitoring workflow during the beginning of Run 3 of the LHC are presented, showing its ability to detect issues missed by the existing system.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Untargeted Spatial Metabolomics and Spatial Proteomics on the Same Tissue Section

An increasing number of spatial multiomic workflows have been recently developed. Some of these approaches have leveraged initial mass spectrometry imaging (MSI)-based spatial metabolomics to inform region of interest (ROI) selection for downstream spatial proteomics. However, these workflows have been limited by varied substrate requirements between modalities or have required analyzing serial sections (i.e., one section per modality). To mitigate these issues, we present a novel multiomic workflow that uses desorption electrospray ionization (DESI)-MSI to identify representative spatial metabolite patterns on-tissue prior to spatial proteomic analyses on the same tissue section. Further, this workflow is demonstrated here with a model mammalian tissue (coronal rat brain section) mounted on a polyethylene naphthalate-membrane slide. Initial DESI-MSI resulted in 160 annotations (SwissLipids) within to the METASPACE platform (≤20% false discovery rate). A segmentation map from the annotated ion images informed downstream ROI selection for spatial proteomics characterization from the same sample. The unspecific substrate requirements and minimal sample disruption inherent to DESI-MSI allowed for an optimized, downstream spatial proteomics assay, resulting in 3888 ± 240 to 4717 ± 48 proteins being confidently directed per ROI (200 µm x 200 µm). Finally, we demonstrate the integration of multiomic information, where we found ceramide localization to be correlated with SMPD3 abundance (ceramide synthesis protein), and we also utilized protein abundance to resolve metabolite isomeric ambiguity. Overall, the integration of DESI-MSI into the multiomic workflow allows for complementary spatial and molecular-level information to be achieved from optimized implementations of each MS assay inherent to the workflow itself.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Chemoproteogenomic stratification of the missense variant cysteinome

Abstract Cancer genomes are rife with genetic variants; one key outcome of this variation is widespread gain-of-cysteine mutations. These acquired cysteines can be both driver mutations and sites targeted by precision therapies. However, despite their ubiquity, nearly all acquired cysteines remain unidentified via chemoproteomics; identification is a critical step to enable functional analysis, including assessment of potential druggability and susceptibility to oxidation. Here, we pair cysteine chemoproteomics—a technique that enables proteome-wide pinpointing of functional, redox sensitive, and potentially druggable residues—with genomics to reveal the hidden landscape of cysteine genetic variation. Our chemoproteogenomics platform integrates chemoproteomic, whole exome, and RNA-seq data, with a customized two-stage false discovery rate (FDR) error controlled proteomic search, which is further enhanced with a user-friendly FragPipe interface. Chemoproteogenomics analysis reveals that cysteine acquisition is a ubiquitous feature of both healthy and cancer genomes that is further elevated in the context of decreased DNA repair. Reference cysteines proximal to missense variants are also found to be pervasive, supporting heretofore untapped opportunities for variant-specific chemical probe development campaigns. As chemoproteogenomics is further distinguished by sample-matched combinatorial variant databases and is compatible with redox proteomics and small molecule screening, we expect widespread utility in guiding proteoform-specific biology and therapeutic discovery.

Desai, Heta (ORCID:0000000343621707)↗

Functional protein mining with conformal guarantees

Molecular structure prediction and homology detection offer promising paths to discovering protein function and evolutionary relationships. However, current approaches lack statistical reliability assurances, limiting their practical utility for selecting proteins for further experimental and in-silico characterization. To address this challenge, we introduce a statistically principled approach to protein search leveraging principles from conformal prediction, offering a framework that ensures statistical guarantees with user-specified risk and provides calibrated probabilities (rather than raw ML scores) for any protein search model. Our method (1) lets users select many biologically-relevant loss metrics (i.e. false discovery rate) and assigns reliable functional probabilities for annotating genes of unknown function; (2) achieves state-of-the-art performance in enzyme classification without training new models; and (3) robustly and rapidly pre-filters proteins for computationally intensive structural alignment algorithms. Our framework enhances the reliability of protein homology detection and enables the discovery of uncharacterized proteins with likely desirable functional properties.

59 BASIC BIOLOGICAL SCIENCES↗

Anomaly Detection Based on Machine Learning for the CMS Electromagnetic Calorimeter Online Data Quality Monitoring

Using a semi-supervised machine learning approach we present a real-time anomaly detection system based on an autoencoder used for online data quality monitoring of the CMS electromagnetic calorimeter operating at the CERN LHC. We introduce a novel method that maximizes the anomaly detection performance making use of the time-dependence of anomalies and the spatial variations in the detector response. The autoencoder-based system efficiently detects anomalies in real time and maintains a very low false discovery rate. We validate the performance of this novel system with anomalies from LHC collision data taken in 2018 and 2022. In addition, results are presented after deploying the autoencoder-based system in the CMS online Data Quality Monitoring workflow at the beginning of LHC Run 3 resulting in the system to detect issues that were missed by the existing system.

Harilal, Abhirami [Carnegie Mellon University, Pit↗

Elastic Changepoint Detection for Globally-indexed Functional Time Series Data with Climate Applications

Changepoint detection is a vital tool in the application of climate data analysis. Numerous types of climate observation data are most properly represented by functional time series, implying a need for accurate changepoint detection methods applicable to functional time series data. Such data taken at a global scale often contain both spatial heterogeneity and dependence as well as phase (time) misalignment. In this report, we present methods which can detect spatially-dependent changepoints while allowing different estimates of change time and change strength depending on location. Additionally, we provide extensions to this spatially-predicted model which controls for phase variability among observations. Our methods provide the ability to detect a single change, or control for epidemic changes (where a “return-to-normal” change is more likely to be detected than the initial change). We showcase results analyzing the June 1991 eruption of Mt. Pinatubo, where our methods demonstrate the ability to accurately detect both single and epidemic changepoints even in the presence of strong seasonal variability. We find that our spatially-predicted model improves the detection of relevant changepoints versus methods which do not take spatial information into account, and we find that controlling for phase variability helps to control the false discovery rate during the detection process.

54 ENVIRONMENTAL SCIENCES↗

A framework to evaluate machine learning crystal stability predictions

The rapid adoption of machine learning in various scientific domains calls for the development of best practices and community agreed-upon benchmarking tasks and metrics. We present Matbench Discovery as an example evaluation framework for machine learning energy models, here applied as pre-filters to first-principles computed data in a high-throughput search for stable inorganic crystals. We address the disconnect between (1) thermodynamic stability and formation energy and (2) retrospective and prospective benchmarking for materials discovery. Alongside this paper, we publish a Python package to aid with future model submissions and a growing online leaderboard with adaptive user-defined weighting of various performance metrics allowing researchers to prioritize the metrics they value most. To answer the question of which machine learning methodology performs best at materials discovery, our initial release includes random forests, graph neural networks, one-shot predictors, iterative Bayesian optimizers and universal interatomic potentials. We highlight a misalignment between commonly used regression metrics and more task-relevant classification metrics for materials discovery. Accurate regressors are susceptible to unexpectedly high false-positive rates if those accurate predictions lie close to the decision boundary at 0 eV per atom above the convex hull. The benchmark results demonstrate that universal interatomic potentials have advanced sufficiently to effectively and cheaply pre-screen thermodynamic stable hypothetical materials in future expansions of high-throughput materials databases.

Riebesell, Janosh↗