Search NASASearch

SEARCH · Search NASA

Results for “omics data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

G2PDeep-v2: A Web-Based Deep-Learning Framework for Phenotype Prediction and Biomarker Discovery for All Organisms Using Multi-Omics Data

Multi-omics data offers rich insights into complex traits across organisms, yet integrating and analyzing these datasets for phenotype prediction and marker discovery remains challenging. Researchers need accessible tools that combine deep learning, hyperparameter optimization, visualization, and downstream analysis in a unified web platform. To address this, we developed G2PDeep-v2, a web-based platform powered by deep learning for phenotype prediction and marker discovery from multi-omics data across a wide range of organisms, including humans and plants. The server provides multiple services for researchers to create deep-learning models through an interactive interface and train these models using an automated hyperparameter tuning algorithm on high-performance computing resources. Users can visualize the results of phenotype and markers predictions and perform Gene Set Enrichment Analysis for the significant markers to provide insights into the molecular mechanisms underlying complex diseases, conditions and other biological phenotypes being studied.

59 BASIC BIOLOGICAL SCIENCES

MODE: A Web Application for Interactive Visualization and Exploration of Omics Data

Studies generating transcriptomics, proteomics, lipidomics, and metabolomics (colloquially referred to as “omics”) data allow researchers to find biomarkers or molecular targets, or understand complex biological structures and functions by identifying changes in biomolecule abundance and expression between experimental conditions. Omics data is multi-dimensional and oftentimes summarization techniques such as principal component analysis (PCA) are used to identify high-level patterns in data. Though useful, these summaries don’t allow exploration of detailed patterns in omics data that may have biological relevance. The use of interactive HTML displays with plots allows researchers to interact with omics data at a detailed level, but building these displays requires significant coding expertise. To overcome this barrier, the software MODE was built to empower users to build their own interactive HTML displays to support scientific discovery. These displays are easily shareable, do not depend on a specific operating system, and allow users to effortlessly sort and filter plots by categorical or numerical variables. MODE allows users to build and share these displays with several options for plot design and meta selection. In conclusion, the MODE web application and its capabilities are presented and then demonstrated on lipidomics data from a leaf wounding study.

lipidomics

OmicsMLMentor: A Web Application for Guided Machine Learning Analysis of Omics Data

Expression-based omics technologies (e.g. proteomics, metabolomics, transcriptomics, etc.) increasingly rely on supervised and unsupervised machine learning (ML) models to find key biomolecules distinguishing conditions, identify natural groupings in biological data, or generate predictions for outcomes of interest. Fitting ML models to omics data presents several challenges, including handling missing data, selecting a normalization method, choosing a valid model, and optimizing hyperparameters, all requiring statistical programming skills to address these challenges. Thus, the open-source web application SLOPE was designed to lower the barrier to ML modeling for omics data. SLOPE supports the fitting of 15 ML models (10 supervised and 5 unsupervised) tailored to omics datasets, such as proteomics, metabolomics, lipidomics, and transcriptomics. SLOPE offers several omics-specific features, including methods for handling missingness (imputation, conversion, removal), normalization tests, ranking of models based on the structure of a user’s data and user input, and optimal hyperparameter selections using cross-validation splits. By streamlining ML workflows for omics analysis, SLOPE address critical gaps in existing online web tools, facilitating a broader adoption of these models for omics research. Here, SLOPE is applied to data from a lignin exposure study to highlight the workflow for fitting both supervised and unsupervised models to data.

lipidomics

Prediction of plant complex traits via integration of multi-omics data

The formation of complex traits is the consequence of genotype and activities at multiple molecular levels. However, connecting genotypes and these activities to complex traits remains challenging. Here, we investigate whether integrating genomic, transcriptomic, and methylomic data can improve prediction for six Arabidopsis traits. We find that transcriptome- and methylome-based models have performances comparable to those of genome-based models. However, models built for flowering time using different omics data identify different benchmark genes. Nine additional genes identified as important for flowering time from our models are experimentally validated as regulating flowering. Gene contributions to flowering time prediction are accession-dependent and distinct genes contribute to trait prediction in different genotypes. Models integrating multi-omics data perform best and reveal known and additional gene interactions, extending knowledge about existing regulatory networks underlying flowering time determination. These results demonstrate the feasibility of revealing molecular mechanisms underlying complex traits through multi-omics data integration.

59 BASIC BIOLOGICAL SCIENCES

Computationally efficient Bayesian estimation of graphical networks for omics data

Graphical networks are useful, widely-used modeling approaches to represent complex biological processes with biological measurements generated by platforms such as mass spectrometry. Bayesian analyses of graphical networks for omics data have several advantages over their frequentist counterparts, such as the inclusion of prior knowledge in the estimation of models. However, Bayesian approaches to date have only been feasible for data with a couple hundred biomolecules due to prohibitive computational time, but omics data often contains tens of thousands of biomolecules. Here, we present and illustrate a more computationally efficient approach named BPlane (Bayesian PseudoLikelihood-based Algorithm for Network Estimation) to extend Bayesian modeling capabilities for larger-sized datasets, such as most untargeted proteomics data. Via simulation, we demonstrate that BPlane produces substantial computational savings over a current state-of-the-art Bayesian algorithm while maintaining competitive edge detection accuracy. On a SARS-CoV2 proteomics data with 7000 proteins, the competing algorithm takes three times as long to complete the first iteration as BPlane takes to converge after over 100 iterations.

EM algorithm

Multi-omics data resource: Data package 25 (Pck025)

This data package comprises omics datasets from human pancreatic islets treated with IL-1β + IFNγ or with estrogen (E2) for 18 h. Two RNA-seq datasets are available: the first is a discovery dataset involving human islets treated with or without IL-1β + IFNγ for 18 hours; the second is a validation dataset, where human islets are treated with or without IL-1β + IFNγ or E2 for 18 hours. DIA proteomic analysis was performed on the same validation dataset samples. Data contributors: Kiersten L. Webster, Sarah Tersey & Raghavendra G. Mirmir: Kovler Diabetes Center and Department of Medicine, The University of Chicago, Chicago, IL, 60637, USA. Soumyadeep Sarkar, Raghavendra Mirmira, Ernesto S. Nakayasu: Biological Sciences Division, Pacific Northwest National Laboratory, Richland, WA, 99354, USA. Data repository: RNA-seq: GSE310965 Proteomics: MSV000101892 Publication: PMID 41279069

Sarkar, Soumyadeep [Pacific Northwest National Lab

HighDimMixedModels.jl: Robust high-dimensional mixed-effects models across omics data

High-dimensional mixed-effects models are an increasingly important form of regression in which the number of covariates rivals or exceeds the number of samples, which are collected in groups or clusters. The penalized likelihood approach to fitting these models relies on a coordinate descent algorithm that lacks guarantees of convergence to a global optimum. Here, we empirically study the behavior of this algorithm on simulated and real examples of three types of data that are common in modern biology: transcriptome, genome-wide association, and microbiome data. Our simulations provide new insights into the algorithm’s behavior in these settings, and, comparing the performance of two popular penalties, we demonstrate that the smoothly clipped absolute deviation (SCAD) penalty consistently outperforms the least absolute shrinkage and selection operator (LASSO) penalty in terms of both variable selection and estimation accuracy across omics data. To empower researchers in biology and other fields to fit models with the SCAD penalty, we implement the algorithm in a Julia package, HighDimMixedModels.jl .

Gorstein, Evan

Finding the missing pieces: filling gaps that impede the translation of omics data into models

High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.

59 BASIC BIOLOGICAL SCIENCES

Ranking Biological Features in Soil-Based Microbial Multi-Omics Data with Integration Modeling

Distinguishing the most important features (e.g. proteins, metabolites, etc.) per group (e.g. control and treatment) is a critical challenge in feature-rich multi-omics experiments, especially in soil data. Traditional feature identification and ranking approaches, such as differential expression, are based on single omics and thus not directly translatable to multi-omics experiments. Here, 5 multi-omics integration models (DIABLO, JACA, MOFA, MultiMLP, and SLIDE) that were not explicitly built for soil data applications were tested using a soil-based multi-omics experiment. The data were obtained from an experimental setup of an autoclaved soil system inoculated with 8 bacteria and using chitin as the carbon source and including samples collected at 0- (control), 4-, 8-, and 12-weeks post-inoculation. The omics data included metaproteomics, 16S rRNA sequencing, and LC-MS/MS metabolomics (in positive and negative mode). Each multi-omics integration model was implemented, and top features were compared to differential univariate statistics per omic type, demonstrating that integration approaches cut the potential number of top features from 2957 identified by differential statistics to 13-224 (a 99.6% to 92.4% reduction). Interestingly, most top features across integration models were not shared; though, scaling and averaging ranks across models shared similar patterns. This work highlights the usefulness of multi-omics integration models in soil-based microbial studies and the power of using multiple integration models together to interpret results.

54 ENVIRONMENTAL SCIENCES

Multi-omics data resource: Data package 22 (Pck022)

In type 1 diabetes (T1D), autoimmune response and inflammation cause the death of pancreatic ß cells, leading to the body’s inability to produce insulin and maintain glucose homeostasis. This process is at least in part mediated by pro-inflammatory cytokines, such as interferon (IFN)a, IFN?, interleukin (IL)-1ß, and tumor necrosis factor (TNF)a, which induce ß-cell dysfunction and apoptosis. A deep understanding of the ß-cell signaling and regulatory networks induced by these cytokines could lead to the identification of therapeutic targets to prevent T1D development. To study cytokine-mediated islets/ß-cell signaling and regulatory networks, a variety of omics experiments have been conducted, including transcriptomics, epigenomics (DNA methylation, UMI-4C, ATAC-seq & ChIP-seq), proteomics (bottom-up, top-down, post-translational modification analysis), lipidomics, and metabolomics. The combination of these datasets can be instrumental in identifying signaling components and regulatory factors involved in ß-cell stress/death. Here, we aggregated these multiple omics datasets into a centralized location, providing a quality-controlled and statistically rigorous resource for investigators seeking to holistically study ß-cell regulation by pro-inflammatory cytokines. The data package consists of isolated pancreatic islets from adult male C57BL6/J mice treated with IL-1β, IFNγ or IL-1β + IFNγ for 6 h and submitted for scRNA-seq. This study focused on understanding the heterogeneity of the cytokine-mediated response. Data contributors: Jennifer S Stancill & John A Corbett: Department of Biochemistry, Medical College of Wisconsin, Milwaukee, WI, USA Data repository: GSE156175 Publication: 10.26508/lsa.202000949

Sarkar, Soumyadeep [Pacific Northwest National Lab

Multi-omics data compendium: Data package 10 (Pck010)

In type 1 diabetes (T1D), autoimmune response and inflammation cause the death of pancreatic ß cells, leading to the body’s inability to produce insulin and maintain glucose homeostasis. This process is at least in part mediated by pro-inflammatory cytokines, such as interferon (IFN)a, IFN?, interleukin (IL)-1ß, and tumor necrosis factor (TNF)a, which induce ß-cell dysfunction and apoptosis. A deep understanding of the ß-cell signaling and regulatory networks induced by these cytokines could lead to the identification of therapeutic targets to prevent T1D development. To study cytokine-mediated islets/ß-cell signaling and regulatory networks, a variety of omics experiments have been conducted, including transcriptomics, epigenomics (DNA methylation, UMI-4C, ATAC-seq & ChIP-seq), proteomics (bottom-up, top-down, post-translational modification analysis), lipidomics, and metabolomics. The combination of these datasets can be instrumental in identifying signaling components and regulatory factors involved in ß-cell stress/death. Here, we aggregated these multiple omics datasets into a centralized location, providing a quality-controlled and statistically rigorous resource for investigators seeking to holistically study ß-cell regulation by pro-inflammatory cytokines.

Sarkar, Soumyadeep [Pacific Northwest National Lab

Multi-omics data compendium: Data package 11 (Pck011)

In type 1 diabetes (T1D), autoimmune response and inflammation cause the death of pancreatic ß cells, leading to the body’s inability to produce insulin and maintain glucose homeostasis. This process is at least in part mediated by pro-inflammatory cytokines, such as interferon (IFN)a, IFN?, interleukin (IL)-1ß, and tumor necrosis factor (TNF)a, which induce ß-cell dysfunction and apoptosis. A deep understanding of the ß-cell signaling and regulatory networks induced by these cytokines could lead to the identification of therapeutic targets to prevent T1D development. To study cytokine-mediated islets/ß-cell signaling and regulatory networks, a variety of omics experiments have been conducted, including transcriptomics, epigenomics (DNA methylation, UMI-4C, ATAC-seq & ChIP-seq), proteomics (bottom-up, top-down, post-translational modification analysis), lipidomics, and metabolomics. The combination of these datasets can be instrumental in identifying signaling components and regulatory factors involved in ß-cell stress/death. Here, we aggregated these multiple omics datasets into a centralized location, providing a quality-controlled and statistically rigorous resource for investigators seeking to holistically study ß-cell regulation by pro-inflammatory cytokines.

Sarkar, Soumyadeep [Pacific Northwest National Lab

Multi-omics data compendium: Data package 12 (Pck012)

In type 1 diabetes (T1D), autoimmune response and inflammation cause the death of pancreatic ß cells, leading to the body’s inability to produce insulin and maintain glucose homeostasis. This process is at least in part mediated by pro-inflammatory cytokines, such as interferon (IFN)a, IFN?, interleukin (IL)-1ß, and tumor necrosis factor (TNF)a, which induce ß-cell dysfunction and apoptosis. A deep understanding of the ß-cell signaling and regulatory networks induced by these cytokines could lead to the identification of therapeutic targets to prevent T1D development. To study cytokine-mediated islets/ß-cell signaling and regulatory networks, a variety of omics experiments have been conducted, including transcriptomics, epigenomics (DNA methylation, UMI-4C, ATAC-seq & ChIP-seq), proteomics (bottom-up, top-down, post-translational modification analysis), lipidomics, and metabolomics. The combination of these datasets can be instrumental in identifying signaling components and regulatory factors involved in ß-cell stress/death. Here, we aggregated these multiple omics datasets into a centralized location, providing a quality-controlled and statistically rigorous resource for investigators seeking to holistically study ß-cell regulation by pro-inflammatory cytokines.

Sarkar, Soumyadeep [Pacific Northwest National Lab

Multi-omics data compendium: Data package 13 (Pck013)

In type 1 diabetes (T1D), autoimmune response and inflammation cause the death of pancreatic ß cells, leading to the body’s inability to produce insulin and maintain glucose homeostasis. This process is at least in part mediated by pro-inflammatory cytokines, such as interferon (IFN)a, IFN?, interleukin (IL)-1ß, and tumor necrosis factor (TNF)a, which induce ß-cell dysfunction and apoptosis. A deep understanding of the ß-cell signaling and regulatory networks induced by these cytokines could lead to the identification of therapeutic targets to prevent T1D development. To study cytokine-mediated islets/ß-cell signaling and regulatory networks, a variety of omics experiments have been conducted, including transcriptomics, epigenomics (DNA methylation, UMI-4C, ATAC-seq & ChIP-seq), proteomics (bottom-up, top-down, post-translational modification analysis), lipidomics, and metabolomics. The combination of these datasets can be instrumental in identifying signaling components and regulatory factors involved in ß-cell stress/death. Here, we aggregated these multiple omics datasets into a centralized location, providing a quality-controlled and statistically rigorous resource for investigators seeking to holistically study ß-cell regulation by pro-inflammatory cytokines.

Sarkar, Soumyadeep [Pacific Northwest National Lab

Multi-omics data compendium: Data package 14 (Pck014)

In type 1 diabetes (T1D), autoimmune response and inflammation cause the death of pancreatic ß cells, leading to the body’s inability to produce insulin and maintain glucose homeostasis. This process is at least in part mediated by pro-inflammatory cytokines, such as interferon (IFN)a, IFN?, interleukin (IL)-1ß, and tumor necrosis factor (TNF)a, which induce ß-cell dysfunction and apoptosis. A deep understanding of the ß-cell signaling and regulatory networks induced by these cytokines could lead to the identification of therapeutic targets to prevent T1D development. To study cytokine-mediated islets/ß-cell signaling and regulatory networks, a variety of omics experiments have been conducted, including transcriptomics, epigenomics (DNA methylation, UMI-4C, ATAC-seq & ChIP-seq), proteomics (bottom-up, top-down, post-translational modification analysis), lipidomics, and metabolomics. The combination of these datasets can be instrumental in identifying signaling components and regulatory factors involved in ß-cell stress/death. Here, we aggregated these multiple omics datasets into a centralized location, providing a quality-controlled and statistically rigorous resource for investigators seeking to holistically study ß-cell regulation by pro-inflammatory cytokines.

Sarkar, Soumyadeep [Pacific Northwest National Lab

Multi-omics data compendium: Data package 15 (Pck015)

In type 1 diabetes (T1D), autoimmune response and inflammation cause the death of pancreatic ß cells, leading to the body’s inability to produce insulin and maintain glucose homeostasis. This process is at least in part mediated by pro-inflammatory cytokines, such as interferon (IFN)a, IFN?, interleukin (IL)-1ß, and tumor necrosis factor (TNF)a, which induce ß-cell dysfunction and apoptosis. A deep understanding of the ß-cell signaling and regulatory networks induced by these cytokines could lead to the identification of therapeutic targets to prevent T1D development. To study cytokine-mediated islets/ß-cell signaling and regulatory networks, a variety of omics experiments have been conducted, including transcriptomics, epigenomics (DNA methylation, UMI-4C, ATAC-seq & ChIP-seq), proteomics (bottom-up, top-down, post-translational modification analysis), lipidomics, and metabolomics. The combination of these datasets can be instrumental in identifying signaling components and regulatory factors involved in ß-cell stress/death. Here, we aggregated these multiple omics datasets into a centralized location, providing a quality-controlled and statistically rigorous resource for investigators seeking to holistically study ß-cell regulation by pro-inflammatory cytokines.

Sarkar, Soumyadeep [Pacific Northwest National Lab

Multi-omics data compendium: Data package 16 (Pck016)

In type 1 diabetes (T1D), autoimmune response and inflammation cause the death of pancreatic ß cells, leading to the body’s inability to produce insulin and maintain glucose homeostasis. This process is at least in part mediated by pro-inflammatory cytokines, such as interferon (IFN)a, IFN?, interleukin (IL)-1ß, and tumor necrosis factor (TNF)a, which induce ß-cell dysfunction and apoptosis. A deep understanding of the ß-cell signaling and regulatory networks induced by these cytokines could lead to the identification of therapeutic targets to prevent T1D development. To study cytokine-mediated islets/ß-cell signaling and regulatory networks, a variety of omics experiments have been conducted, including transcriptomics, epigenomics (DNA methylation, UMI-4C, ATAC-seq & ChIP-seq), proteomics (bottom-up, top-down, post-translational modification analysis), lipidomics, and metabolomics. The combination of these datasets can be instrumental in identifying signaling components and regulatory factors involved in ß-cell stress/death. Here, we aggregated these multiple omics datasets into a centralized location, providing a quality-controlled and statistically rigorous resource for investigators seeking to holistically study ß-cell regulation by pro-inflammatory cytokines.

Sarkar, Soumyadeep [Pacific Northwest National Lab

Multi-omics data compendium: Data package 17 (Pck017)

In type 1 diabetes (T1D), autoimmune response and inflammation cause the death of pancreatic ß cells, leading to the body’s inability to produce insulin and maintain glucose homeostasis. This process is at least in part mediated by pro-inflammatory cytokines, such as interferon (IFN)a, IFN?, interleukin (IL)-1ß, and tumor necrosis factor (TNF)a, which induce ß-cell dysfunction and apoptosis. A deep understanding of the ß-cell signaling and regulatory networks induced by these cytokines could lead to the identification of therapeutic targets to prevent T1D development. To study cytokine-mediated islets/ß-cell signaling and regulatory networks, a variety of omics experiments have been conducted, including transcriptomics, epigenomics (DNA methylation, UMI-4C, ATAC-seq & ChIP-seq), proteomics (bottom-up, top-down, post-translational modification analysis), lipidomics, and metabolomics. The combination of these datasets can be instrumental in identifying signaling components and regulatory factors involved in ß-cell stress/death. Here, we aggregated these multiple omics datasets into a centralized location, providing a quality-controlled and statistically rigorous resource for investigators seeking to holistically study ß-cell regulation by pro-inflammatory cytokines.

Sarkar, Soumyadeep [Pacific Northwest National Lab