Search NASA⌕ Search

SEARCH · Search NASA

Results for “bioinformatics workflows”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Challenges in Bioinformatics Workflows for Processing Microbiome Omics Data at Scale

The nascent field of microbiome science is transitioning from a descriptive approach of cataloging taxa and functions present in an environment to applying multi-omics methods to investigate microbiome dynamics and function. A large number of new tools and algorithms have been designed and used for very specific purposes on samples collected by individual investigators or groups. While these developments have been quite instructive, the ability to compare microbiome data generated by many groups of researchers is impeded by the lack of standardized application of bioinformatics methods. Additionally, there are few examples of broad bioinformatics workflows that can process metagenome, metatranscriptome, metaproteome and metabolomic data at scale, and no central hub that allows processing, or provides varied omics data that are findable, accessible, interoperable and reusable (FAIR). Here, we review some of the challenges that exist in analyzing omics data within the microbiome research sphere, and provide context on how the National Microbiome Data Collaborative has adopted a standardized and open access approach to address such challenges.

NMDC, Microbiome↗

Enabling Open and Interoperable Science: Multi-Omics Data Processing Platform with NASA GeneLab Standardized Bioinformatics Workflows for Space and Earth Research

Multi-omics biological data continues to be generated at an astounding pace. Genomics, transcriptomics, metabolomics, and proteomics, or collectively known as multi-omics data, are used to assess biological functions, and provide invaluable insights into human, animal, plant, and environmental health both on Earth and in Space. Despite the abundance of these valuable data, the need for bioinformatics expertise, particularly as it relates to the niche filed of space biology, and a lack of accessible resources for processing these data limit their usefulness in deriving biological insights. The NASA Open Science Data Repository (OSDR) provides access to omics data from various spaceflight and analog studies. To enhance the accessibility and reusability of these data, GeneLab (part of OSDR) designs and implements standardized, community-driven, open-source bioinformatics workflows to transform raw omics data into standardized processed data. Currently, GeneLab-processed data from hundreds of space studies have been reused for meta-analyses. This has led to new insights and scientific publications that extend beyond the initial research, thereby enriching our understanding of molecular-scale biological responses to the space environment. To make these bioinformatics workflows open and accessible, GeneLab teamed up with DOE-funded initiatives, including the National Microbiome Data Collaborative (NMDC), to create the NASA EDGE [Empowering the Development of Genomics Expertise] Bioinformatics web-based platform. NASA EDGE utilizes shared compute resources to run the GeneLab standardized bioinformatics workflows, which eliminates the need for researchers to have their own high performance computing cluster. The web-based platform makes complicated biological analyses incredibly easy to perform, thus expanding the reach of these analyses to bioinformatics novices, students, and even citizen scientists enabling them to contribute to scientific discoveries and progress. The authors will demonstrate how the NASA EDGE platform can be used to process microbial omics data hosted on OSDR as well as user-generated omics datasets using GeneLab’s standard workflows.

Amanda M. Saravia-Butler↗

Recommendations for Uniform Variant Calling of SARS-CoV-2 Genome Sequence across Bioinformatic Workflows

Genomic sequencing of clinical samples to identify emerging variants of SARS-CoV-2 has been a key public health tool for curbing the spread of the virus. As a result, an unprecedented number of SARS-CoV-2 genomes were sequenced during the COVID-19 pandemic, which allowed for rapid identification of genetic variants, enabling the timely design and testing of therapies and deployment of new vaccine formulations to combat the new variants. However, despite the technological advances of deep sequencing, the analysis of the raw sequence data generated globally is neither standardized nor consistent, leading to vastly disparate sequences that may impact identification of variants. Here, we show that for both Illumina and Oxford Nanopore sequencing platforms, downstream bioinformatic protocols used by industry, government, and academic groups resulted in different virus sequences from same sample. These bioinformatic workflows produced consensus genomes with differences in single nucleotide polymorphisms, inclusion and exclusion of insertions, and/or deletions, despite using the same raw sequence as input datasets. Here, we compared and characterized such discrepancies and propose a specific suite of parameters and protocols that should be adopted across the field. Consistent results from bioinformatic workflows are fundamental to SARS-CoV-2 and future pathogen surveillance efforts, including pandemic preparation, to allow for a data-driven and timely public health response.

60 APPLIED LIFE SCIENCES↗

Marine Microeukaryote Metatranscriptomics: Sample Processing and Bioinformatic Workflow Recommendations for Ecological Applications

Microeukaryotes (protists) serve fundamental roles in the marine environment as contributors to biogeochemical nutrient cycling and ecosystem function. Their activities can be inferred through metatranscriptomic investigations, which provide a detailed view into cellular processes, chemical-biological interactions in the environment, and ecological relationships among taxonomic groups. Established workflows have been individually put forth describing biomass collection at sea, laboratory RNA extraction protocols, and bioinformatic processing and computational approaches. Here, we present a compilation of current practices and lessons learned in carrying out metatranscriptomics of marine pelagic protistan communities, highlighting effective strategies and tools used by practitioners over the past decade. We anticipate that these guidelines will serve as a roadmap for new marine scientists beginning in the realms of molecular biology and/or bioinformatics, and will equip readers with foundational principles needed to delve into protistan metatranscriptomics.

Cohen, Natalie R.↗

An empirical DNA-based identification of morphologically similar snappers (Lutjanus campechanus, Lutjanus purpureus) using a versatile bioinformatics workflow for the discovery and analysis of informative single-nucleotide polymorphisms

The commercially important speciesLutjanus campechanus(Northern/Gulf red snapper) andLutjanus purpureus(Southern/Caribbean red snapper) are the protagonists of a decade’s long taxonomic debate over their species delimitation, due in part to partial habitat overlap, extensive morphological similarity, and the lack of resolution when applying canonically reliable DNA barcoding approaches. In this study, we leveraged publicly available RAD-Seq data forL. campechanusandL. purpureusto identify species-informative single‐nucleotide polymorphisms (SNPs) at the genome scale that were successful in distinguishing the Northern and Southern red snappers, while also detecting individuals exhibiting introgression. This 4-step empirical approach demonstrates the value of applying novel bioinformatics pipelines to existing genome-scale data to maximize the distillation of informative subsets. Our results facilitate economically relevant species identification in addition to confirming or challenging species identifications for specimens with data in public databases. These findings and their applications will benefit future sustainability strategies and broader research questions surrounding these overfished and evolutionarily entangled snapper species.

Environmental Sciences & Ecology↗

Accelerating Scientific Workflows on HPC Platforms with In Situ Processing

Scientific workflows drive most modern large-scale science breakthroughs by allowing scientists to define their computations as a set of jobs executed in a given order based on their data dependencies. Workflow management systems (WMSs) have become key to automating scientific workflows-executing computational jobs and orchestrating data transfers between those jobs running on complex high-performance computing (HPC) platforms. Traditionally, WMSs use files to communicate between jobs: a job writes out files that are read by other jobs. However, HPC machines face a growing gap between their storage and compute capabilities. To address that concern, the scientific community has adopted a new approach called in situ, which bypasses costly parallel filesystem I/O operations with faster in-memory or in-network communications. When using in situ approaches, communication and computations can be interleaved. In this work, we leverage the Decaf in situ dataflow framework to accelerate task-based scientific workflows managed by the Pegasus WMS, by replacing file communications with faster MPI messaging. We propose a new execution engine that uses Decaf to manage communications within a sub-workflow (i.e., set of jobs) to optimize inter-job communications. We consider two workflows in this study: (i) a synthetic workflow that benchmarks and compares file- and MPI-based communication; and (ii) a realistic bioinformatics workflow that computes mu-tational overlaps in the human genome. Experiments show that in situ communication can improve the bioinformatics workflow execution time by 22% to 30% compared with file communication. Our results motivate further opportunities and challenges for bridging traditional WMSs with in situ frameworks.

Decaf↗

EDGE COVID-19: a web platform to generate submission-ready genomes from SARS-CoV-2 sequencing efforts

Abstract Summary Genomics has become an essential technology for surveilling emerging infectious disease outbreaks. A range of technologies and strategies for pathogen genome enrichment and sequencing are being used by laboratories worldwide, together with different and sometimes ad hoc, analytical procedures for generating genome sequences. A fully integrated analytical process for raw sequence to consensus genome determination, suited to outbreaks such as the ongoing COVID-19 pandemic, is critical to provide a solid genomic basis for epidemiological analyses and well-informed decision making. We have developed a web-based platform and integrated bioinformatic workflows that help to provide consistent high-quality analysis of SARS-CoV-2 sequencing data generated with either the Illumina or Oxford Nanopore Technologies (ONT). Using an intuitive web-based interface, this workflow automates data quality control, SARS-CoV-2 reference-based genome variant and consensus calling, lineage determination and provides the ability to submit the consensus sequence and necessary metadata to GenBank, GISAID and INSDC raw data repositories. We tested workflow usability using real world data and validated the accuracy of variant and lineage analysis using several test datasets, and further performed detailed comparisons with results from the COVID-19 Galaxy Project workflow. Our analyses indicate that EC-19 workflows generate high-quality SARS-CoV-2 genomes. Finally, we share a perspective on patterns and impact observed with Illumina versus ONT technologies on workflow congruence and differences. Availability and implementation https://edge-covid19.edgebioinformatics.org, and https://github.com/LANL-Bioinformatics/EDGE/tree/SARS-CoV2. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Development of Computational Environmental Microbiome Workflows for the Laboratory and the International Space Station

Identification of microorganisms in the spaceflight environment is critical for crew health risk assessment on the International Space Station (ISS). Since 2017, nanopore sequencing technology has been used to support thein situ identification of microbial species during spaceflight. Beginning in 2018, a culture-independent, swab-to-sequencer method was implemented onboard the ISS to provide a more thorough insight of the ISS microbiome. Eliminating microbial culture enables identification of difficult-to-culture organisms, reduces risks associated with potentially pathogenic cultures, and could significantly reduce the time from sample-to-answer. However, this molecular-based approach generates large metagenomic datasets that require substantial computational resources for analysis. To process nanopore-generated sequencing data, the JSC Microbiology Laboratory established a bioinformatics workflow on Amazon EC2 under the security guidance of the NASA Science Managed Cloud Environment (SMCE).This resource allows for the development, testing, and accessing of computational tools for processing large and complex datasets. The work described here will address the downlinking of data from the ISS, the automated pipeline developed to identify targeted bacterial and fungal organisms, and the time from sampling onboard to microbial identification. The pipelines have been enhanced to address high and low biomass samples using optimization based on sample source (air, water, or surface) and type of collection (filter, colony, or swab).The resulting microbiome data can be assessed beyond microbial identifications to gain understanding toward population changes over time, potential selective environmental pressures, and evaluating correlations with a wide range of additional data sets. Metagenome analysis pipelines in development could allow for simultaneous identification of microbial species, gene function, and gene pathways present in the environment. Beyond the ground processing, the developed analysis pipeline is currently deployed onboard the ISS to allow for near real-time assessments of the ISS microbiome. This study serves as a critical foundation for exploration missions, where rapid microbiome analyses will be required.

G. Marie Sharp↗

DOE BSSD Performance Management Metrics Report Q1

Microbes play key roles in our biosphere, from driving global nutrient cycling to impacting plant, animal and human health and disease. Complex data from microbial genomes, proteins, and metabolites provide a window into these tiny engines that drive life on our planet. Yet these data are dispersed among researchers’ laboratories and various repositories, making it difficult to access. This calls for new ways of managing data, improving data interoperability, advancing community standards, and creating an infrastructure where data are shared efficiently. We have built the National Microbiome Data Collaborative (NMDC) to advance how scientists create, use, and reuse data to redefine the way we understand and harness the power of microbes. The vision of the National Microbiome Data Collaborative (NMDC) is to drive a microbiome data sharing network connecting data, people, and ideas to advance microbiome innovation and discovery. The NMDC was launched in 2019 and brought together DOE National Laboratories to collaborate across resources, capabilities, and expertise. The NMDC team was strategically assembled to include software developers, microbial researchers, metadata experts, and multi-omics specialists. The diversity of the NMDC team reflects the inherently interdisciplinary nature of microbiome science, and we leverage the strengths of the DOE National Laboratory system. Towards BER’s goal of advancing an iterative systems biology approach to the understanding of microbial genomes, the NMDC serves as a foundation for infrastructure, data standards, and community building. Together with the flagship DOE User Facilities, the Joint Genome Institute (JGI) and the Environmental Molecular Sciences Laboratory (EMSL), we are developing core capabilities in metadata standards for environmental descriptors and sample handling and processing; standardized bioinformatic workflows; an interface for data search and access; and robust community engagement activities. The NMDC production platform supports long-term data infrastructure and community building for BER’s bioenergy and environmental research goals. Our approach leverages lessons learned and an ambitious framework for collaborative, interdisciplinary data infrastructure to support microbiome research. The NMDC supports data, information, and knowledge access through three defined software tools – the Submission Portal, NMDC EDGE, and the Data Portal – driven by community needs. Herein, we describe the value proposition for the microbiome research community, our overarching strategy, and challenges and opportunities for developing the NMDC as both an infrastructure and community engagement program.

59 BASIC BIOLOGICAL SCIENCES↗

Modularization of EDGE Workflows Using Nextflow: Improving the Efficiency and Maintainability of Bioinformatics Software

EDGE is a bioinformatics platform developed in 2016 by researchers at Los Alamos National Laboratory (LANL) to facilitate the analysis of next-generation sequencing data by researchers with varying levels of experience in bioinformatics (Li et al., 2017). Users with single-end, paired-end or long-read sequencing data can provide their reads as input to EDGE and select the combination of workflows to run that are most useful for their research (e.g., quality control of reads, genome assembly, or the taxonomic classification of input reads). Table 1 summarizes the modules available in EDGE. EDGE is available as a web platform at https://edgebioinformatics.org, as installable source code maintained on GitHub under a GPLv3 license, and as a publicly hosted Docker image.

59 BASIC BIOLOGICAL SCIENCES↗

Expanded Coverage of Phytocompounds by Mass Spectrometry Imaging Using On-Tissue Chemical Derivatization by 4-APEBA

Probing the entirety of any species metabolome is an analytical grand challenge, especially at a cellular scale. Where spatial metabolomics, completed primarily by matrix-assisted laser desorption/ionization (MALDI), has limited molecular coverage for several reasons. To expand the scope of spatial metabolomics, we developed an on-tissue chemical derivatization (OTCD) workflow using 4-APEBA for confident identification of several dozen elusive phytocompounds, including several phytohormones, which have various roles within stress responses and cellular communication. Superiority of 4-APEBA is established in comparison to other derivatization agents with (1) broad specificity towards carbonyls, (2) low background, and (3) introduction of bromine isotopes, where the latter two facilitate confident bioinformatics. In conclusion, the outlined workflow trailblazes a path towards spatial hormonomics within plant samples, enhancing detection of carboxylates, aldehydes, ketones, and plausibly phenols.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Critical Assessment of MetaProteome Investigation (CAMPI): a multi-laboratory comparison of established workflows

Metaproteomics has matured into a powerful tool to assess functional interactions in microbial communities. While many metaproteomic workflows are available, the impact of method choice on results remains unclear. Here, we carry out a community-driven, multi-laboratory comparison in metaproteomics: the critical assessment of metaproteome investigation study (CAMPI). Based on well-established workflows, we evaluate the effect of sample preparation, mass spectrometry, and bioinformatic analysis using two samples: a simplified, laboratory-assembled human intestinal model and a human fecal sample. We observe that variability at the peptide level is predominantly due to sample processing workflows, with a smaller contribution of bioinformatic pipelines. These peptide-level differences largely disappear at the protein group level. While differences are observed for predicted community composition, similar functional profiles are obtained across workflows. CAMPI demonstrates the robustness of present-day metaproteomics research, serves as a template for multi-laboratory studies in metaproteomics, and provides publicly available data sets for benchmarking future developments.

59 BASIC BIOLOGICAL SCIENCES↗

Earth Science Data Processing With Nextflow

Earth science data processing tasks present many challenges. These tasks often process large input datasets and require scores of CPU-hours to generate results. All but the simplest tasks will be decomposed into a series of computational or data manipulation steps, also known as a scientific workflow. In order to reduce the burden of orchestrating and running the dependent processing steps, a workflow execution engine is required. This poster describes the lessons learned by the CLARREO Pathfinder (CPF) team while developing multiple scientific workflows and utilizing the open-source Nextflow engine to execute them in a cloud computing environment. The Nextflow engine is designed with the following stated goals: first, the engine does not dictate how individual steps in the task are implemented (i.e. it is language and interface agnostic); second, the engine supports easy configuration and modularity at the workflow level so that others can easily execute our workflows to reproduce results; lastly, the engine eases development by transparently scaling execution from local to remote environments. Nextflow was developed for the bioinformatics domain but is a good fit for other scientific workflows where the overall task is well-described by a dataflow diagram. The CPF team has developed Nextflow pipelines (i.e. scientific workflows) to simulate CLARREO radiance, generate large look-up tables for inter-calibration algorithms, and generate L4 intercalibration data products. These pipelines consume from single-digits to hundreds of thousands of CPU-hours. In the development and evolution of these pipelines we have discovered many design patterns, pitfalls, and solutions to common problems. Our goal is to demonstrate important aspects of how to design, implement, run, and ultimately share Nextflow pipelines in the domain of Earth science.

Aron D Bartle↗

Toward designing effective exascale scientific computing workflows: experiences and best practices

Many fields within scientific computing have embraced advances in big-data analysis and machine learning, which often requires the deployment of large, distributed and complicated workflows that may combine training neural networks, performing simulations, running inference, and performing database queries and data analysis in asynchronous, parallel and pipelined execution frameworks. Such a shift has brought into focus the need for scalable, efficient workflow management solutions with reproducibility, error and provenance handling, traceability, and checkpoint-restart capabilities, among other needs. Here, we discuss challenges and best-practices for deploying exascale-generation computational science workflows on resources at the Oak Ridge Leadership Computing Facility (OLCF). We present our experiences with large-scale deployment of distributed workflows on the Summit supercomputer, including for bioinformatics and computational biophysics, materials science, and deep learning model optimization. We also present problems and solutions created by working within a Python-centric software base on traditional HPC systems, and discuss steps that will be required before the convergence of HPC, AI, and data science can be fully realized. Our results point to a wealth of exciting new possibilities for harnessing this convergence to tackle new scientific challenges.

Coletti, Mark↗

Performance of methods for SARS-CoV-2 variant detection and abundance estimation within mixed population samples

The accurate identification of SARS-CoV-2 (SC2) variants and estimation of their abundance in mixed population samples (e.g., air or wastewater) is imperative for successful surveillance of community level trends. Assessing the performance of SC2 variant composition estimators (VCEs) should improve our confidence in public health decision making. Here, we introduce a linear regression based VCE and compare its performance to four other VCEs: two re-purposed DNA sequence read classifiers (Kallisto and Kraken2), a maximum-likelihood based method (Lineage deComposition for Sars-Cov-2 pooled samples (LCS)), and a regression based method (Freyja). We simulated DNA sequence datasets of known variant composition from both Illumina and Oxford Nanopore Technologies (ONT) platforms and assessed the performance of each VCE. We also evaluated VCEs performance using publicly available empirical wastewater samples collected for SC2 surveillance efforts. Bioinformatic analyses were performed with a custom NextFlow workflow (C-WAP, CFSAN Wastewater Analysis Pipeline). Relative root mean squared error (RRMSE) was used as a measure of performance with respect to the known abundance and concordance correlation coefficient (CCC) was used to measure agreement between pairs of estimators. Based on our results from simulated data, Kallisto was the most accurate estimator as it had the lowest RRMSE, followed by Freyja. Kallisto and Freyja had the most similar predictions, reflected by the highest CCC metrics. We also found that accuracy was platform and amplicon panel dependent. For example, the accuracy of Freyja was significantly higher with Illumina data compared to ONT data; performance of Kallisto was best with ARTICv4. However, when analyzing empirical data there was poor agreement among methods and variations in the number of variants detected (e.g., Freyja ARTICv4 had a mean of 2.2 variants while Kallisto ARTICv4 had a mean of 10.1 variants). This work provides an understanding of the differences in performance of a number of VCEs and how accurate they are in capturing the relative abundance of SC2 variants within a mixed sample (e.g., wastewater). Such information should help officials gauge the confidence they can have in such data for informing public health decisions.

60 APPLIED LIFE SCIENCES↗