Search NASA⌕ Search

SEARCH · Search NASA

Results for “bioinformatics workflows”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Challenges in Bioinformatics Workflows for Processing Microbiome Omics Data at Scale

The nascent field of microbiome science is transitioning from a descriptive approach of cataloging taxa and functions present in an environment to applying multi-omics methods to investigate microbiome dynamics and function. A large number of new tools and algorithms have been designed and used for very specific purposes on samples collected by individual investigators or groups. While these developments have been quite instructive, the ability to compare microbiome data generated by many groups of researchers is impeded by the lack of standardized application of bioinformatics methods. Additionally, there are few examples of broad bioinformatics workflows that can process metagenome, metatranscriptome, metaproteome and metabolomic data at scale, and no central hub that allows processing, or provides varied omics data that are findable, accessible, interoperable and reusable (FAIR). Here, we review some of the challenges that exist in analyzing omics data within the microbiome research sphere, and provide context on how the National Microbiome Data Collaborative has adopted a standardized and open access approach to address such challenges.

NMDC, Microbiome↗

Recommendations for Uniform Variant Calling of SARS-CoV-2 Genome Sequence across Bioinformatic Workflows

Genomic sequencing of clinical samples to identify emerging variants of SARS-CoV-2 has been a key public health tool for curbing the spread of the virus. As a result, an unprecedented number of SARS-CoV-2 genomes were sequenced during the COVID-19 pandemic, which allowed for rapid identification of genetic variants, enabling the timely design and testing of therapies and deployment of new vaccine formulations to combat the new variants. However, despite the technological advances of deep sequencing, the analysis of the raw sequence data generated globally is neither standardized nor consistent, leading to vastly disparate sequences that may impact identification of variants. Here, we show that for both Illumina and Oxford Nanopore sequencing platforms, downstream bioinformatic protocols used by industry, government, and academic groups resulted in different virus sequences from same sample. These bioinformatic workflows produced consensus genomes with differences in single nucleotide polymorphisms, inclusion and exclusion of insertions, and/or deletions, despite using the same raw sequence as input datasets. Here, we compared and characterized such discrepancies and propose a specific suite of parameters and protocols that should be adopted across the field. Consistent results from bioinformatic workflows are fundamental to SARS-CoV-2 and future pathogen surveillance efforts, including pandemic preparation, to allow for a data-driven and timely public health response.

60 APPLIED LIFE SCIENCES↗

Marine Microeukaryote Metatranscriptomics: Sample Processing and Bioinformatic Workflow Recommendations for Ecological Applications

Microeukaryotes (protists) serve fundamental roles in the marine environment as contributors to biogeochemical nutrient cycling and ecosystem function. Their activities can be inferred through metatranscriptomic investigations, which provide a detailed view into cellular processes, chemical-biological interactions in the environment, and ecological relationships among taxonomic groups. Established workflows have been individually put forth describing biomass collection at sea, laboratory RNA extraction protocols, and bioinformatic processing and computational approaches. Here, we present a compilation of current practices and lessons learned in carrying out metatranscriptomics of marine pelagic protistan communities, highlighting effective strategies and tools used by practitioners over the past decade. We anticipate that these guidelines will serve as a roadmap for new marine scientists beginning in the realms of molecular biology and/or bioinformatics, and will equip readers with foundational principles needed to delve into protistan metatranscriptomics.

Cohen, Natalie R.↗

An empirical DNA-based identification of morphologically similar snappers (Lutjanus campechanus, Lutjanus purpureus) using a versatile bioinformatics workflow for the discovery and analysis of informative single-nucleotide polymorphisms

The commercially important speciesLutjanus campechanus(Northern/Gulf red snapper) andLutjanus purpureus(Southern/Caribbean red snapper) are the protagonists of a decade’s long taxonomic debate over their species delimitation, due in part to partial habitat overlap, extensive morphological similarity, and the lack of resolution when applying canonically reliable DNA barcoding approaches. In this study, we leveraged publicly available RAD-Seq data forL. campechanusandL. purpureusto identify species-informative single‐nucleotide polymorphisms (SNPs) at the genome scale that were successful in distinguishing the Northern and Southern red snappers, while also detecting individuals exhibiting introgression. This 4-step empirical approach demonstrates the value of applying novel bioinformatics pipelines to existing genome-scale data to maximize the distillation of informative subsets. Our results facilitate economically relevant species identification in addition to confirming or challenging species identifications for specimens with data in public databases. These findings and their applications will benefit future sustainability strategies and broader research questions surrounding these overfished and evolutionarily entangled snapper species.

Environmental Sciences & Ecology↗

Accelerating Scientific Workflows on HPC Platforms with In Situ Processing

Scientific workflows drive most modern large-scale science breakthroughs by allowing scientists to define their computations as a set of jobs executed in a given order based on their data dependencies. Workflow management systems (WMSs) have become key to automating scientific workflows-executing computational jobs and orchestrating data transfers between those jobs running on complex high-performance computing (HPC) platforms. Traditionally, WMSs use files to communicate between jobs: a job writes out files that are read by other jobs. However, HPC machines face a growing gap between their storage and compute capabilities. To address that concern, the scientific community has adopted a new approach called in situ, which bypasses costly parallel filesystem I/O operations with faster in-memory or in-network communications. When using in situ approaches, communication and computations can be interleaved. In this work, we leverage the Decaf in situ dataflow framework to accelerate task-based scientific workflows managed by the Pegasus WMS, by replacing file communications with faster MPI messaging. We propose a new execution engine that uses Decaf to manage communications within a sub-workflow (i.e., set of jobs) to optimize inter-job communications. We consider two workflows in this study: (i) a synthetic workflow that benchmarks and compares file- and MPI-based communication; and (ii) a realistic bioinformatics workflow that computes mu-tational overlaps in the human genome. Experiments show that in situ communication can improve the bioinformatics workflow execution time by 22% to 30% compared with file communication. Our results motivate further opportunities and challenges for bridging traditional WMSs with in situ frameworks.

Decaf↗

EDGE COVID-19: a web platform to generate submission-ready genomes from SARS-CoV-2 sequencing efforts

Abstract Summary Genomics has become an essential technology for surveilling emerging infectious disease outbreaks. A range of technologies and strategies for pathogen genome enrichment and sequencing are being used by laboratories worldwide, together with different and sometimes ad hoc, analytical procedures for generating genome sequences. A fully integrated analytical process for raw sequence to consensus genome determination, suited to outbreaks such as the ongoing COVID-19 pandemic, is critical to provide a solid genomic basis for epidemiological analyses and well-informed decision making. We have developed a web-based platform and integrated bioinformatic workflows that help to provide consistent high-quality analysis of SARS-CoV-2 sequencing data generated with either the Illumina or Oxford Nanopore Technologies (ONT). Using an intuitive web-based interface, this workflow automates data quality control, SARS-CoV-2 reference-based genome variant and consensus calling, lineage determination and provides the ability to submit the consensus sequence and necessary metadata to GenBank, GISAID and INSDC raw data repositories. We tested workflow usability using real world data and validated the accuracy of variant and lineage analysis using several test datasets, and further performed detailed comparisons with results from the COVID-19 Galaxy Project workflow. Our analyses indicate that EC-19 workflows generate high-quality SARS-CoV-2 genomes. Finally, we share a perspective on patterns and impact observed with Illumina versus ONT technologies on workflow congruence and differences. Availability and implementation https://edge-covid19.edgebioinformatics.org, and https://github.com/LANL-Bioinformatics/EDGE/tree/SARS-CoV2. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

DOE BSSD Performance Management Metrics Report Q1

Microbes play key roles in our biosphere, from driving global nutrient cycling to impacting plant, animal and human health and disease. Complex data from microbial genomes, proteins, and metabolites provide a window into these tiny engines that drive life on our planet. Yet these data are dispersed among researchers’ laboratories and various repositories, making it difficult to access. This calls for new ways of managing data, improving data interoperability, advancing community standards, and creating an infrastructure where data are shared efficiently. We have built the National Microbiome Data Collaborative (NMDC) to advance how scientists create, use, and reuse data to redefine the way we understand and harness the power of microbes. The vision of the National Microbiome Data Collaborative (NMDC) is to drive a microbiome data sharing network connecting data, people, and ideas to advance microbiome innovation and discovery. The NMDC was launched in 2019 and brought together DOE National Laboratories to collaborate across resources, capabilities, and expertise. The NMDC team was strategically assembled to include software developers, microbial researchers, metadata experts, and multi-omics specialists. The diversity of the NMDC team reflects the inherently interdisciplinary nature of microbiome science, and we leverage the strengths of the DOE National Laboratory system. Towards BER’s goal of advancing an iterative systems biology approach to the understanding of microbial genomes, the NMDC serves as a foundation for infrastructure, data standards, and community building. Together with the flagship DOE User Facilities, the Joint Genome Institute (JGI) and the Environmental Molecular Sciences Laboratory (EMSL), we are developing core capabilities in metadata standards for environmental descriptors and sample handling and processing; standardized bioinformatic workflows; an interface for data search and access; and robust community engagement activities. The NMDC production platform supports long-term data infrastructure and community building for BER’s bioenergy and environmental research goals. Our approach leverages lessons learned and an ambitious framework for collaborative, interdisciplinary data infrastructure to support microbiome research. The NMDC supports data, information, and knowledge access through three defined software tools – the Submission Portal, NMDC EDGE, and the Data Portal – driven by community needs. Herein, we describe the value proposition for the microbiome research community, our overarching strategy, and challenges and opportunities for developing the NMDC as both an infrastructure and community engagement program.

59 BASIC BIOLOGICAL SCIENCES↗

Modularization of EDGE Workflows Using Nextflow: Improving the Efficiency and Maintainability of Bioinformatics Software

EDGE is a bioinformatics platform developed in 2016 by researchers at Los Alamos National Laboratory (LANL) to facilitate the analysis of next-generation sequencing data by researchers with varying levels of experience in bioinformatics (Li et al., 2017). Users with single-end, paired-end or long-read sequencing data can provide their reads as input to EDGE and select the combination of workflows to run that are most useful for their research (e.g., quality control of reads, genome assembly, or the taxonomic classification of input reads). Table 1 summarizes the modules available in EDGE. EDGE is available as a web platform at https://edgebioinformatics.org, as installable source code maintained on GitHub under a GPLv3 license, and as a publicly hosted Docker image.

59 BASIC BIOLOGICAL SCIENCES↗

Expanded Coverage of Phytocompounds by Mass Spectrometry Imaging Using On-Tissue Chemical Derivatization by 4-APEBA

Probing the entirety of any species metabolome is an analytical grand challenge, especially at a cellular scale. Where spatial metabolomics, completed primarily by matrix-assisted laser desorption/ionization (MALDI), has limited molecular coverage for several reasons. To expand the scope of spatial metabolomics, we developed an on-tissue chemical derivatization (OTCD) workflow using 4-APEBA for confident identification of several dozen elusive phytocompounds, including several phytohormones, which have various roles within stress responses and cellular communication. Superiority of 4-APEBA is established in comparison to other derivatization agents with (1) broad specificity towards carbonyls, (2) low background, and (3) introduction of bromine isotopes, where the latter two facilitate confident bioinformatics. In conclusion, the outlined workflow trailblazes a path towards spatial hormonomics within plant samples, enhancing detection of carboxylates, aldehydes, ketones, and plausibly phenols.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Critical Assessment of MetaProteome Investigation (CAMPI): a multi-laboratory comparison of established workflows

Metaproteomics has matured into a powerful tool to assess functional interactions in microbial communities. While many metaproteomic workflows are available, the impact of method choice on results remains unclear. Here, we carry out a community-driven, multi-laboratory comparison in metaproteomics: the critical assessment of metaproteome investigation study (CAMPI). Based on well-established workflows, we evaluate the effect of sample preparation, mass spectrometry, and bioinformatic analysis using two samples: a simplified, laboratory-assembled human intestinal model and a human fecal sample. We observe that variability at the peptide level is predominantly due to sample processing workflows, with a smaller contribution of bioinformatic pipelines. These peptide-level differences largely disappear at the protein group level. While differences are observed for predicted community composition, similar functional profiles are obtained across workflows. CAMPI demonstrates the robustness of present-day metaproteomics research, serves as a template for multi-laboratory studies in metaproteomics, and provides publicly available data sets for benchmarking future developments.

59 BASIC BIOLOGICAL SCIENCES↗

Toward designing effective exascale scientific computing workflows: experiences and best practices

Many fields within scientific computing have embraced advances in big-data analysis and machine learning, which often requires the deployment of large, distributed and complicated workflows that may combine training neural networks, performing simulations, running inference, and performing database queries and data analysis in asynchronous, parallel and pipelined execution frameworks. Such a shift has brought into focus the need for scalable, efficient workflow management solutions with reproducibility, error and provenance handling, traceability, and checkpoint-restart capabilities, among other needs. Here, we discuss challenges and best-practices for deploying exascale-generation computational science workflows on resources at the Oak Ridge Leadership Computing Facility (OLCF). We present our experiences with large-scale deployment of distributed workflows on the Summit supercomputer, including for bioinformatics and computational biophysics, materials science, and deep learning model optimization. We also present problems and solutions created by working within a Python-centric software base on traditional HPC systems, and discuss steps that will be required before the convergence of HPC, AI, and data science can be fully realized. Our results point to a wealth of exciting new possibilities for harnessing this convergence to tackle new scientific challenges.

Coletti, Mark↗

Performance of methods for SARS-CoV-2 variant detection and abundance estimation within mixed population samples

The accurate identification of SARS-CoV-2 (SC2) variants and estimation of their abundance in mixed population samples (e.g., air or wastewater) is imperative for successful surveillance of community level trends. Assessing the performance of SC2 variant composition estimators (VCEs) should improve our confidence in public health decision making. Here, we introduce a linear regression based VCE and compare its performance to four other VCEs: two re-purposed DNA sequence read classifiers (Kallisto and Kraken2), a maximum-likelihood based method (Lineage deComposition for Sars-Cov-2 pooled samples (LCS)), and a regression based method (Freyja). We simulated DNA sequence datasets of known variant composition from both Illumina and Oxford Nanopore Technologies (ONT) platforms and assessed the performance of each VCE. We also evaluated VCEs performance using publicly available empirical wastewater samples collected for SC2 surveillance efforts. Bioinformatic analyses were performed with a custom NextFlow workflow (C-WAP, CFSAN Wastewater Analysis Pipeline). Relative root mean squared error (RRMSE) was used as a measure of performance with respect to the known abundance and concordance correlation coefficient (CCC) was used to measure agreement between pairs of estimators. Based on our results from simulated data, Kallisto was the most accurate estimator as it had the lowest RRMSE, followed by Freyja. Kallisto and Freyja had the most similar predictions, reflected by the highest CCC metrics. We also found that accuracy was platform and amplicon panel dependent. For example, the accuracy of Freyja was significantly higher with Illumina data compared to ONT data; performance of Kallisto was best with ARTICv4. However, when analyzing empirical data there was poor agreement among methods and variations in the number of variants detected (e.g., Freyja ARTICv4 had a mean of 2.2 variants while Kallisto ARTICv4 had a mean of 10.1 variants). This work provides an understanding of the differences in performance of a number of VCEs and how accurate they are in capturing the relative abundance of SC2 variants within a mixed sample (e.g., wastewater). Such information should help officials gauge the confidence they can have in such data for informing public health decisions.

60 APPLIED LIFE SCIENCES↗

The need for standardization and improved open (meta)data practices in metaproteomics

Metaproteomics enables functional insight into microbial communities by identifying and quantifying proteins in complex samples. Yet, heterogeneous analytical workflows and the lack of standardization across experimental and bioinformatics stages hinder reproducibility and comparability, limiting integration with other omics data. We here present a community-developed reporting checklist tailored to the specific needs of metaproteomics. We also outline current efforts to enable structured and interoperable metadata capture, drawing on standards from proteomics and microbiome research wherever possible. By promoting transparent reporting and advancing metadata practices, our recommendations aim to align metaproteomics more closely with FAIR principles and support reproducible and interoperable research practices.

Armengaud, Jean [Universite Paris-Saclay, France]↗

Spatial Proteomics towards cellular Resolution

Introduction: Spatial biology is an emerging interdisciplinary field facilitating biological discoveries through the use of spatial omics technologies. Recent advancements in spatial transcriptomics, spatial genomics (e.g. genetic mutations and epigenetic marks), multiplexed immunofluorescence, and spatial metabolomics/lipidomics have enabled high-resolution spatial profiling of gene expression, genetic variation, protein expression, and metabolites/lipids profiles in tissue. These developments contribute to a deeper understanding of the spatial organization within tissue microenvironments at the molecular level. Areas covered: This report provides an overview of the untargeted, bottom-up mass spectrometry (MS)-based spatial proteomics workflow. It highlights recent progress in tissue dissection, sample processing, bioinformatics, and liquid chromatography (LC)-MS technologies that are advancing spatial proteomics toward cellular resolution. Expert opinion: The field of untargeted MS-based spatial proteomics is rapidly evolving and holds great promise. To fully realize the potential of spatial proteomics, it is critical to advance data analysis and develop automated and intelligent tissue dissection at the cellular or subcellular level, along with high-throughput LC-MS analyses of thousands of samples. In conclusion, achieving these goals will necessitate significant advancements in tissue dissection technologies, LC-MS instrumentation, and computational tools.

59 BASIC BIOLOGICAL SCIENCES↗

Metagenomic Methods for Addressing NASA's Planetary Protection Policy Requirements on Future Missions: A Workshop Report

Molecular biology methods and technologies have advanced substantially over the past decade. These new molecular methods should be incorporated among the standard tools of planetary protection (PP) and could be validated for incorporation by 2026. To address the feasibility of applying modern molecular techniques to such an application, NASA conducted a technology workshop with private industry partners, academics, and government agency stakeholders, along with NASA staff and contractors. The technical discussions and presentations of the Multi-Mission Metagenomics Technology Development Workshop focused on modernizing and supplementing the current PP assays. The goals of the workshop were to assess the state of metagenomics and other advanced molecular techniques in the context of providing a validated framework to supplement the bacterial endospore-based NASA Standard Assay and to identify knowledge and technology gaps. In particular, workshop participants were tasked with discussing metagenomics as a stand-alone technology to provide rapid and comprehensive analysis of total nucleic acids and viable microorganisms on spacecraft surfaces, thereby allowing for the development of tailored and cost-effective microbial reduction plans for each hardware item on a spacecraft. Workshop participants recommended metagenomics approaches as the only data source that can adequately feed into quantitative microbial risk assessment models for evaluating the risk of forward (exploring extraterrestrial planet) and back (Earth harmful biological) contamination. Participants were unanimous that a metagenomics workflow, in tandem with rapid targeted quantitative (digital) PCR, represents a revolutionary advance over existing methods for the assessment of microbial bioburden on spacecraft surfaces. The workshop highlighted low biomass sampling, reagent contamination, and inconsistent bioinformatics data analysis as key areas for technology development. Finally, it was concluded that implementing metagenomics as an additional workflow for addressing concerns of NASA's robotic mission will represent a dramatic improvement in technology advancement for PP and will benefit future missions where mission success is affected by backward and forward contamination.

59 BASIC BIOLOGICAL SCIENCES↗

wastewater_virus

This repo contains software used to clean and assemble high-throughput sequencing data containing viruses. The input is raw illumina sequencing reads and the output is a database of high-quality viral genomes. The specific application is to wastewater viral concentrates but it is not restricted to that sample type. The software is composed of Nextflow workflows and a set of custom Python and bash scripts that call publicly available bioinformatics tools to accomplish obvious tasks in data analysis in a high performance computing environment. For detailed information, please see the repo's README file.

Kantor, Rose [Lawrence Livermore National Laborato↗