Search NASASearch

Engineering topics

Sylvain Costes

Publications and source records attributed to Sylvain Costes.

At least 19 records

Combining Large Datasets - Cancer Moonshot Task Group Final Summary

In February 2022, President Biden re-ignited the Cancer Moonshot with bold new goals: to reduce the cancer death rate by half within 25 years and improve the lives of people with cancer and cancer survivors. To achieve these ambitious goals, the White House convened the first-ever Cancer Cabinet, bringing together departments and agencies from across the federal government to end cancer as we know it.The Cancer Cabinet convened three task forces and supporting task groups, including the Data and Innovation Task Force, which supported the Cancer Moonshot priority to “Deliver innovation to patients and communities.” In early 2023, the “Combining Large Datasets” (CoLD) Task Group was created within the Data and Innovation Task Force. The scope of the CoLD Task Group was how federal agencies combine large datasets for broad applications across cancer prevention and control, including nutrition, epidemiology, and military/Veteran health. Within this scope, the group sought to better leverage the immense potential of data and power of data tools to increase our understanding of cancer incidence, causes, mortality, treatments, prevention, outcomes, costs, and all other aspects of the burden of cancer.

data integration

GeneLab: The NASA Systems Biology Platform for Space Omics Repository, Analysis and Visualization

The NASA GeneLab project capitalizes on multi-omic technologies to maximize the return on spaceflight experiments. To do this, GeneLab maintains a publicly accessible database (GLDS) that houses spaceflight and spaceflight relevant multi-omics data, and collaborates with NASA principal investigators and projects to generate additional omics data. GeneLab houses more than 220 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, animal and microbial experiments, with a growing number of these having been produced by the GeneLab sample processing lab. The GLDS contains rich metadata about each experiment and has recently integrated radiation dosimetery data from experiments flown on the Space Shuttle. GeneLab has also recently implemented an effort to present processed data in the GLDS in addition to the raw omics data. The processed data will enable interpretation of the data by a larger group of students, scientists and the general public. Standard pipelines for the transformation of raw data into visualizations were developed by four GeneLab Analysis Working Groups (animals, plants, microbes, multi-omics) comprised of over 120 scientists from NASA, industry, and academia. To explore the data, the GLDS provides users various tools for data analysis, collaborative workspace for file storage and sharing, and a visualization portal. The analysis platform built using the Galaxy toolshed provides access to a broad variety of users including those with limited bioinformatics experience and students to learn how to analyze spaceflight omics data. The visualization portal takes GeneLab one step closer to data democratization by removing all bioinformatics requisites to interpret transcriptomics data hosted in the repository. Discoveries made using GeneLab have begun and will continue to deepen our understanding of biology, advance the field of genomics, and help to discover cures for diseases, create better diagnostic tools, and ultimately allow astronauts to better withstand the rigors of long-duration spaceflight.

Samrawit Gebre

NASA GeneLab: The NASA Systems Biology Platform for Spaceomics Repository, Analysis and Visualization

At NASA Ames Research Center, the GeneLab Open Science Project is on a mission to gather all large -omics datasets relevant to space biology research. These datasets come from various organisms flown in multiple space habitats such as the International Space Station or the Space Shuttle, in addition to mimicking space-like conditions on ground. Researchers and citizen scientists all around the world have used the data and the analytical tools put together by the GeneLab team to start deciphering new biological impact of microgravity, space ionizing radiation and other space stressors.

GeneLab

Combining RNA-SEQ Datasets from NASA GENELAB: An Evaluation of Correction Methods

Background: Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets. Methods: In the present study, NASA GeneLab-hosted RNAseq datasets from mouse liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, the median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC to determine statistical differences between datasets before and after correction, Principal Component Analysis to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. Results: The results showed that the reference-based approach introduced several additional (and likely artificial) differentially expressed genes when compared with the respective standard approach. Conclusions: Of the methods tested, standard ComBat_seq and DESeq2 were identified as the most robust correction methods for combining spaceflight mouse liver RNAseq datasets hosted on GeneLab.

Finsam Samson

Genomic and Phenotypic Associations to Predict Human Sensitivity to SpaceRadiation

High-linear energy transfer (LET) ionizing radiation is a major health hazard for astronauts who will be exposed to galactic cosmic rays during upcoming lunar and Mars missions. Predicting and mitigating this risk requires understanding the factors underlying individual radiation sensitivity. We started to address this challenge by identifying the genomic and phenotypic associations with sensitivity to low and high-LET ionizing radiation ex vivo in over 750 healthy human donors. We exposed primary human blood mononuclear cells to simulated galactic cosmic ray components: 350MeV/n 28Si, 350MeV/n 40Ar and 600MeV/n 56Fe particles, at 1.1 and 3 particles/100 sq.m fluences, as well as 0.1 Gy and 1 Gy doses of gamma rays, and analyzed the outcomes at 4 and 24 hours post-irradiation. We quantified DNA damage and repair responses based on 53BP1+ radiation-induced foci formation, together with oxidative stress and changes in secreted factors including immune cytokines and exosomes. We also analyzed phenotypic associations with spontaneous DNA repair foci at baseline prior to irradiation. We identified an increase in spontaneous DNA repair associated with age and latent viral infection, and observed that human spontaneous DNA repair foci at baseline can serve as a negative predictor of DNA repair and immunoregulatory cytokine production after irradiation. Furthermore, we observed a wide variability of subject- and LET-dependent radiation responses, with radiation-induced DNA repair foci increasing by dose and LET, and high-LET particle radiation resulting in more residual DNA damage compared to low-LET particles and gamma rays. We have developed multiple metrics to quantify human radio sensitivity across the spectrum of conditions. Here we present their dependence on radiation quality and phenotypic variables, including an age-dependent decrease in DNA repair after irradiation, and genomic associations with radio sensitivity based on low-coverage whole genome sequencing. We anticipate that our work will pave the way for understanding the spectrum of human radio sensitivity and identifying targets for countermeasure development to reduce DNA and cellular damage and radiation carcinogenesis during deep space exploration.

GWAS radiation human

Harnessing Large Language Models for Scientific Endeavors

The rapid proliferation of Large Language Models (LLMs) such as GPT, Bard, and Llama has revolutionized various sectors, including the scientific community. These models, with their potential to automate and augment tasks, are increasingly being recognized as both a valuable asset and a potential challenge in the realm of scientific research and data management. However, the current LLMs, primarily trained on general corpora, exhibit a limited understanding of scientific concepts and terminologies due to the lack of scientific corpus in their training data. Recognizing this gap, several groups are now advocating for the development of LLMs specifically tailored for scientific applications. A notable initiative in this direction is the Large Language Model effort initiated by NASA's CSDO. This endeavor aims to align LLM efforts across NASA’s Science Mission Directorate, develop a science-specific corpus and validation test set for model training, and create an encoder-only model for various downstream tasks. Moreover, the initiative also plans to develop a decoder-only model to explore the potential benefits and risks associated with a generative LLM for science. Lastly, the project aims to create a science evaluation suite, encompassing various categories of downstream scientific tasks, to serve as a benchmark for assessing the value of any LLM for future use. This presentation will provide an overview and current status of this ongoing initiative, highlighting its potential to reshape the use of LLMs in the scientific domain.

Rahul Ramachandran

NASA Open Science Data Repository: Biomedical FAIR Data, Analysis Tools, User Communities, and Discoveries for Deep Space Missions

Increased biomedical risks and challenges associated with deep space missions require new knowledge discovery, new health countermeasures, and development of novel ecosystems, life support, crop production, and biomedical support capabilities. To meet NASA’s Moon to Mars strategic program goals for Human and Biological Sciences, findable, accessible, interoperable, reusable (FAIR), and maximally open-access data is going to be required to enable humanity to thrive in deep space. Indeed, this cornerstone perspective on FAIR and maximally open access data was also recommended in the recent 2023-2032 Decadal Survey from the National Academies of Sciences, Engineering, and Medicine. The NASA Open Science Data Repository (OSDR) is a maximally open access and FAIR database, and meets various scientific, technical, and operational spaceflight needs. It offers public users and submitters the ability to upload, download, search, share, analyze, and visualize data across ‘omics, physiological, phenotypic, behavioral, bioimaging, video, and environmental monitoring telemetry datasets. OSDR includes NASA GeneLab, NASA Ames Life Sciences Data Archive, and the NASA Biological Institutional Scientific Collection. OSDR has >455 studies with datasets from model organisms and non-NASA human astronauts. There are ~12 datasets from the Inspiration 4 (I4) mission, spanning metagenomics, comprehensive metabolic panels, clonal hematopoiesis, spatial transcriptomics, proteomics, and cytokine panels. In the interest of data privacy, two I4 datasets have raw FASTQ and FASTA files relating to the epitranscriptome, and a new request feature is live in OSDR (with a backend review process established) which was developed based on industry norms. OSDR also recently began a collaboration with the European Space Agency (ESA) to scientifically curate and make available >200 terabytes of human and model organism space-relevant data. The OSDR submission portal is designed to ingest and curate ~25 ‘omics assay data types, and ~50 physiological-phenotypic-imaging assay data types, spanning ultrasonography, micro-computed tomography, histology, morphometric photography, rebound tonometry, gait analysis, optical coherence tomography, novel object recognition, flow cytometry, and immunohistochemistry. A suite of analysis tools are available for OSDR users including: 1) an Environmental Data Application to compare radiation, CO2, relative humidity, temperature, and other telemetry across missions and subjects, 2) the RadLab database, a collaboration between NASA, ESA, the German and Italian Space Agencies, and the Bulgarian Academy of Sciences, which compiles radiation measurements relevant to human spaceflight and provides tools for accessing and manipulating the data, and 3) a Multi-study visualization tool which enables users to look across and combine GeneLab’s omics datasets across different experiments and missions. There are ~600 volunteer OSDR Analysis Working Group (AWG) members who: 1) provide feedback on scientific standards for reuse (subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability), and 2) collaborate to mine-reuse OSDR data conducting scientific analysis. OSDR has enabled 60 publications as of September 2023, many directly from AWG collaborations most notably the Cell Press package in 2020. Lastly, there are at least 15 articles which mine OSDR data part of a package of ~50 articles across Nature Portfolio with research stemming from I4, the Japan Aerospace Exploration Agency, NASA Space Biology, and the NASA Human Research Program.

open access