Search NASASearch

Engineering topics

Sylvain V. Costes

Publications and source records attributed to Sylvain V. Costes.

At least 19 records

iGCE and MitoFlyght Spaceflight Missions: Unraveling Oxidative Stress Responses in Space

Thriving In DEep Space (TIDES) initiative aims to comprehensively understand how hostile environments such as the Moon and Mars affect human physiology. Here we present two NASA-selected spaceflight experiments under the TIDES portfolio, iGCE (Integrated Gravity Continuum Experiment) and MitoFlyght (Mitochondrial Investigation of Oxidative Stress in Flies). We hypothesize that exposure to spaceflight conditions induces oxidative stress responses that negatively impact physiology. The iGCE mission employs two well-established spaceflight models, Drosophila melanogaster and C.elegans, using Redwire’s Multi-use Variable-g Platform (MVP) hardware to assess changes in cardiac, muscle, and nervous systems across five different gravities: Hypergravity (2g), Earth (1g), Mars (0.37g), Moon (0.16g), and microgravity (ug). This mission focuses on uncovering alterations in protein homeostasis, autophagy, and mitochondrial function conserved across species. In the MitoFlyght mission to the ISS, Drosophila will be housed in the Vented Fly Box (VFB). This mission evaluates the oxidative stress response and autophagic pathway in muscle, heart, and nervous system. Additionally, we will (a) use the genetic mutant, Tor7/P to test whether increased autophagy is beneficial or a maladaptive response to the spaceflight stressors, and (b) utilize fly lines with tissue-specific expression (neuronal, muscle, and cardiac) of an antioxidant gene, SOD2 (superoxide dismutase) as a potential countermeasure. Data from these missions will be compared with previous LEO-based datasets to identify shared signatures. Furthermore, cross-species analysis of the transcriptomic data from other invertebrate and vertebrate spaceflight studies will help determine evolutionarily conserved pathways perturbed by space stressors. Overall, both these missions aim to provide crucial insights into the mechanisms underlying oxidative stress responses, synaptic changes, and heart and muscle deficits, facilitating the identification of diagnostic and therapeutic targets to mitigate the adverse health effects of long-duration space habitation. Ultimately, this research will enhance our ability to thrive in deep space and inform future missions.

Janani Iyer

Machine Intelligence for Radiation Science: Summary of the Radiation Research Society 67th Annual Meeting Symposium

The era of high-throughput techniques created big data in the medical field and research disciplines. Machine intelligence (MI) approaches can overcome critical limitations on how those large-scale data sets are processed, analyzed, and interpreted. The 67 th Annual Meeting of the Radiation Research Society featured a symposium on MI approaches to highlight recent advancements in the radiation sciences and their clinical applications. This article summarizes three of those presentations regarding recent developments for metadata processing and ontological formalization, data mining for radiation outcomes in pediatric oncology, and imaging in lung cancer.

radiation

WEBINAR, May 6: New Discoveries Using GeneLab

The NASA GeneLab project capitalizes on multi-omic technologies to maximize the return on spaceflight experiments. To do this, GeneLab maintains a publicly accessible database (GLDS) that houses spaceflight and spaceflight relevant multi-omics data and collaborates with NASA principal investigators and projects to generate additional omics data. GeneLab houses more than 220 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, animal and microbial experiments, with a growing number of these having been produced by the GeneLab sample processing lab. The GLDS contains rich metadata about each experiment and has recently integrated radiation dosimetry data from experiments flown on the Space Shuttle. GeneLab has also recently implemented an effort to present processed data in the GLDS in addition to the raw omics data. The processed data will enable interpretation of the data by a larger group of students, scientists and the general public. Standard pipelines for the transformation of raw data into visualizations were developed by four GeneLab Analysis Working Groups (animals, plants, microbes, multi-omics) comprised of over 120 scientists from NASA, industry, and academia. To explore the data, the GLDS provides users various tools for data analysis, collaborative workspace for file storage and sharing, and a visualization portal. The analysis platform built using the Galaxy toolshed provides access to a broad variety of users including those with limited bioinformatics experience and students to learn how to analyze spaceflight omics data. The visualization portal takes GeneLab one step closer to data democratization by removing all bioinformatics requisites to interpret transcriptomics data hosted in the repository. Discoveries made using GeneLab have begun and will continue to deepen our understanding of biology, advance the field of genomics, and help to discover cures for diseases, create better diagnostic tools, and ultimately allow astronauts to better withstand the rigors of long-duration spaceflight.

Sylvain V. Costes

Dose, LET, time and strain dependence of radiation-induced 53BP1 foci in 15 mouse strains ex vivo and associations to in vivo radiation susceptibility

We present a comparative analysis on the repair of radiation-induced DNA damage ex vivo in 15 strains of mice, including 5 inbred reference strains and 10 collaborative-cross strains, of both sexes. Non-immortalized primary skin fibroblasts derived from 76 mice were subjected to both low- and high-LET radiation (0.1, 1 and 4 Gy of X rays; 1.1 and 3 particles/100μm2 of 350 MeV/n 40Ar and 600 MeV/n 56Fe). Automated image quantification of 53BP1 radiation-induced foci (RIF) during the first 4-48 h post-irradiation was performed as a function of dose and LET. Similarly to what we had previously reported for immortalized human cell lines [1], we observed a saturation of RIF number with dose at 4h post-irradiation, with more RIF/Gy for lower LET (X rays and 40Ar) compared to 56Fe. However at later time points (24h and above), the trend was inverted with more RIF/Gy for higher LET. Our data suggest that multiple DSBs cluster into RIF: as the linear density of DSBs increases with LET, so does the probability of having more DSBs per RIF, which makes it more difficult for cells to fully resolve high-LET-induced RIF, explaining the hypersensitivity to high-LET radiation despite a low number of RIF. Taking into account the amount of clustering at a given dose and LET, but also the kinetics of DNA damage repair, we introduced a novel mathematical formalism to evaluate the number of remaining RIF over time. We showed that the newly introduced kinetic metrics can be used as surrogate biomarkers for in vivo radiation toxicity, with potential applications in radiotherapy and human space exploration. In particular, we observed an association between the repairable fraction of RIF measured in vitro and survival levels of immune cells collected from irradiated mice. Moreover, the speed of DNA damage repair correlated with spontaneous cancer incidence data collected from the Mouse Tumor Biology database, suggesting a relationship between the efficiency of DSB repair after irradiation and cancer risk. In addition to the efficacy of repair and persistent RIF levels, even the amount of spontaneous foci without irradiation was shown to be strain dependent, indicating that these phenotypes are at least partially driven by genetics, and supporting their potential as indicators of individual radiation sensitivity. [1] Neumaier, T., et al., PNAS, 2012 (8) 109:443

Radiation, DNA damage, repair kinetics

BLOOD-BASED MULTI-SCALE MODEL FOR CANCER RISK FROM GCR IN GENETICALLY DIVERSE POPULATIONS

OBJECTIVES AND METHODS This project addresses the challenge of understanding and predicting individual radiation sensitivity by integrating genetics, demographics and biomarker characteristics across species (mice and humans). We hypothesize that ex vivo DNA repair response to GCR components is a central determinant of cancer risk from space radiation and can serve as a biomarker of radiation risk in combination with genetics. Automated image quantification of 53BP1+ radiation-induced foci (RIF) during the first 4-48 h post-irradiation was performed as a function of dose and LET in non-immortalized primary skin fibroblasts derived from 76 mice across 15 strains (5 inbred reference strains and 10 collaborative-cross strains) exposed to X rays (0.1, 1 and 4 Gy), 350 MeV/n 40Ar and 600 MeV/n 56Fe (1.1 and 3 particles/100μm2), as well as in peripheral blood mononuclear cells (PBMCs) from 768 healthy donors (matched ethnicity, 50/50 male/female, 18-70 years old) exposed to gamma rays (0.1 and 1 Gy), 350 MeV/n 28Si, 350 MeV/n 40Ar and 600 MeV/n 56Fe (1.1 and 3 particles/100μm2). QUANTIFICATION OF 53BP1+ FOCI IN VITRO AND ASSOCIATIONS TO IN VIVO RADIATION SUSCEPTIBILITY IN 15 MOUSE STRAINS We reported in vitro repair kinetic and repairable fractions of RIF for the 15 mouse strains and introduced a mathematical model for RIF as a function of time, dose and LET. We noted that the metabolic activity of cells modulates the RIF response, and we introduced the open access tool terRIFic (Tool for Enhanced Results of RIF In Cells, https://radbiolab.shinyapps.io/terrific/) to correct for such bias using confluence level. Notably, at 4h post-irradiation, RIF/Gy decreased with dose or LET: as the dose or LET increases, so does the proximity of DNA double-strand-breaks (DSB) and our data suggest that proximal DSBs are brought together inside isolated RIF for repair. The RIF/Gy trend was inverted at 24h, suggesting RIF with high DSB content are more difficult to repair. We showed that in vitro metrics correlate with in vivo measurements in the same 15 mouse strains, such as survival levels of immune cells or spontaneous cancer incidence, suggesting a relationship between the efficiency of DSB repair and cancer risk or radiation toxicity. In addition to the efficiency of repair and persistent RIF, the amount of spontaneous foci before irradiation was also found to be strain dependent. Finally, we performed genome-wide association study in the same 15 mouse strains using all RIF phenotypes measured in vitro, identifying genes of interest and validating RIF as an ideal biomarker for individual radiation sensitivity. BASELINE 53BP1+ FOCI PREDICTS INDIVIDUAL HUMAN RESPONSE TO GCR COMPONENTS Based on the analysis of radiation responses of 576 donor PBMCs (using quantification of 53BP1+ foci, oxidative stress and cell death), we observed a wide variability of subject- and LET-dependent radiation responses, with radiation-induced DNA repair foci increasing with LET, though oxidative stress being notably reduced by high-LET irradiation, potentially due to a switch between hydrogen peroxide and oxygen radical-based mechanisms. We identified a relationship between few spontaneous DNA foci at baseline and increased DNA repair after irradiation, accompanied by an alteration in immunoregulatory cytokine secretion, which might be adapted as biomarkers to predict ionizing radiation sensitivity. Among demographic variables, only latent cytomegalovirus infection and age were predictive of high baseline foci formation. Finally, we have performed low-throughput whole genome sequencing of all samples and are currently in the process of identifying the genes and pathways associated with low and high-LET ionizing radiation sensitivity in humans.

53BP1

New developments in space radiation research at NASA: Annotating data using a novel radiation biology ontology

Like many interdisciplinary sciences, data producers and consumers in the field of radiation biology often use a wide variety of terminology to describe their experiments and data. Furthermore, space systems and technologies are rapidly evolving, and a shared understanding and common terminology for these is also lacking. The efficiency of research organizations can be enhanced by standardizing metadata through the use of knowledge resources like ontologies. Employing a sophisticated model such as a formal ontology to standardize metadata enables automated data acquisition processes and supports more complete, accurate meta-analysis through more efficient and complete data discovery and retrieval, particularly when using multiple data sources. Thus, we developed the Radiation Biology Ontology (RBO) in order to improved radiation biology metadata uniformity and transparency. We used open-source software (the Ontology Development Kit, Protégé and WebProtégé) and worked within the OBO Foundry framework, which includes a set of ontology development principles and practices for ontology consistency, uniformity, and accountability. The RBO has now been incorporated into two radiation research data repositories, NASA’s GeneLab omics database (https://genelab.nasa.gov), and the European Commission STORE database (https://www.storedb.org/). Continuous build integration tools allowed our international RBO collaboration to be more efficient and focus its efforts on semantic model design. Currently, the RBO contains over 300 annotated classes and individuals specific to the study of radiation on biological systems, as well as imports of many additional classes from other OBO Foundry ontologies that relate to and/or provide context for these RBO entities. We publish the RBO through the OBO Foundry, so that it is available for browsing, download, and querying through NCBI Bioportal web site and application programming interface. The NASA Ames Life Science Data Archive (ALSDA) is also in the process of adopting use of the RBO, taking NASA one step closer to a knowledge-based system for space biology data. It is our hope that the global communities of radiation research Investigators, data curators and data analysts can similarly leverage the RBO and will contribute to its further development.

radiation

RadBREAD: Radiation Biology Research at an Elevated Altitude through Dosimetry – A student-designed payload

NASA uses extreme environment platforms (ground testing facilities, high-altitude balloons and aircraft, and CubeSats) to provide greater understanding of the conditions and limitations of extra-terrestrial environments. As part of a two-week flight planned for summer 2021, RadBREAD (Radiation Biology Research at an Elevated Altitude through Dosimetry) will fly as a secondary payload consisting of a M-42C (German Aerospace Center, DLR) ionizing radiation dosimeter, UV micro-logger, and multiple desiccated yeast samples. The platform is a novel high-altitude solar-powered aircraft: the Swift Engineering High-Altitude samples. The platform is a novel high-altitude solar-powered aircraft: the Swift Engineering High-Altitude Long-Endurance Unmanned Aircraft System (HALE UAS), which offers significantly longer flight durations than other high-altitude platforms. The yeast Saccharomyces cerevisiae will provide meaningful biological correlation for the sensor readings, due to its resistance to extremely low temperature and pressure when desiccated, ease of genetic manipulation, and homology to human genes. The RadBREAD team comprises the 2020 cohort of NASA’s Space Life Sciences Training Program (SLSTP) research associates as well as NASA scientists, engineers and radiation experts from NASA and the DLR. Yeast survival, metabolic, and transcriptomic changes will be correlated with environmental data collected during long-term exposure to the upper atmosphere. Additionally, the team will evaluate the upper atmospheric environment (radiation, pressure, and temperature) provided by the HALE UAS platform as a Mars surface analog for biological payloads. We hypothesize that exposure to upper atmospheric conditions during the HALE UAS flight will alter the survival, metabolism, and transcriptome of desiccated wild-type S. cerevisiae upon rehydration compared to sensitive and tolerant yeast strains exposed to the same conditions, and between the flight samples compared to asynchronous ground controls.

radiation exposure

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, the re-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA’s Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related ‘omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata ‘omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data re-use, resulting in 38 additional publications derived from the original 67 publication over the past four years. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA “Open Science Data Repositories (OSDR)” and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Fluorescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to “big data” from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology. Several other talks will cover these topics in this conference.

life sciences

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, there-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA's Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomatic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related 'omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata 'omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data-use, resulting in 40 enabled publications by open data. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA "Open Science Data Repositories (OSDR)" and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Flourescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to "big data" from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology.

omics

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching. The use of health countermeasures and biomonitoring systems for space missions are required to counteract space health hazards and to support life to thrive in deep space (e.g., humans, animals, plants, crops; entire ecosystems within spacecrafts/habitats/spacesuits). The development of these mission components will be highly dependent on our understanding of basic biological and health responses to myriad space hazards (ionizing radiation, altered gravitational fields, altered day-night cycles, confined isolation, hostile-closed environments, distance-duration from Earth, planetary dust-regolith, and extreme temperatures/atmospheres). The fast-growing array of space biological and mission telemetry data, which in the past was simply archived after minimal analysis, holds great potential once applied to these mission challenges if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its multi-hierarchical, multi-modal, and heterogenous nature (molecular, cellular, tissue, organ, whole organism, behavior, ecosystem, microbiome; tabular, omics, imaging, video, biospecimen, environmental physical-chemical telemetry). This session focuses on current approaches in this domain such as: making space biological data FAIR (findable, accessible, interoperable, reusable), effective data ingestion/dissemination, observational versus experimental data, Open Science collaborations, data analysis techniques, AI/ML/knowledge graph/modeling methods, and data integration/discovery tools.

open science

Using Federated Learning to Overcome Data Gravity in Space

Humans intend to take longer missions to outer space. Understanding the impact that space has on human health is paramount to the success of these missions. Controlled experiments with model organisms are run to infer the impact of space conditions on human health, but the data these experiments generate are too large to transfer to Earth for building models. The same is true for space-relevant data generated on Earth. Ideally, these datasets should be combined to improve statistical power and model accuracy without having to transfer data. Federated learning is such a method which trains an algorithm across decentralized computing systems, each of which has their own local copy of training and testing data. In this research, made possible by NASA@Work, the AI for Life in Space group at NASA demonstrates the use of federated learning to train an ensemble of causality inference models on a combination of data residing on the International Space Station (ISS) and in the cloud. Our work leverages CRISP, a causal inference platform developed during the 2020 Frontier Development Lab’s “Astronaut Health Challenge.” We also leverage the OpenFL federated learning library which was collaboratively developed at Intel and UPenn. We used publicly available data from the NASA Ames Life Sciences Data Archive to identify features in ionizing radiation experiments as causal of changes in cardiac blood velocity. This research demonstrates, for the first time, the possibility of running machine learning algorithms on datasets separated by astronomical distances. In this experiment, all the data were generated in terra, half of which were transferred to the ISS and analyzed on the Spaceborne Computer. In the future, our research will leverage federated learning on data generated in situ on the ISS with data generated terrestrially to predict the impact of spaceflight on mammalian female reproductive capacity.

James Casaletto

A Pipeline for Assessing the Quality of Rna-Seq Datasets in GeneLab

Transcriptome profiling by RNA sequencing (RNA-seq) is a powerful approach to identify gene expression changes in organisms exposed to unique environments such as spaceflight. One of the challenges of evaluating RNA-seq data both within and across different space-relevant studies is the ability to control for technical differences, including the use of different library preparation kits, sequencing platforms, RNA yield, and person-to-person variation. To help address this issue, the National Institute of Standards and Technology (NIST, nist.gov) initiated a consortium, at the request of industry and academia, to develop a set of controls for gene expression measurements. The result was a set of 92 unlabeled, polyadenylated transcripts that range from 250 – 2,000 nucleotides in length to mimic natural eukaryotic mRNAs. These External RNA Controls Consortium (ERCC) genes can be used in any RNA-seq experiment, by adding known concentrations of the ERCC genes to samples after RNA extraction, to offer a standard measurement for data comparison. At NASA GeneLab, we employ these controls as part of our standard operating procedures for every in-house RNA-seq study to assess the limit of detection, dynamic range, and power of differential expression analysis both within and across experiments. Here we will discuss the use, benefits, and limitations of ERCC genes and other types of controls, such as universal RNA references, to generate quality control information for RNA-seq studies conducted at GeneLab.

GeneLab

GL4U: Using Space Omics Data to Provide Bioinformatics Training for Students and Educators

NASA’s GeneLab project provides researchers open access to space-relevant experiment multi-omics data that can be mined to understand the effects of spaceflight on biological systems. To maximize the number of scientists who understand and utilize GeneLab data and data processing pipelines, GeneLab has created GeneLab for Colleges and Universities (GL4U). GL4U provides space biology-relevant training in bioinformatics to the next generation of scientists through direct and indirect approaches. The GeneLab team plans to host two annual data processing bootcamps, one for college-level students (direct) and one for college educators (indirect – training of trainers), in which participants learn to analyze GeneLab’s space-relevant omics data. The GL4U direct training pilot program was conducted in June 2021. During the pilot, students participated in a week-long bootcamp consisting of space biology-specific lectures and hands-on instruction using Jupyter Notebooks to analyze RNA sequence data. This pilot demonstrated the capacity of GL4U for training young scientists and encouraging data re-use. In June 2022, GL4U partnered with Jet Propulsion Laboratory’s (JPL) Planetary Protection Center of Excellence to conduct the indirect training pilot program by training educators at historically black colleges and universities (HBCUs) and minority serving institutions (MSIs). During the educator pilot, participants received materials, training, and will be provided the necessary compute resources to enable them to run the bootcamp at their home institutions or alternatively to adapt the content to implement within existing courses, thereby extending the reach of this initiative. The GL4U training program provides undergraduate students from underrepresented groups the opportunity to learn about NASA and Space Biology, and to enhance their career prospects by gaining hands-on experience analyzing omics data, a skillset that is highly applicable and marketable in the life sciences. Pre- and post-bootcamp surveys were completed by all participants and show the overwhelming success of the bootcamps.

Amanda M. Saravia-Butler

Increasing accessibility to deep learning-based analytics for space biology: pretrained models, transfer learning, and analytics platform development

Biological systems react in complex ways to the stressors of spaceflight, and the data capturing these relationships is concomitantly high-dimensional and complex. Deep learning and machine learning approaches are increasingly popular as an analytical approach for space biosciences, due to their ability to model complex relationships in complex data. However, such approaches often require large datasets and extensive computational resources. New approaches that minimize data sizes and computational power needed to leverage machine learning, and resources that make these approaches accessible, are needed to increase accessibility and adoption of machine learning in the space biosciences. Transfer learning, in which a pretrained model of broad utility is trained on a large dataset, and subsequently reused on downstream applications for which data is more limited, is one approach to minimizing data and computational intensity of deep learning applications. This transfer learning approach results in more performant models in high-dimensional, low-sample-size settings such as space biology, as compared to training models on limited data from scratch. This presentation will outline efforts to generate pretrained models for the space biology community, and highlight transfer learning applications modeling microbial antibiotic resistance during spaceflight. Finally, in order to increase accessibility of these models and tools, as well as others, for the broader space biology community, we present a modeling and analysis platform facilitating machine learning applications in space biology. This platform streamlines machine learning training and analysis in a notebook format, facilitates download and use of space biology data from the NASA GeneLab database, and can be utilized on NASA-hosted servers or downloaded and hosted locally. This effort, as part of the AI4LS (Artificial Intelligence for Life in Space) working group, will increase accessibility, feasibility, and performance of machine learning approaches for the space biology community.

Adrienne Hoarfrost