Search NASASearch

SEARCH · Search NASA

Results for “omics data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

GL4U: Training the next generation of bioinformaticians, one omics datatype at a time

Spaceflight modifies gene expression in every organism examined to date, including humans. Understanding how these gene expression changes affect physiology is crucial for the development of countermeasures to enable long-duration manned missions. NASA’s GeneLab project provides researchers open access to multi-omics data, including genetic and gene expression data, from spaceflight experiments that can be mined to understand the effects of spaceflight on biological systems. To ensure new knowledge generation through data re-use, it is important to maximize the number of scientists who utilize GeneLab data. Training students on the GeneLab platform is the best way to create long-term adopters of this NASA database and its tools. Turning students into future instructors and advocates will also accelerate the dissemination of these data and tools to the broader scientific community. Therefore, in collaboration with the GeneLab Educational Working Group (EWG), GeneLab has created GeneLab for Colleges and Universities (GL4U). GL4U provides space biology-relevant training in bioinformatics to the next generation of scientists through direct and indirect approaches. The GeneLab team plans to host two annual data processing bootcamps, one for college-level students (direct) and one for college educators (indirect – training of trainers), in which participants learn to analyze GeneLab’s space-relevant omics data. During the bootcamp, educators will receive materials and training to enable them to run the bootcamp at their home institutions or alternatively to adapt the content to implement within existing courses, thereby extending the reach of this initiative. The GL4U direct training pilot program was conducted in June 2021 in collaboration with USRA and San Jose State University (SJSU). During the pilot, SJSU students participated in a week-long bootcamp consisting of space biology-specific lectures and hands-on instruction using Jupyter Notebooks to analyze RNA sequence data. This pilot demonstrates the capacity of GL4U for training young scientists and encouraging data re-use.

Jonathan Matthew Galazka

Prediction of non-intuitive metabolic targets with bayesian metabolic control analysis to improve 3-hydroxypropionic acid production in Aspergillus niger

Development of efficient bioconversion processes is limited by the ability to predictably improve metabolic flux. Here we deployed Bayesian Metabolic Control Analysis as a platform to integrate multi-omics data with metabolic modeling and evaluated its ability to predict genetic interventions that improve metabolic flux. Global Metabolomics and proteomics data was collected from 17 Aspergillus niger strains engineered to produce the platform biochemical 3-hydroxypropionic acid from which seven actional genetic interventions were predicted from significant flux control coefficients. Of the suggested genetic interventions, two were present within the intuitively designed strains used for training (malonic semialdehyde dehydrogenase and pyruvate carboxylase) while five predicted targets were present within non-intuitive areas of the metabolic network including 5-formyltetrahydrofolate deformylase and four mitochondrial enzymes, alcohol dehydrogenase, succinyl-CoA ligase, aspartate aminotransferase, and malate dehydrogenase. Six of the targets were validated in the highest performing 3-HP strain used for multi-omics data generation which contained a prior disruption of the highest scoring target malonic semialdehyde dehydrogenase. Predicted directional perturbation of five of the six tested targets significantly improved titer and rate of 3-HP production and two significantly improved yield. The greatest improvements were observed following disruption of the non-intuitive target succinyl-CoA ligase which increased titer by 39% and yield by 29% (to 20.4 g/L 3-HP and 0.31 g 3-HP/g glucose) over the strains used for training. This study demonstrates the utility of Bayesian Metabolic Control Analysis and highlights the ability to predict meaningful genetic targets in unexpected areas of metabolism to improve engineered strains for bioconversion.

3-hydroxypropionic acid

NASA GeneLab: Open Science for Life in Space

The NASA GeneLab project (genelab.nasa.gov) seeks to get the most from space-relevant biology experiments by providing and maintaining a public database consisting of DNA, RNA, protein, and metabolite data from spaceflight experiments. Since these types of data, referred to as omics data, are difficult to understand for non-bioinformaticians, the GeneLab data processing team works with the scientific community to develop methods to process these data. The processed data found on GeneLab reveals information about which genes are turned on and turned off in the space environment, which helps us understand how space changes our biology and how we can best mitigate these effects to travel deeper into space.

Jonathan Oribello

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge, reduce risk, and support safe, productive human space missions. Through the powerful emerging computer science approaches of artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in biomedical science and engineered astronaut health systems, to enable Earth-independence and autonomy of mission operations. We present a decadal view of AI/ML architecture to support deep space mission goals, developed in concert with leaders in the field. We describe current AI/ML methods to support 1) fundamental biology, 2) in situ analytics, 3) high performance computing hardware, 4) automated science, 5) self-driving labs, 6) remote data management, 7) integrated real-time mission biomonitoring, and 8) a Precision Space Health system. Cutting-edge AI/ML approaches that can be integrated to support these domains include active learning, explainable AI, adaptive learning, causal inference, knowledge graphs, federated learning, transfer learning, and large language models. Finally, we present results from several current ML projects that are underway in the field to address key challenges of small sample n, high feature count, heterogeneity, and sparse data. These include 1) connecting omics data to phenotypic data using an ensemble model to infer causality of spaceflight rodent liver health disruption, 2) usage of explainable ML to interrogate the muscular underpinnings of spaceflight muscle atrophy, 3) ML models analyzing and determining directed acyclic graphs of human space health risk leveraging rodent bone datasets, 4) usage of large pre-trained models connecting biomedical knowledgebases with small spaceflight datasets to understand gene-to-gene interaction networks, and 5) a suite of benchmarked open science datasets (spaceflight mouse liver; radiation DNA damage) enabling programmers to identify the best ML algorithms to answer space biological science questions.

space biology

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge, reduce risk, and support safe, productive human space missions. Through the powerful emerging computer science approaches of artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in biomedical science and engineered astronaut health systems, to enable Earth-independence and autonomy of mission operations. We present a decadal view of AI/ML architecture to support deep space mission goals, developed in concert with leaders in the field. We describe current AI/ML methods to support 1) fundamental biology, 2) in situ analytics, 3) high performance computing hardware, 4) automated science, 5) self-driving labs, 6) remote data management, 7) integrated real-time mission biomonitoring, and 8) a Precision Space Health system. Cutting-edge AI/ML approaches that can be integrated to support these domains include active learning, explainable AI, adaptive learning, causal inference, knowledge graphs, federated learning, transfer learning, and large language models. Finally, we present results from several current ML projects that are underway in the field to address key challenges of small sample n, high feature count, heterogeneity, and sparse data. These include 1) connecting omics data to phenotypic data using an ensemble model to infer causality of spaceflight rodent liver health disruption, 2) usage of explainable ML to interrogate the muscular underpinnings of spaceflight muscle atrophy, 3) ML models analyzing and determining directed acyclic graphs of human space health risk leveraging rodent bone datasets, 4) usage of large pre-trained models connecting biomedical knowledgebases with small spaceflight datasets to understand gene-to-gene interaction networks, and 5) a suite of benchmarked open science datasets (spaceflight mouse liver; radiation DNA damage) enabling programmers to identify the best ML algorithms to answer space biological science questions.

space biology

Murine Host-gut Microbiota Interactions are Modulated During Spaceflight

The rodent habitat on the International Space Station has provided critical insight into the impact of spaceflight on mammalian physiology. These effects include dysfunction of carbohydrate, steroid and lipid metabolism, and immune response, as well as induction of symptoms characteristic of liver disease, insulin resistance, osteopenia and myopathy, which are anticipated to intensify over long-duration spaceflight. Although these physiological responses can involve the microbiome, the host-microorganism interactions during spaceflight are still largely unknown. NASA GeneLab curates a wide range of space research data and the current work harnesses GeneLab multi’omic data from recent Rodent Research studies to explore changes to gut microbiota during spaceflight and their associations with host physiology when compared to ground controls. Using a hybrid analysis of DNA barcoding and whole genome shotgun data, an array of bacteria, fungi and nematodes could be identified at species level, and significant differences in relative abundances associated with spaceflight. Functional prediction based on differential abundance of species and metagenome gene inventories as well as metatranscriptomic gene expression at the host-gut microbiome interface implicate microbiota interactions could contribute to spaceflight pathology. Harnessing carefully curated publicly available data, such as from Genelab, to generate multi‘omic space science discoveries can help decipher the complex host-microbiome interactions that influence both health on Earth and the feasibility of long-duration spaceflight.

Microbiome

Carbon source–driven metabolic and regulatory remodeling defines phenomic states in Lipomyces starkeyi

Lipomyces is a genus of oleaginous yeasts with potential for contributing to reliable biomanufacturing supply chains. However, progress in advanced strain designs and engineering efforts are still constrained by a lack of understanding of the underlying molecular drivers of Lipomyces phenotypes. To address this gap, we collected a suite of multi-omic data to dissect how carbon source availability reshapes the metabolic network, lipid allocation, and regulatory architecture of Lipomyces starkeyi. We observed that glucose promotes biosynthetic and proliferative processes supported by abundant energy and carbon intermediates, xylose enhances redox-balancing mechanisms centered on the pentose phosphate pathway, and glycerol activates respiratory metabolism, ß-oxidation, and the glyoxylate cycle. Lipid species distributions remained consistent in both nitrogen replete and depleted conditions across the carbon sources, indicating robust production mechanisms. Regulatory protein identification and network analysis revealed glycerol-driven respiratory growth favors regulatory programs integrating stress tolerance, redox balance, and lipid-associated metabolism, whereas xylose growth activates compensatory transcriptional responses aimed at maintaining mitochondrial function. Nitrogen limitation modulates the strength of these responses but does not fundamentally alter their direction, reinforcing carbon source as the dominant driver of regulatory architecture. Taken together, this data enhances the understanding of Lipomyces molecular rearrangements and provides a foundation for further development of predictive phenotypic tools in this genus.

Biotechnology

GeneLab: Omics Database for Spaceflight Experiments

Motivation - To curate and organize expensive spaceflight experiments conducted aboard space stations and maximize the scientific return of investment, while democratizing access to vast amounts of spaceflight related omics data generated from several model organisms. Results - The GeneLab Data System (GLDS) is an open access database containing fully coordinated and curated "omics" (genomics, transcriptomics, proteomics, metabolomics) data, detailed metadata and radiation dosimetry for a variety of model organisms. GLDS is supported by an integrated data system allowing federated search across several public bioinformatics repositories. Archived datasets can be queried using full-text search (e.g., keywords, Boolean and wildcards) and results can be sorted in multifactorial manner using assistive filters. GLDS also provides a collaborative platform built on GenomeSpace for sharing files and analyses with collaborators. It currently houses 172 datasets and supports standard guidelines for submission of datasets, MIAME (for microarray), ENCODE Consortium Guidelines (for RNA-seq) and MIAPE Guidelines (for proteomics).

omics

Using the NASA GeneLab Data System to Study the Metagenomes of Spaceships and Their Occupants

With humans pushing to live further off Earth for longer periods of time, it is increasingly important to understand the changes that occur in biological systems during spaceflight whether these be astronauts, their microbial commensals, or their plant-based life support systems. In a three-part presentation, we discuss GeneLab and recent discoveries regarding the microbiota of spacecrafts and space-flown animals. Part 1: GeneLab: Open Science for Life in Space, Jonathan Galazka, NASA Ames Research Center To accelerate the pace of discovery from precious spaceflight biological experiments, NASA as develop the GeneLab data system (genelab.nasa.gov), which allows unfettered access to omics data from spaceflight and spaceflight relevant experiments. GeneLab houses metagenomic datasets from spacecraft and relevant spacecraft models. Users can download this data and associated metadata to make new discoveries about how microbial communities may change and adapt to spaceflight.

Galazka, Jonathan M.

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, there-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA's Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomatic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related 'omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata 'omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data-use, resulting in 40 enabled publications by open data. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA "Open Science Data Repositories (OSDR)" and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Flourescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to "big data" from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology.

omics

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, the re-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA’s Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related ‘omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata ‘omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data re-use, resulting in 38 additional publications derived from the original 67 publication over the past four years. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA “Open Science Data Repositories (OSDR)” and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Fluorescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to “big data” from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology. Several other talks will cover these topics in this conference.

life sciences

NASA Omics Archive Project

The space environment consists of a complex set of hazards including altered gravity, radiation, psychological/physiological stress, isolation, and confinement leading to complex biological responses. Advances in biotechnology capabilities offer considerable potential to provide novel insight into those responses as well as innovative diagnostic, treatment, and countermeasure solutions for astronauts as NASA begins to travel beyond low Earth orbit. Omics data (genomics, transcriptomics, proteomics, etc.) is one example that can provide NASA with critical knowledge of how a crewmember’s genetics, environment, and lifestyle can be used to develop individualized approaches for disease prevention, advance diagnostics, and improve treatment strategies. NASA ventured into the field of omics on human subjects with the successful completion of the NASA Twins Study which was the first step in mapping the multi-omic profile of astronauts to understand and mitigate the health consequences of spaceflight. The Human Research Program aims to build upon the success of the Twins Study with the NASA Omics Archive flight study, establishing a longitudinal biospecimen archive and efficiently generating a comprehensive high-quality multi-omic dataset from astronauts for the purpose of studying molecular, metabolic, and microbial changes associated with longduration spaceflight missions. The goal is to facilitate scientific and medical research community efforts to characterize and mitigate spaceflight health and performance risks. In this presentation, we will review details regarding the biospecimen and data archive to be generated by the NASA Omics Archive flight study. Data generated as part of this project will be archived in the NASA Life Sciences Portal (NLSP) and be made available for future hypothesis-driven research efforts or occupational surveillance through Institutional Review Board-approved data sharing and retrospective data requests submitted to the Life Sciences Data Archive (LSDA) team. We will also present results of a ground study performed to evaluate in-house procedures, new sample collection hardware, and vendor capabilities. The data repository generated and the samples to be archived by this study will enable future research efforts to assess an astronauts’ unique molecular and genetic profile with respect to individual spaceflight responses. Results of which will be instrumental in enabling precision health capabilities to better assess and mitigate spaceflight risks, detect disease states earlier, and actively monitor countermeasure treatments, ultimately improving clinical outcomes during future exploration class missions.

C. A. Theriot

NASA's GeneLab Phase II: Federated Search and Data Discovery

GeneLab is currently being developed by NASA to accelerate 'open science' biomedical research in support of the human exploration of space and the improvement of life on earth. Phase I of the four-phase GeneLab Data Systems (GLDS) project emphasized capabilities for submission, curation, search, and retrieval of genomics, transcriptomics and proteomics ('omics') data from biomedical research of space environments. The focus of development of the GLDS for Phase II has been federated data search for and retrieval of these kinds of data across other open-access systems, so that users are able to conduct biological meta-investigations using data from a variety of sources. Such meta-investigations are key to corroborating findings from many kinds of assays and translating them into systems biology knowledge and, eventually, therapeutics.

exobiology

NASAs GeneLab Phase II: Federated Search and Data Discovery

GeneLab is currently being developed by NASA to accelerate open science biomedical research in support of the human exploration of space and the improvement of life on earth. Phase I of the four-phase GeneLab Data Systems (GLDS) project emphasized capabilities for submission, curation, search, and retrieval of genomics, transcriptomics and proteomics (omics) data from biomedical research of space environments. The focus of development of the GLDS for Phase II has been federated data search for and retrieval of these kinds of data across other open-access systems, so that users are able to conduct biological meta-investigations using data from a variety of sources. Such meta-investigations are key to corroborating findings from many kinds of assays and translating them into systems biology knowledge and, eventually, therapeutics.

genome

Multi-Omics Analysis of Mouse Retina Following Low Dose Radiation and/or Hindlimb Unloading

Rodent models have been used as analogs for studying the effects of spaceflight. NASA’s GeneLab provides access to omics datasets generated from spaceflight and ground-based experiments allowing for additional retrospective analysis. We used GeneLab’s GLDS-203, a dataset generated by researchers at Loma Linda University to study the impact of prolonged unloading and/or low-dose radiation on mouse retina. The purpose of this study was to understand the effect of gamma radiation and/or hindlimb unloading on mice retinas through a multi-omics analysis. In the experiment that generated the omics data, mice were irradiated with gamma-ray and/or subjected to hindlimb unloading for 21 days and multi-omics analysis was performed at 7 days, 1 month, or 4 months post exposure. In the current study, for each of the three timepoints, we compared epigenomic profiles for retinas from exposed mice against timepoint-matched controls. We identified a total of 5,271 differentially methylated loci (DML) and 321 differentially methylated regions (DMR; using a sliding window and step size of 500 bp) with methylation difference > 10% and q-value < 0.05 (sliding linear model corrected p-value) across the nine exposure groups. Highest correlation in methylation difference was seen for significant DMLs (q-value < 0.05) across different conditions at same post exposure timepoint (Figure 1).The location of DMLs and DMRs were characterized with respect to CpG islands and shores, putative promoters, gene body, and intergenic regions (Table 1). We analyzed RNA-seq counts and performed gene set enrichment analysis using differential expression results from comparing each exposure group to its timepoint-matched control group. Significant pathways (adjusted p-value <0.05) enriched in all three microgravity-only groups were related to morphogenesis of a branching epithelium, skeletal muscle cell differentiation, and response to fibroblast growth factor. Common processes across all timepoints in the radiation-only groups were retina homeostasis, synaptic vesicle exocytosis-endocytosis, and chemotaxis. In the combination groups, regulation of trans-synaptic signaling, and Rho protein signal transduction were enriched at all three timepoints. Processes related to purine nucleotide metabolism were enriched in all nine exposure groups, with activation at 1 month, and suppression at 7 days and 4 months. A total of 14 genes contained at least one DML and were differentially expressed at adjusted p-value < 0.05, including genes implicated in cataract development (Sipa1l3, Crybb3) and those involved in cytoskeletal organization (Plec, Flnb, Eef1a1). This analysis is part of a larger effort to understand the molecular mechanisms following spaceflight exposures that can help translate effects observed in animal models to human impacts.

Prachi Kothiyal

GL4U: GeneLab for Colleges and Universities

GeneLab for Colleges and Universities (GL4U) will provide space biology-relevant training in bioinformatics to the next generation of scientists through direct and indirect approaches. The GeneLab (GL) team will host two annual data processing bootcamps, one for college-level students (direct) and one for college educators (indirect – Training of Trainers), in which participants learn to analyze space-relevant omics data hosted on GL. The first bootcamp took place in early June 2021 with about 30 SJSU undergraduate students and covered space biology-specific lectures and hands-on instruction using Jupyter Notebooks (JNs) for RNA sequence (RNAseq) data analysis. All training materials including the enclosed files listed below will be made publicly available on GitHub. RNAseq Bootcamp Lectures (attached in combined file): Introduction to NASA, Space Biology, GeneLab, and the Command Line: NASA_GL_CL_Intro_FINAL.pdf - DRAFT from initial submission NASA_SB_GL_CL_Intro_FULL.pdf - FINAL version presented during the bootcamp - only minor edits from the draft version RNAseq and Data Processing Overview: RNAseq_Overview_FINAL.pdf - DRAFT from initial submission RNAseq_Overview_FULL.pdf - FINAL version presented during the bootcamp - only minor edits from the draft version Overview of the Statistics Used for RNAseq Data Analysis: SJSU_Statistics_Intro_Lecture_FINAL.pdf - DRAFT from initial submission Statistics_Overview_FULL.pdf - FINAL version presented during the bootcamp - only minor edits from the draft version Completed JNs in HTML format (attached in combined file): Unix_Intro_JN_06-2021_completed.html R_Intro_JN_06-2021_completed.html RNAseq_fastq_to_counts_JN_06-2021_completed.html RNAseq_DGE_JN_06-2021_completed.html RNAseq Bootcamp Recordings (attached): GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day1_Part_1_of_5.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day1_Part_2_of_5.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day1_Part_3_of_5.mp4 *There were issues with the part 4 recording so that is not available GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day1_Part_5_of_5.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day2_Part_1_of_3.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day2_Part_2_of_3.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day2_Part_3_of_3.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day3_Part_1_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day3_Part_2_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day3_Part_3_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day3_Part_4_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day4_Part_1_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day4_Part_2_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day4_Part_3_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day4_Part_4_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day5_Part_1_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day5_Part_2_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day5_Part_3_of_4.mp4 GL4U_RNAseq_Bootcamp_June_2021_Pilot_Day5_Part_4_of_4.mp4

GeneLab

The need for standardization and improved open (meta)data practices in metaproteomics

Metaproteomics enables functional insight into microbial communities by identifying and quantifying proteins in complex samples. Yet, heterogeneous analytical workflows and the lack of standardization across experimental and bioinformatics stages hinder reproducibility and comparability, limiting integration with other omics data. We here present a community-developed reporting checklist tailored to the specific needs of metaproteomics. We also outline current efforts to enable structured and interoperable metadata capture, drawing on standards from proteomics and microbiome research wherever possible. By promoting transparent reporting and advancing metadata practices, our recommendations aim to align metaproteomics more closely with FAIR principles and support reproducible and interoperable research practices.

Armengaud, Jean [Universite Paris-Saclay, France]

Knowledge Network Embedding of Transcriptomic Data From Spaceflown Mice Uncovers Signs and Symptoms Associated With Terrestrial Diseases

There has long been an interest in understanding how the hazards from spaceflight may trigger or exacerbate human diseases. With the goal of advancing our knowledge on physiological changes during space travel, NASA GeneLab provides an open-source repository of multi-omics data from real and simulated spaceflight studies. Alone, this data enables identification of biological changes during spaceflight, but cannot infer how that may impact an astronaut at the phenotypic level. To bridge this gap, SPOKE, a heterogeneous knowledge graph connecting biological and clinical data from over 30 databases, was used in combination with GeneLab transcriptomic data from six studies. This integration identified critical symptoms and physiological changes incurred during spaceflight.

spaceflight