Search NASA⌕ Search

SEARCH · Search NASA

Results for “GeneLab”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Spaceflight Biospecimen and Data Sharing in Support of Science Discovery and Exploration

For decades, NASA and international partners have conducted biological experiments in space to understand effects of spaceflight and address potential hazards. To enable spaceflight back to the Moon, and then to Mars and beyond, it is imperative to further understand basic science and health risks associated with spaceflight, along with developing countermeasures. The sending of experiments and organisms into space is a costly endeavor. To maximize scientific return, sharing with the scientific community both space-flown biospecimens and data from completed experiments is essential. New fundamental, applied, and bioinformatic science insights can be gained from specimen and data sharing efforts. Data reuse enables spaceflight health risk modeling, analyzing adverse outcomes across spaceflight hazards, and deep space autonomous support for the flight medical officer. Space-flown biospecimens not required by mission Principal Investigators are regularly archived and made available for scientific request. The largest biorepository of these samples are found within NASA’s Institutional Scientific Collection at Ames Research Center (ISC-ARC), which stores over 32,000 specimens mostly from Shuttle and International Space Station (ISS) missions, but also some ground-based analog samples. The Ames Life Sciences Data Archive manages the ISC-ARC. Tissues are predominantly from mice and rats, though samples are also available from bacteria and quail. Only a handful of other similar collections exist worldwide. Rodent biospecimens exposed to simulated space radiation at Brookhaven National Laboratory are archived under the purview of NASA HRP Space Radiation Element. Microbial collection and analyses from 20 years of routine environmental monitoring of air, surfaces, and water systems of the ISS were performed to ensure a safe environment for astronauts. Samples from the ISC-ARC, space radiation and microbial collections are searchable and requestable through the NASA Life Sciences Data Archive (LSDA). Decades of planetary protection microbial isolates derived from spacecraft bioburden are archived in JPL’s microbial collection. Rodent biospecimens from spaceflight investigations conducted by the Japan Aerospace Exploration Agency (JAXA) are archived and available at the JAXA Biorepository at Tsukuba Space Center. The Russian Institute of Biomedical Problems also has a collection of animal, microbial, cellular, and fungi available for research from ground analog experiments. Several data repositories exist for scientists to utilize. The LSDA is the primary NASA source of life sciences research data and information. It contains decades of spaceflight and ground-analog research involving human, microbial, cellular, plant, and animal subjects. Data is collected from NASA-funded investigations through the Human Research Program and the Space Biology Program. The NASA Lifetime Surveillance of Astronaut Health collects and grants access to clinical and occupational health monitoring data from astronauts, with a list and description of data collected available for request through the LSDA. NASA GeneLab at ARC collects genomic, transcriptomic, proteomic, and metabolomic data from any species. It is a repository and platform for collaborative open-science bioinformatic approaches. JAXA is establishing an ‘omics-based repository in collaboration with the Tohoku Medical Megabank (ToMMo), called the JAXA-ToMMo Integrated Biobank for Space Life Science. Overall, the sharing of these biospecimen and data resources can assist researchers worldwide in understanding spaceflight effects on biology, along with enabling next generation data science applications for space exploration platforms. Websites: https://lsda.jsc.nasa.gov/ ; https://www.nasa.gov/ames/research/space-biosciences/isc-bsp ; https://www.nasa.gov/ames/research/space-biosciences/alsda

Ryan T. Scott↗

Overview of the Translational Radiation Research and Countermeasures (TRRaC) Project in Space Radiation

The Translational Radiation Research and Countermeasures (TRRaC) Project was initiated by the Space Radiation (SR) Element within NASA’s Human Research Program (HRP) to support the SR mission to understand and characterize the space radiation environment, understand and quantify radiation-associated risks, and mitigate the impacts of radiation-induced adverse health outcomes to enable human space exploration. TRRaC’s mission is to translate radiation research results from experimental studies and epidemiological data to humans and astronauts using bioinformatics and computational modeling. TRRaC will leverage existing datasets such as those available from NASA’s GeneLab repository and human medical radiation exposure registries, along with relevant data generated from ground-based research at molecular, cellular, tissue, and system levels, to: (1) refine the radiation dose rate effectiveness factor, (2) characterize radiation quality effects to improve estimates of the risk of exposure-induced death (REID) from space-relevant radiation exposures, and (3) identify pathways and biomarkers for cancer, cardiovascular disease, and central nervous system changes that are impacted by space radiation exposure.

Parastou Eslami↗

New developments in space radiation research at NASA: Annotating data using a novel radiation biology ontology

Like many interdisciplinary sciences, data producers and consumers in the field of radiation biology often use a wide variety of terminology to describe their experiments and data. Furthermore, space systems and technologies are rapidly evolving, and a shared understanding and common terminology for these is also lacking. The efficiency of research organizations can be enhanced by standardizing metadata through the use of knowledge resources like ontologies. Employing a sophisticated model such as a formal ontology to standardize metadata enables automated data acquisition processes and supports more complete, accurate meta-analysis through more efficient and complete data discovery and retrieval, particularly when using multiple data sources. Thus, we developed the Radiation Biology Ontology (RBO) in order to improved radiation biology metadata uniformity and transparency. We used open-source software (the Ontology Development Kit, Protégé and WebProtégé) and worked within the OBO Foundry framework, which includes a set of ontology development principles and practices for ontology consistency, uniformity, and accountability. The RBO has now been incorporated into two radiation research data repositories, NASA’s GeneLab omics database (https://genelab.nasa.gov), and the European Commission STORE database (https://www.storedb.org/). Continuous build integration tools allowed our international RBO collaboration to be more efficient and focus its efforts on semantic model design. Currently, the RBO contains over 300 annotated classes and individuals specific to the study of radiation on biological systems, as well as imports of many additional classes from other OBO Foundry ontologies that relate to and/or provide context for these RBO entities. We publish the RBO through the OBO Foundry, so that it is available for browsing, download, and querying through NCBI Bioportal web site and application programming interface. The NASA Ames Life Science Data Archive (ALSDA) is also in the process of adopting use of the RBO, taking NASA one step closer to a knowledge-based system for space biology data. It is our hope that the global communities of radiation research Investigators, data curators and data analysts can similarly leverage the RBO and will contribute to its further development.

radiation↗

Inception of a Spaceflight-specific Mouse to Human Expression Profiling Translation Model

Rodents are foundational model organisms often utilized due to their seemingly analogous morphologies and biological responses to humans. However, recent studies have demonstrated that murine model data are limited in their applicability, particularly in inflammatory disease. In space studies, accurately predicting human response from mouse data is critical due to extreme limiting factors in both rodent and human spaceflight research. With successful prediction, spaceflight ailments can be predicted and prevented while respecting the constraints of the spaceflight industry and minimizing danger to humans. To do so, novel methodologies must be developed that predict human response from murine data after considering biological differences between rodents and humans in spaceflight. After considering terrestrial models, we determined that a spaceflight-based expression profiting translation tool should be created to accurately capture predictions of human gene expression in spaceflight from mouse data. To prepare to build this model, we organized known human spaceflight risks, chose analog human diseases as training data categories, then identified existing RNASeq disease datasets from GEO as potential training data. In addition, we classified existing Genelab mouse differential gene expression datasets for use as experimental data.

Translation↗

INCREASING THE TRANSPARENCY AND REPRODUCIBILITY OF SPACE RADIATION SCIENCE: THE RADIATION BIOLOGY ONTOLOGY

Among the primary objectives of the Open/Open-Source Science paradigm are making scientific investigation data transparent and results reproducible [1], objectives shared by the FAIR principles [2]. To accomplish this, the conceptual framework that includes all the investigation objects needs to be accurately captured and communicated to all data consumers. A large part of this requires using metadata standards to annotate data collected. These standards should be readily accessible, informed by scientific community consensus and sufficiently specific to encompass all of the important aspects of the investigation. Starting in 2020 we have been co-leading an open consortium to develop a new metadata standard, the Radiation Biology Ontology (RBO), through the Open Biological and Biomedical Ontologies (OBO) Foundry [3]. We began by transforming many of the terms from the National Council on Radiation Protection and Measurement into concepts that can be formally related to existing OBO Foundry classes or attributes. We then identified and imported into the RBO existing OBO Foundry classes that have obvious relevance for radiation biomedicine (for example, concepts from the Environment Ontology that describe radiative processes, and concepts from the Gene Ontology dealing with molecular and cellular responses to radiation). Finally, we scrutinized datasets from investigations of radiation effects held in NASA GeneLab and LSDA repositories and added additional classes, instances, and attributes into the RBO that should be used to annotate these data. We developed the RBO using the open-source tools of GitHub and publish the RBO periodically through the NIH/NCBI BioPortal website, so systems worldwide can leverage the knowledge it contains [4]. This initial phase of concept modeling has yielded an RBO that at present has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies. While this first phase has focused on concepts for annotating samples, environments, exposures, and measurements, the next phase will center on supporting annotation of results and findings, such as concept models of molecular, cellular and tissue effects. The value of the RBO will be determined in part by our ability to engage the community in its development, and we have established a Radiobiology Informatics Consortium with unrestricted membership as the owner of the RBO in order to encourage investigators, system owners and other to join in this effort. Anyone can report issues or request new concept modeling or other features directly on GitHub. By using the BioPortal application programming interface, systems can pose dynamic queries to the latest version of the RBO for information on individual classes or entire hierarchies; this design eliminates the need for systems to be updated in order to use newer versions of the RBO. We hope to contribute to the advancement of open radiobiological science through the continued, open development of the RBO, that will provide more precise, machine-interpretable descriptions of investigations, as well as support data meta-analysis through machine learning or other artificial intelligence methods. REFERENCES [1] Open science in space. Nature Medicine, 2021. 27(9): p. 1485-1485. [2] Wilkinson, M.D., et al., The FAIR Guiding Principles for scientific data management and stewardship. Sci Data, 2016. 3: p. 160018. [3] Smith, B., et al., The OBO Foundry: coordinated evolution of ontologies to support biomedical data integration. Nat Biotechnol, 2007. 25(11): p. 1251-5. [4] Whetzel, P.L., et al., BioPortal: enhanced functionality via new Web services from the National Center for Biomedical Ontology to access and use ontologies in software applications. Nucleic Acids Res, 2011. 39(Web Server issue): p. W541-5.

informatics↗

Expanding Repository Data Available For Sharing and Knowledge Discovery

Some of the hardest space biology and space health challenges require data-intensive, bioinformatic, meta-analytical, and computer-assisted research approaches. These challenges include examining interdisciplinary space life science research across experiments and across interacting spaceflight hazards (radiation, altered gravity, confinement, hostile-closed environments, distance-duration from Earth). The approaches to confront these challenges involve mining multiple datasets simultaneously from various hierarchical organizations of biological complexity, all while concurrently evaluating how experimental design factors affect endpoints of standard assays. To enable this field, it is essential that principal investigators (PIs) submit data in a structure so it can be maximally re-used. The purpose of the NASA Ames Life Sciences Data Archive (ALSDA) is to collect, curate, and make publicly available all non-human space-relevant biological data. ALSDA must also ensure data are open-access, and maximally findable, accessible, interoperable, and reusable (FAIR). The scope of ALSDA data collected and submitted by PIs include subject and study design metadata, assay metadata parameters, raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). ALSDA recently integrated into a collaborative group of Open Science projects to facilitate a suite of new tools and workflows that will improve data submission, accessibility, and reusability by implementing digital data submission agreements, and adopting the data management system originally developed by NASA GeneLab. ALSDA intends to bring current biological repository data and all future collected data into this new scientific data reuse reality. This new suite of tools will enable ALSDA to deploy a science curation system using scientific assay configurations for the data submission portal. It will capture essential assay parameters according to established standards in each sub-field within biology. The submission portal expedites data collection by enhancing ease of PI data submission, providing a user interface and specificity for which data is to be submitted. Data submissions can be brought into cutting-edge informatic analysis portals to enable mining of physiological, behavioral, biochemical, and imaging datasets in conjunction with ‘omics-level datasets. As ALSDA datasets are submitted, curated, and published (e.g., micro-computed tomography, histology, pulse oximetry, serum metabolites, magnetic resonance imaging, intraocular pressure, novel object recognition, etc.), the merging together of spaceflight data along this multi-hierarchical complexity of biology will enable informatics and data-intensive approaches resulting in knowledge discoveries across missions, space hazards, and biological disciplines.

Biology↗

Understanding amyloids to prevent biofilm formation in space

There is a pressing need to search for novel approaches to combat biofilm formation, both in space and in medical applications. Many proteins have the ability to form ordered aggregates called amyloids. Amyloids are known to be an important part of biofilms. The use of anti-amyloid drugs is a novel venue for the development of antimicrobial agents. The ultrastructure of the amyloid aggregate shows a high packing of proteins, the second-order structure of which is dominated by β-sheets. The ability to form an amyloid aggregate is especially typical for proteins containing domains (protein fragments) with sufficient lability to arrange themselves in a tight β-sheet structure. Bioinformatics tools allow the prediction of such behavior of proteins in genomic data. We use GeneLab data of microbial populations identified aboard the International Space Station and other spacecraft to look for bacterial species that utilize amyloid aggregation in biofilm formation. We use a combined bioinformatic approach with a relatively high throughput molecular biology assay and biophysical assays to evaluate the anti-amyloid anti-biofilm approach. The significance of the research extends from understanding basic microbial community responses to spaceflight, to biofouling of the built environments in space as well as the long-term health of astronauts. Bioinformatics shows that onboard the ISS, bacterial species produce far more amyloid and prion proteins than are currently verified, hence their role in bacterial ecosystems is largely unknown. As we propose there is a link between amyloid formation in space and biofilm production, this research should lead to new paths for biofilm remediation in space.

Tomasz Zajkowski↗

INCREASING THE TRANSPARENCY AND REPRODUCIBILITY OF SPACE RADIATION SCIENCE: THE RADIATION BIOLOGY ONTOLOGY

Among the primary objectives of the Open/Open-Source Science paradigm are making scientific investigation data transparent and results reproducible [1], objectives shared by the FAIR principles [2]. To accomplish this, the conceptual framework that includes all the investigation objects needs to be accurately captured and communicated to all data consumers. A large part of this requires using metadata standards to annotate data collected. These standards should be readily accessible, informed by scientific community consensus and sufficiently specific to encompass all of the important aspects of the investigation. Starting in 2020 we have been co-leading an open consortium to develop a new metadata standard, the Radiation Biology Ontology (RBO), through the Open Biological and Biomedical Ontologies (OBO) Foundry [3]. We began by transforming many of the terms from the National Council on Radiation Protection and Measurement into concepts that can be formally related to existing OBO Foundry classes or attributes. We then identified and imported into the RBO existing OBO Foundry classes that have obvious relevance for radiation biomedicine (for example, concepts from the Environment Ontology that describe radiative processes, and concepts from the Gene Ontology dealing with molecular and cellular responses to radiation). Finally, we scrutinized datasets from investigations of radiation effects held in NASA GeneLab and LSDA repositories and added additional classes, instances, and attributes into the RBO that should be used to annotate these data. We developed the RBO using the open-source tools of GitHub and publish the RBO periodically through the NIH/NCBI BioPortal website, so systems worldwide can leverage the knowledge it contains [4]. This initial phase of concept modeling has yielded an RBO that at present has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies. While this first phase has focused on concepts for annotating samples, environments, exposures, and measurements, the next phase will center on supporting annotation of results and findings, such as concept models of molecular, cellular and tissue effects. The value of the RBO will be determined in part by our ability to engage the community in its development, and we have established a Radiobiology Informatics Consortium with unrestricted membership as the owner of the RBO in order to encourage investigators, system owners and other to join in this effort. Anyone can report issues or request new concept modeling or other features directly on GitHub. By using the BioPortal application programming interface, systems can pose dynamic queries to the latest version of the RBO for information on individual classes or entire hierarchies; this design eliminates the need for systems to be updated in order to use newer versions of the RBO. We hope to contribute to the advancement of open radiobiological science through the continued, open development of the RBO, that will provide more precise, machine-interpretable descriptions of investigations, as well as support data meta-analysis through machine learning or other artificial intelligence methods.

knowledge↗

Expanding Repository Data Available For Sharing And Knowledge Discovery

Some of the hardest space biology and space health challenges require data-intensive, bioinformatic, meta-analytical, and computer-assisted research approaches. These challenges include examining interdisciplinary space life science research across experiments and across interacting spaceflight hazards (radiation, altered gravity, confinement, hostile-closed environments, distance-duration from Earth). The approaches to confront these challenges involve mining multiple datasets simultaneously from various hierarchical organizations of biological complexity, all while concurrently evaluating how experimental design factors affect endpoints of standard assays. To enable this field, it is essential that principal investigators (PIs) submit data in a structure so it can be maximally re-used. The purpose of the NASA Ames Life Sciences Data Archive (ALSDA) is to collect, curate, and make publicly available all non-human space-relevant biological data. ALSDA must also ensure data are open-access, and maximally findable, accessible, interoperable, and reusable (FAIR). The scope of ALSDA data collected and submitted by PIs include subject and study design metadata, assay metadata parameters, raw and processed assay data, assay imagery/video, and subject-experienced mission data telemetry (radiation, temperature, humidity, acoustics, vibrations, etc.). ALSDA recently integrated into a collaborative group of Open Science projects to facilitate a suite of new tools and workflows that will improve data submission, accessibility, and reusability by implementing digital data submission agreements, and adopting the data management system originally developed by NASA GeneLab. ALSDA intends to bring current biological repository data and all future collected data into this new scientific data reuse reality. This new suite of tools will enable ALSDA to deploy a science curation system using scientific assay configurations for the data submission portal. It will capture essential assay parameters according to established standards in each sub-field within biology. The submission portal expedites data collection by enhancing ease of PI data submission, providing a user interface and specificity for which data is to be submitted. Data submissions can be brought into cutting-edge informatic analysis portals to enable mining of physiological, behavioral, biochemical, and imaging datasets in conjunction with ‘omics-level datasets. As ALSDA datasets are submitted, curated, and published (e.g., micro-computed tomography, histology, pulse oximetry, serum metabolites, magnetic resonance imaging, intraocular pressure, novel object recognition, etc.), the merging together of spaceflight data along this multi-hierarchical complexity of biology will enable informatics and data-intensive approaches resulting in knowledge discoveries across missions, space hazards, and biological disciplines.

life science↗

Open Science for Life in Space: Bioimaging, Data Sharing, and Tools for Knowledge Discovery

Precious space-flown biological experiments have both multi-omic and phenotypic data which NASA strives to make maximally open access for reuse. Currently a number of these space-relevant bioimaging datasets are being reused for AI/ML approaches. NASA Ames Life Science Data Archive and NASA GeneLab are working to make all current and future bioimaging data even more accessible and reusable. Standards for collection and curation are being implemented to enable scientists worldwide access to these data for further discovery and use.

data science↗

Looking into Ocular Risks of Spaceflight through the Mouse Retina

Ocular alterations have been observed at anatomical levels in astronauts on long duration spaceflight missions, such as what would be required for missions to Mars. These alterations cause an array of signs which together constitute the Spaceflight-Associated Neuro-ocular Syndrome (SANS), one of the top risk priorities of the NASA Human Research Program. Not much is known about SANS at the cellular and molecular level, but studies in mice and rats have recently begun to yield observations on how the spaceflight environment might affect the eye’s biology. Preliminary data from shuttle mouse experiments, and more recently experiments on ISS, have shown changes in retinal physiology via histology and gene expression analysis. This study utilizes samples from the CASIS sponsored Rodent Research 8 Experiment (RRRM-1) tissue sharing opportunity, delivered to the ISS by SpaceX CRS-16 on 12/08/2018. Female BALB/cAnNTac mice were on the ISS for 45 days, while ground controls consisted ofa standard vivarium group and spaceflight habitat group. Here we investigate the molecular response of the mouse retina to identify genes and pathways affected by spaceflight conditions using histology and transcriptomic RNAseq data. This Differentially Expressed Gene (DEG) data was used for pathway analysis with Galaxy (Genelab) and Ingenuity Pathway Analysis (IPA). We identified pathways related to neuronal differentiation, cellular transport/movement, and wound healing. Some of the top DEGs have known relation to ophthalmic diseases. Though there were DEGs throughout the comparisons we tested, there was no clear effect of spaceflight. This could be due to sample processing, which required mice to be returned to Earth about a day before they were sacrificed, possibly allowing for readaptation affecting the retinal transcriptome. However, there was a clear effect of age, between the young (10-12 weeks) and old (32 weeks) groups, and between the baseline and end of experiment, about 46 days.

SANS↗

Translational Line of Sight: a multi-omics longitudinal study of the murine retinal response to spaceflight hazard analogs

The space environment includes unique hazards like radiation and microgravity which adversely affect physiology and behavior of humans and rodent models. To better characterize the retinal response to spaceflight, we assessed a multi-omics NASA GeneLab dataset where 6-month-old female mice were gamma irradiated and/or hindlimb unloaded for 21 days followed by whole transcriptome shotgun sequencing (RNA-Seq) and reduced representation bisulfite sequencing (RRBS) of retina samples collected at 7 days, 1 month or 4 months post-exposure. We compared time-matched epigenomic and transcriptomic retinal profiles revealing a total of 4,178 differentially methylated loci or regions, and 457 differentially expressed genes. Highest correlation in methylation differences was seen across different conditions at the same time point (e.g., between radiation exposure and hindlimb unloaded at 7 days). Biological processes related to nucleotide metabolism were enriched in all groups with activation at 1 month and suppression at 7 days and 4 months. Genes and processes related to Notch and Wnt signaling showed alterations 4 months post-exposure. Interestingly, Notch3 and Lrg1 showed differential patterns in the NASA Twins Study in-flight samples and in response to stressors in the murine retina in the current study. A total of 23 genes were both differentially methylated and expressed, including genes involved in retinal disease or cataract development (Crybb3, Fgfr1, Pitpnm3, Sipa1l3, Sox9) and inflammatory response (B4galt6, Ppm1a, Sphk1). To our knowledge, the current multi-omics analysis is the first multi-omics study to interrogate the epigenomic and transcriptomic impacts of radiation and hindlimb unloading on the retina in isolation and in combination. The results provide an insight into the retinal response to individual spaceflight hazard analogs and their interplay at different post-exposure stages and contributes towards a mechanistic understanding of spaceflight-induced vision impairment using ground-based models.

Prachi Kothiyal↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), has typically limited machine learning (ML) in space studies and further study of radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNAseq) data from 6 mouse liver GeneLab datasets (GLDS) with a total of 113 spaceflight and ground-control samples to determine top features relevant to spaceflight including the effect of radiation exposure. Data was normalized within each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. The top MRMR features were used to predict spaceflight vs. ground-control samples using a Random Forest (RF) classifier with 5-fold cross validation (CV). The ML-based gene sets were further compared against differential gene expression results from individual GLDS. CV training using the top 100 MRMR genes show averages of 86% accuracy and 0.95 AUC value on the validation set over 5 folds (Figure 1A). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 811 or 68 DEGs overlapping between at least 2 or 3 studies, respectively (Figure 1B). Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism. Set analysis between the MRMR features and the DEGs showed 60 or 8 genes overlapping with at least 1 or 2 studies, respectively. MRMR feature selection and ensemble ML methods (e.g. RF) improve performance relative to a Naïve Bayes classifier when NGS data sets are analyzed. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise ratio. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from RNASeq analysis. Non-intersecting sets introduce opportunity to explore spaceflight relevant genes and implementing ML methods across existing NGS datasets may overcome sample size limitations. ML coupled with existing analytical methods enhances understanding of disease by revealing common underlying pathways across datasets.

Machine Learning↗

Mining the Gravity Mutants of Arabidopsis

Gravity mutants are a valuable resource for understanding gravity perception, signaling, and response in plants. A review of 190 publications resulted in a list of 97 loci with mutations that caused a gravitropic phenotype. While all these mutants show some form of gravity phenotype, several are also generally defective in growth. After removing these nonspecific mutants, 76 loci were deemed to have true gravity mutants. The gene list was then used to create a sortable database containing key factors of gravity signaling as well as data on methodologies and the development of mutant lines. Indexing this database has allowed us to pull out trends that were not visible in individual publications. For example, mutants that are unable to rearrange starch statoliths make up a larger proportion of the described inflorescence stems mutants than any other organ. Comparing these statolith mutants across organs shows that root and inflorescence stems consistently display different phenotypes for the same class of defect. The more severe gravity defects are most commonly described in root tissue while other tissues are more likely to show only a delayed or reduced response to gravity. The experience of members of the GeneLab plant Analytics Working Group (AWG) in data visualization and gene mapping have sparked new ideas for utilizing this database. Knowledge of shared traits, mutant development, and growth can provide a new resource for the production of seed lines specialized for their response to gravity.

gravitropism↗

Increasing accessibility to deep learning-based analytics for space biology: pretrained models, transfer learning, and analytics platform development

Biological systems react in complex ways to the stressors of spaceflight, and the data capturing these relationships is concomitantly high-dimensional and complex. Deep learning and machine learning approaches are increasingly popular as an analytical approach for space biosciences, due to their ability to model complex relationships in complex data. However, such approaches often require large datasets and extensive computational resources. New approaches that minimize data sizes and computational power needed to leverage machine learning, and resources that make these approaches accessible, are needed to increase accessibility and adoption of machine learning in the space biosciences. Transfer learning, in which a pretrained model of broad utility is trained on a large dataset, and subsequently reused on downstream applications for which data is more limited, is one approach to minimizing data and computational intensity of deep learning applications. This transfer learning approach results in more performant models in high-dimensional, low-sample-size settings such as space biology, as compared to training models on limited data from scratch. This presentation will outline efforts to generate pretrained models for the space biology community, and highlight transfer learning applications modeling microbial antibiotic resistance during spaceflight. Finally, in order to increase accessibility of these models and tools, as well as others, for the broader space biology community, we present a modeling and analysis platform facilitating machine learning applications in space biology. This platform streamlines machine learning training and analysis in a notebook format, facilitates download and use of space biology data from the NASA GeneLab database, and can be utilized on NASA-hosted servers or downloaded and hosted locally. This effort, as part of the AI4LS (Artificial Intelligence for Life in Space) working group, will increase accessibility, feasibility, and performance of machine learning approaches for the space biology community.

Adrienne Hoarfrost↗

Space Flown Rodent Liver RNA Sequencing Data for Machine Learning in Space Biology Research

High-throughput nucleic acid sequencing (DNA-seq, RNA-seq) has become widespread in biomedical research due to the growing availability and affordability of these assays. Data analysis has been accelerated in recent years by the adoption of artificial intelligence (AI) and machine learning (ML) techniques by biomedical researchers. In space biology research, RNAseq datasets from space-flown experimental samples are critical for characterizing the gene expression aberrations associated with exposure to spaceflight stressors. However, space biological experiments tend to be very low sample size, so identifying proper AI/ML algorithms for sequencing data analysis is an ongoing challenge since these algorithms typically require large sample size. The NASA Science Mission Directorate (SMD) has started the “Benchmark Initiative for AI/ML”, focused on creating datasets meant for three main applications: 1) scientific benchmarking, which finds the best algorithm for a specific problem; 2) application benchmarking, which measures algorithm performance against a set of parameters; and 3) system benchmarking, which evaluates performance of hardware and software architecture. These scientific benchmarks consist of an AI-ready dataset and a reference implementation on a specific scientific question. In this work, we focused on generating standardized datasets to allow the scientific community to benchmark AI/ML algorithms in the domain of space biology. We present here a standardized, AI-ready, publicly available benchmark dataset for space biology RNA-seq data as a collaboration between the NASA AI4LS (Artificial Intelligence for Life Sciences) working group. and NASA’s SMD. This dataset consists of space-flown and ground control mouse liver found in the NASA GeneLab omics database. However, to amplify the small sample number (n=112 samples) for ML purposes, we employ Gaussian noise and a generative adversarial network to extend this dataset to 6,000 synthetic samples, matching the original gene expression characteristics.

James Casaletto↗

MULTI-OMICS ANALYSIS OF THE IMPACT OF CHRONIC LOW-DOSE RADIATION AND HINDLIMB SUSPENSION ON MURINE BRAIN AND RETINA

The space environment includes hazards like radiation and microgravity which can adversely affect biological systems. We assessed multi-omics multi-tissue NASA GeneLab datasets where 6-month-old female mice were gamma irradiated (IR) and/or hindlimb unloaded (HLU) for 21 days. Whole transcriptome shotgun sequencing (RNA-Seq) and reduced representation bisulfite sequencing (RRBS) of brain and retina samples collected at 4 months post-exposure was performed to better characterize the retinal and neurological responses to spaceflight. We compared epigenomic and transcriptomic profiles within each exposure group for both tissue types to identify correlation (Pearson’s correlation test; p-value < 0.05) between gene expression and DNA methylation levels that may be related to transcriptional regulation. We then obtained genes with methylation-expression correlation that also showed differences in mean expression or dispersion between exposed and control groups (adjusted p-value < 0.25; relaxed to denote ‘hypothesis’) in the brain (37 genes in HLU, 4 in IR, and 156 in HLU+IR) or retina (92 genes in HLU, 1 in IR, and 55 in HLU+IR). Enriched Gene Ontology (GO) terms for these genes are listed in Table 1 for HLU and HLU+IR for both tissue types. No enriched terms and only a few genes were detected with IR-only exposure in the brain (Chmp1a, Limd1, Rab40b, Ubc) and retina (retinoblastoma binding protein Rbbp7). Cellular components related to synapse were enriched in both tissue types. Previous analysis of differentially expressed genes in the retina after 1 month of HLU+IR showed enrichment in the somatodendritic compartment of the neuron, which was also observed in the brain 4 months post-exposure. Interestingly, genes related to ubiquitination showed correlation between methylation and expression and were differentially expressed or dispersed in different exposure groups indicating that this pathway may play an important role in multi-stressor response (Figure 1). The current multi-omics and multi-tissue analysis interrogates the epigenomic and transcriptomic impacts of radiation and hindlimb unloading, in isolation and in combination, on the retina and the brain. The results provide insight into the adaptive response to individual spaceflight hazard analogs and their interplay, as well as hypotheses to be further tested for understanding spaceflight-induced neurological and vision effects.

Prachi Kothiyal↗

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗