Search NASA⌕ Search

SEARCH · Search NASA

Results for “omics data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Space Radiation Induces Long Term Impact on the Cardiovascular System by the Activation of FYN Through Reactive Oxygen Species

Space radiation can damage the cardiovascular system and thus is an important health risk factor for astronauts during long-term space missions. We utilized publicly available transcriptomic data through NASA's GeneLab platform (genelab.nasa.gov) to determine cardiovascular system response to space radiation. GeneLab is an open repository that houses all NASA related omics experiments including on the International Space Station (ISS) and related radiation ground studies. We analyzed 3 datasets from GeneLab: GLDS-117 and GLDS-109, which are ground studies of cardiomyocytes followed-up for 28 days after exposure to 90cGy of proton at 1GeV and 15cGy of 56Fe at 1GeV; and GLDS-52, human endothelial cells (HUVECs) that were cultured for 10 days on the ISS. The ground studies were designed to characterize the long-term impact following space irradiation on cardiomyocytes for 5 different time points up to 28 days after irradiation. Our analysis was guided by the hypothesis that there are common persistent molecules affecting the cardiovascular system due to radiation effects during spaceflight. Endothelial cells are known to directly regulate the development and activity of cardiomyocytes, and thus their response to spaceflight should be highly correlated with cardiomyocytes. To investigate our hypothesis, we identified the molecular pathways that were modified for all time points compared across both radiation on the ground and the pathways found in HUVECs flown in the ISS. We found the following key results related to the cardiovascular systems: 1) space radiation downregulate ROS functions; and 2) the key/driving genes: FYN, LCK, AKT1 are upregulated and LYN and FOS are downregulated with FYN being the central driver/hub for the cardiovascular response to space radiation. It is worth noting the activation of FYN is a key event which prevents cardiac cell death and ROS production. From our study we thus hypothesize that a feedback loop occurs from the oxidative stress caused by space radiation that upregulates FYN which in turn reduces ROS levels and thus ROS pathways, preventing cardiomyocyte and endothelial cell death and thus protecting the cardiovascular systems. We believe that this is a novel mechanism for space radiation induced cardiovascular risk directly linking radiation ground studies to spaceflight

Beheshti, Afshin↗

MINE: a new way to design genetics experiments for discovery

Abstract The Maximally Informative Next Experiment or MINE is a new experimental design approach for experiments, such as those in omics, in which the number of effects or parameters p greatly exceeds the number of samples n (p > n). Classical experimental design presumes n > p for inference about parameters and its application to p > n can lead to over-fitting. To overcome p > n, MINE is an ensemble method, which makes predictions about future experiments from an existing ensemble of models consistent with available data in order to select the most informative next experiment. Its advantages are in exploration of the data for new relationships with n < p and being able to integrate smaller and more tractable experiments to replace adaptively one large classic experiment as discoveries are made. Thus, using MINE is model-guided and adaptive over time in a large omics study. Here, MINE is illustrated in two distinct multiyear experiments, one involving genetic networks in Neurospora crassa and a second one involving a genome-wide association study in Sorghum bicolor as a comparison to classic experimental design in an agricultural setting.

Biochemistry & Molecular Biology↗

1 × 1 km maps of abundances of eight enzyme functional classes for soil C, N, and P cycling across the CONUS

This dataset includes eight 1 × 1 km maps of the abundances of eight enzyme functional classes (EFC) for soil C, N, and P cycling across the CONUS. These mappings are predicted by the machine learning model trained using metagenomics and the corresponding environmental data. This item corresponds to our article: Fan, C., Song, Y., Mishra, U., Gautam, S., & Mayes, M. A. (2025). Harnessing the Power of Machine Learning and Omics to Identify Environmental Regulation on Microbial Functional Composition for Soil C, N, and P Cycling. Journal of Geophysical Research: Biogeosciences, 130(10).

1 × 1 km↗

Harnessing the Power of Machine Learning and Omics to Identify Environmental Regulation on Microbial Functional Composition for Soil C, N, and P Cycling

Microbial enzyme-mediated soil organic matter (SOM) decomposition regulates many key ecosystem functions, such as elemental cycling, soil carbon sequestration, and soil fertility. However, representing microbial processes in Earth system models (ESMs) remains challenging due to a limited understanding of the spatial patterns of diverse microbial functions responsible for soil carbon (C), nitrogen (N), and phosphorus (P) cycling as well as the underlying mechanisms regulating their relative abundances across various environments. We collected published metagenomics data across the continental US (CONUS) to identify hundreds of microbial genes involved in soil C, N, and P cycling and grouped them into eight enzyme functional classes (EFCs). Each EFC represented a group of gene-encoded potential enzymes that decompose similar soil compounds. By integrating the abundances of omics-informed EFCs with the corresponding environmental information, we trained a machine learning (ML) model to identify key edaphic, climate, and vegetation factors regulating the abundances of each EFC. Quantitative analysis of effects of these factors revealed that the spatial distribution of eight EFCs for soil C, N, and P cycling across CONUS reflected potential resource optimization strategies of microbial communities under nutrient limitation, preferential organic-mineral associations, and climatological stresses. This insight, together with the interpreted ML tool and the CONUS-level benchmark for EFCs abundances, paves the way for parameterizing environmental-regulated microbial functional dynamics in biogeochemical models.

machine learning↗

Altering translation allows E. coli to overcome G-quadruplex stabilizers

The data included in this Dryad submission was collected in order to understand how the model organism* Escherichia coli* overcomes stabilized G-quadruplexes. This work involved a multi-omics approach to studying how the G-quadruplex stabilizers NMM and Braco-19 impact growth, gene importance, and mRNA/proteomic abundance in G-quadruplex stabilizing conditions.

Bacteria↗

Expanding Biological Repository Data Available for Sharing and Knowledge Discovery

Biology has developed next-generation data science and alternative analytical approaches with methodologies which require principal investigator (PI) experimental assay data be re-used. This new approach involves mining multiple datasets at once from various hierarchical organizations of biological complexity, while concurrently evaluating how experimental factors affect endpoints of standard assays. The purpose of the NASA Ames Life Sciences Data Archive (ALSDA) is to collect, curate, and make findable, accessible, interoperable, and reusable (FAIR) all non-human space-relevant biological data. These data include mission metadata, subject metadata, assay metadata (parameters), raw and processed assay data, assay imagery, and subject-experienced telemetry (radiation, temperature, humidity, acoustics, vibrations). ALSDA has transformed to bring current biological repository data and all future collected data into this new scientific data mining reality. It has integrated into the ‘NASA Open Science’ group of projects to facilitate a suite of new tools and workflows to improve data accessibility and reusability by implementing data management plans, automating data submission agreements, and adopting the single-point-of-entry data submission portal, originally developed by NASA GeneLab. These systems required ALSDA to develop science assay configurations for the submission portal, capturing essential assay parameters according to established norms in each sub-field within biology. The submission portal expedites data collection by enhancing ease of PI data submission, providing a user interface and specificity for which data is to be submitted. ALSDA datasets are curated to maintain rich metadata, accuracy of datasets, data transparency, provenance, and additionally ensure data are machine-readable (e.g., R and Python languages). ALSDA integration with GeneLab and its analysis portals enable higher-order physiological-level datasets be mined in conjunction with -omics datasets. As ALSDA physiological-level datasets are published (micro-computed tomography, histology, intraocular pressure, hormonal assays, immunostaining, ultrasonography), the merging of hierarchical organizations of biological complexity from spaceflight will enable new knowledge discovery approaches.

Ryan T Scott↗

2024 NMDC Ambassador Training Materials [Slides]

The NMDC is a sustainable data discovery platform that promotes open science and shared-ownership across a broad and diverse community of researchers, funders, publishers, societies, and other collaborators. The NMDC aims to enable multi-omic microbiome research to accelerate scientific discovery. The NMDC is a Department of Energy funded program that is a collaboration between 3 National Laboratories: Lawrence Berkeley National Laboratory (LBNL), Los Alamos National Laboratory (LANL), and Pacific Northwest National Laboratory (PNNL).

54 ENVIRONMENTAL SCIENCES↗

Transcriptomics and Proteomics Discussion

This presentation will cover the the basic pipelines for transcriptomics and proteomics that the GeneLab Analysis Working Groups (AWGs) have so far determined to be optimal. Basic transcriptomic pipelines will first be presented from primary analysis to higher-order systems analysis. Examples of how the data has been analyzed will be presented. Proteomics pipelines will also be presented compiled from various AWG members. Discussion will be generated from the AWG members to reach a consensus for each omic type.

Transcriptomics↗

RWRtoolkit: multi-omic network analysis using random walks on multiplex networks in any species

Abstract We introduce RWRtoolkit, a multiplex generation, exploration, and statistical package built for R and command-line users. RWRtoolkit enables the efficient exploration of large and highly complex biological networks generated from custom experimental data and/or from publicly available datasets, and is species agnostic. A range of functions can be used to find topological distances between biological entities, determine relationships within sets of interest, search for topological context around sets of interest, and statistically evaluate the strength of relationships within and between sets. The command-line interface is designed for parallelization on high-performance cluster systems, which enables high-throughput analysis such as permutation testing. Several tools in the package have also been made available for use in reproducible workflows via the KBase web application.

Kainer, David (ORCID:0000000172714676)↗

BioNutrients-1: Development of an On-Demand Nutrient Production System for Long-Duration Missions

Future long-duration missions beyond low-Earth orbit will require advances in food technologies to address the documented problem of degradation of vitamins and nutrients in supplied foods stored long-term. To begin to address the issue of nutrient degradation, we are developing and flight-testing a platform biomanufacturing technology for in situ production of target nutrients. This technology is being tested over a five-year duration on the International Space Station (ISS). As part of the BioNutrients-1 project we have developed an on-demand system for the production of two carotenoids, β-carotene and zeaxanthin, by genetically engineering distinct strains of Saccharomyces cerevisiae, more commonly known as baker’s yeast. The on-orbit nutrient production packs contain a desiccated yeast strain and edible growth substrate. Once hydrated, the contents of the production packs are intended to grow and produce a desired amount of ready-to-consume nutrients. In this current version the production packs will not be consumed and future missions will require an inactivation of microorganisms before consumption. In addition to the on-orbit hydration of the production packs a series of valuable microorganisms are currently being stored in stasis packs on the ISS including probiotics organisms, bacterial strains used in yogurt production, and organisms with potential use for future biomanufacturing. Analysis of returned ISS stasis packs and ground controls will include multi-omics studies and provide insight into long-term survival of organisms stored in a space environment. Both stasis packs and hydrated production packs will be intermittently returned to Earth for analysis. Preliminary data from long-term storage studies of stasis packs stored on the ISS for 47 days versus their ground control counterparts have shown no significant difference in viability. Currently no production packs have been processed.

Yeast↗

Nitrogen starvation causes lipid remodeling in Rhodotorula toruloides

Abstract Background The oleaginous yeast Rhodotorula toruloides is a promising chassis organism for the biomanufacturing of value-added bioproducts. It can accumulate lipids at a high fraction of biomass. However, metabolic engineering efforts in this organism have progressed at a slower pace than those in more extensively studied yeasts. Few studies have investigated the lipid accumulation phenotype exhibited by R. toruloides under nitrogen limitation conditions. Consequently, there have been only a few studies exploiting the lipid metabolism for higher product titers. Results We performed a multi-omic investigation of the lipid accumulation phenotype under nitrogen limitation. Specifically, we performed comparative transcriptomic and lipidomic analysis of the oleaginous yeast under nitrogen-sufficient and nitrogen deficient conditions. Clustering analysis of transcriptomic data was used to identify the growth phase where nitrogen-deficient cultures diverged from the baseline conditions. Independently, lipidomic data was used to identify that lipid fractions shifted from mostly phospholipids to mostly storage lipids under the nitrogen-deficient phenotype. Through an integrative lens of transcriptomic and lipidomic analysis, we discovered that R. toruloides undergoes lipid remodeling during nitrogen limitation, wherein the pool of phospholipids gets remodeled to mostly storage lipids. We identify specific mRNAs and pathways that are strongly correlated with an increase in lipid levels, thus identifying putative targets for engineering greater lipid accumulation in R. toruloides . One surprising pathway identified was related to inositol phosphate metabolism, suggesting further inquiry into its role in lipid accumulation. Conclusions Integrative analysis identified the specific biosynthetic pathways that are differentially regulated during lipid remodeling. This insight into the mechanisms of lipid accumulation can lead to the success of future metabolic engineering strategies for overproduction of oleochemicals.

59 BASIC BIOLOGICAL SCIENCES↗

Space Flown Rodent Liver RNA Sequencing Data for Machine Learning in Space Biology Research

High-throughput nucleic acid sequencing (DNA-seq, RNA-seq) has become widespread in biomedical research due to the growing availability and affordability of these assays. Data analysis has been accelerated in recent years by the adoption of artificial intelligence (AI) and machine learning (ML) techniques by biomedical researchers. In space biology research, RNAseq datasets from space-flown experimental samples are critical for characterizing the gene expression aberrations associated with exposure to spaceflight stressors. However, space biological experiments tend to be very low sample size, so identifying proper AI/ML algorithms for sequencing data analysis is an ongoing challenge since these algorithms typically require large sample size. The NASA Science Mission Directorate (SMD) has started the “Benchmark Initiative for AI/ML”, focused on creating datasets meant for three main applications: 1) scientific benchmarking, which finds the best algorithm for a specific problem; 2) application benchmarking, which measures algorithm performance against a set of parameters; and 3) system benchmarking, which evaluates performance of hardware and software architecture. These scientific benchmarks consist of an AI-ready dataset and a reference implementation on a specific scientific question. In this work, we focused on generating standardized datasets to allow the scientific community to benchmark AI/ML algorithms in the domain of space biology. We present here a standardized, AI-ready, publicly available benchmark dataset for space biology RNA-seq data as a collaboration between the NASA AI4LS (Artificial Intelligence for Life Sciences) working group. and NASA’s SMD. This dataset consists of space-flown and ground control mouse liver found in the NASA GeneLab omics database. However, to amplify the small sample number (n=112 samples) for ML purposes, we employ Gaussian noise and a generative adversarial network to extend this dataset to 6,000 synthetic samples, matching the original gene expression characteristics.

James Casaletto↗

Untargeted, tandem mass spectrometry (LC/MS-MS) metaproteomes from soil samples in control and warming plots in Blodgett Forest, CA (2014-2021)

The pathways of carbon transport and loss through and from soils—soil organic matter (SOM) depolymerization to dissolved organic carbon and mineralization to carbon dioxide (CO2)—are fundamentally driven by microbial activity, which is strongly regulated by environmental conditions. As part of Lawrence Berkeley National Laboratory (LBNL) Terrestrial Ecosystem Science (TES) Belowground Biogeochemistry Science Focus Area (SFA), we have established a novel whole-soil long-term warming experiment at the University of California (UC) Blodgett Forest Research Station (Sierra Nevada) in 2014, where we study the role of biogeochemical, microbial and geochemical process interactions in SOM decomposition and stabilization. This package contains soil metaproteomics data in the context of site specific metagenomes from soil depth profiles in three paired control and warming plots from a temperate mixed forest in Northern California. Each paired plot had been subjected to experimental warming since June 2014 to simulate a predicted climate change scenario for northern California. These metaproteomes were collected in 2018 after 4.5 years of warming from five depth intervals (0-10 cm, 10-30 cm, 30-45 cm, 45-60 cm, 60-80 cm). For protein identification, the collected spectra were searched following a target-decoy search strategy against a database of metagenome predicted proteins (covering 96 samples from 2014 to 2021) representing the complete sequence diversity at the site. Data was searched with mass spectrometry database search tool (MS-GF+) using Pacific Northwest National Laboratory (PNNL)'s Data Management System (DMS) Processing pipeline. The metagenomes are published as part of another data package. Raw metaproteomic data and the data products from MS-GF+ are deposited in the Mass Spectrometry Interactive Virtual Environment (MassIVE) database under accession no. MSV000097826. Here we present a dataset that includes spectral counts for the detected proteins across samples (EMSL50964_BrodieAllMAGs_Globals_SC.txt), the sequences of the detected proteins, and sample metadata file that contains site information for the soil metaproteome samples.

Belowground Biogeochemistry Science Focus Area↗

Omics-Based Comparison of Fungal Virulence Genes, Biosynthetic Gene Clusters, and Small Molecules in Penicillium expansum and Penicillium chrysogenum

Penicillium expansum is a ubiquitous pathogenic fungus that causes blue mold decay of apple fruit postharvest, and another member of the genus, Penicillium chrysogenum, is a well-studied saprophyte valued for antibiotic and small molecule production. While these two fungi have been investigated individually, a recent discovery revealed that P. chrysogenum can block P. expansum-mediated decay of apple fruit. To shed light on this observation, we conducted a comparative genomic, transcriptomic, and metabolomic study of two P. chrysogenum (404 and 413) and two P. expansum (Pe21 and R19) isolates. Global transcriptional and metabolomic outputs were disparate between the species, nearly identical for P. chrysogenum isolates, and different between P. expansum isolates. Further, the two P. chrysogenum genomes revealed secondary metabolite gene clusters that varied widely from P. expansum. This included the absence of an intact patulin gene cluster in P. chrysogenum, which corroborates the metabolomic data regarding its inability to produce patulin. Additionally, a core subset of P. expansum virulence gene homologues were identified in P. chrysogenum and were similarly transcriptionally regulated in vitro. Molecules with varying biological activities, and phytohormone-like compounds were detected for the first time in P. expansum while antibiotics like penicillin G and other biologically active molecules were discovered in P. chrysogenum culture supernatants. Our findings provide a solid omics-based foundation of small molecule production in these two fungal species with implications in postharvest context and expand the current knowledge of the Penicillium-derived chemical repertoire for broader fundamental and practical applications.

Bartholomew, Holly P. (ORCID:0000000292726399)↗

Transcriptomic Analysis of Irradiated Mouse Retina Following Readaptation

Rodent models are used as analogs for studying the effects of spaceflight. NASA GeneLab provides access to relevant omics datasets generated from spaceflight and ground-based experiments allowing for additional retrospective analysis. In this study, we used GeneLab’s GLDS-203, a dataset generated by researchers at Loma Linda University to study the impact of prolonged unloading and/or low-dose radiation on mouse retina. We analyzed transcriptomics data from retina of mice irradiated with gamma-rays for 21 days followed by 7 days, 1 month, or 4 months of readaptation. We obtained raw gene counts from GeneLab and performed differential gene expression analysis after data normalization. For each of the three timepoints, we performed differential expression analysis to compare transcriptional profiles for retina from irradiated vs. non-irradiated (controls) mice, all exposed to gravity. We observed the highest number of differentially expressed genes at 7 days, followed by 1 month and 4 months. Enrichment analysis showed top pathways (adjusted p-value < 0.05) were related to transport along microtubule and photoreceptor cell development in the 7-day readaptation group. Fewer significantly enriched pathways were observed for the 1-month group and included mRNA metabolic processes and neuron differentiation. No significantly enriched pathways were found in the 4-month group. The Gene Ontology biological processes common between the 7 days and 1-month groups include visual perception, synapse organization, and perception of light stimulus. This analysis is part of a larger effort to characterize the molecular mechanisms involved in retinal readaptation following radiation exposure. Future analyses will include other related retina datasets in GeneLab repository to assess whether gene expression patterns are consistent across different study cohorts.

Prachi Kothiyal↗

2024 IUFRO Tree Biotechnology Conference (Aug 4-8, 2024)

The 2024 IUFRO Tree Biotechnology Conference is the biennial meeting on genomics, molecular biology, and biotechnology of forest trees, associated with the IUFRO Working Party 2.04.06. This year's meeting was held in Annapolis, MD, USA from August 4th to 8th and was hosted by Yiping Qi (University of Maryland), Edward Eisenstein (University of Maryland), Gary Coleman (University of Maryland), and Heather Coleman (Syracuse University). The conference covered seven topics over the course of five days: 1) Biological and ecological insights from OMICS, 2) Advancing technologies for targeted trait manipulation and acceptability to diverse tree species, 3) Genes, development, and physiology, 4) Translating genomics and biotechnology to practice, 5) Trees in a changing world, 6) Genetic and phenotypic diversity for breeding and genomic selection, and 7) Biotechnology for biomaterials and bioeconomy. In addition to the sessions, there were two plenary sessions, provided by John Ralph (University of Wisconsin) and Tanja Pyrhäjärvi (University of Helsinki). The meeting celebrated the second awardees of the newly created IUFRO WG 2.04.06 Award: Excellence in Forest Molecular Biology and Genomics, which was presented to Chung-Jui (C.J.) Tsai (University of Georgia). Greg Goralogia (Oregon State University) was the recipient of the associated Early Career Award. The scientific presentations at the conference highlighted cutting-edge advancements in many facets of forest biotechnology research, including applications of genomic selection in forest genetics and breeding, the use of genetic editing, tree physiology, stress response, molecular breeding, wood development, "omics" technologies, and the social and economic impacts of genetically modified (GM) trees. Scientific take homes from the meeting include the power of NMR to dissect the composition of lignin, the genomic diversity of forest trees that has enormous potential for tree improvement and the integration of systems biology with climate and geographical data. The conference attracted a mix of students (25), postdoctoral fellows (32), and scientists from academia (66) and industry (18). In all, the conference was attended by 141 registered participants, representing 20 countries that participated in 23 invited lectures (including 6 'early-career' keynotes), 27 voluntary talks and 61 poster presentations. Support for the conference was drawn from a wide variety of Academia, Industry, and Government sources, and included financial support from several tree improvement companies. Overall, the conference was a great success, providing an exceptional mix of science and social activities in a relaxed and collegial atmosphere. More information about the meeting can be found at treebiotech.org. The next meeting will be held in Stellenbosch, South Africa, in 2026, hosted jointly by Zander Myburg, Dave Drew (University of Stellenbosch,) and Sanushka Naidoo (University of Pretoria, FABI).

59 BASIC BIOLOGICAL SCIENCES↗

Research from the NASA Twins Study and Omics in Support of Mars Missions

The NASA Twins Study, NASA's first foray into integrated omic studies in humans, illustrates how an integrated omics approach can be brought to bear on the challenges to human health and performance on a Mars mission. The NASA Twins Study involves US Astronaut Scott Kelly and his identical twin brother, Mark Kelly, a retired US Astronaut. No other opportunity to study a twin pair for a prolonged period with one subject in space and one on the ground is available for the foreseeable future. A team of 10 principal investigators are conducting the Twins Study, examining a very broad range of biological functions including the genome, epigenome, transcriptome, proteome, metabolome, gut microbiome, immunological response to vaccinations, indicators of atherosclerosis, physiological fluid shifts, and cognition. A novel aspect of the study is the integrated study of molecular, physiological, cognitive, and microbiological properties. Major sample and data collection from both subjects for this study began approximately six months before Scott Kelly's one year mission on the ISS, continue while Scott Kelly is in flight and will conclude approximately six months after his return to Earth. Mark Kelly will remain on Earth during this study, in a lifestyle unconstrained by this study, thereby providing a measure of normal variation in the properties being studied. An overview of initial results and the future plans will be described as well as the technological and ethical issues raised for spaceflight studies involving omics.

Kundrot, C.↗

Growth rate as a link between microbial diversity and soil biogeochemistry

The growth rate of a microorganism is a simple yet profound way to quantify its impact on the world. The absolute growth rate of a microbial population reflects rates of resource assimilation, biomass production, and element transformation, some of the many ways that organisms affect Earth’s ecosystems and climate. Microbial fitness in the environment depends on the ability to reproduce quickly when conditions are favorable and adopt a survival physiology when conditions worsen, which cells coordinate by adjusting their relative growth rate. At the population level, relative growth rate is a sensitive metric of fitness, linking survival and reproduction to the ecology and evolution of populations. Techniques combining ‘omics and stable isotope probing enable sensitive measurements of growth rates of microbial assemblages and individual taxa in soil. Microbial ecologists can explore how the growth rates of taxa with known traits and evolutionary histories respond to changes in resource availability, environmental conditions, and interactions with other organisms. We anticipate that quantitative and scalable data on the growth rates of soil microorganisms, coupled with measurements of biogeochemical fluxes, will allow scientists to test and refine ecological theory and advance process-based models of carbon flux, nutrient uptake, and ecosystem productivity. Finally, measurements of in situ microbial growth rates provide insights into the ecology of populations and can be used to quantitatively link microbial diversity to soil biogeochemistry.

54 ENVIRONMENTAL SCIENCES↗