Search NASA⌕ Search

SEARCH · Search NASA

Results for “structural proteomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

The microbiologist's guide to metaproteomics

Metaproteomics is an emerging approach for studying microbiomes, offering the ability to characterize proteins that underpin microbial functionality within diverse ecosystems. As the primary catalytic and structural components of microbiomes, proteins provide unique insights into the active processes and ecological roles of microbial communities. By integrating metaproteomics with other omics disciplines, researchers can gain a comprehensive understanding of microbial ecology, interactions, and functional dynamics. This review, developed by the Metaproteomics Initiative (www.metaproteomics.org), serves as a practical guide for both microbiome and proteomics researchers, presenting key principles, state-of-the-art methodologies, and analytical workflows essential to metaproteomics. Topics covered include experimental design, sample preparation, mass spectrometry techniques, data analysis strategies, and statistical approaches.

bioinformatics↗

Algorithms and file structures to extend and enhance liquid chromatography and ion mobility mass spectrometry workflows (CRADA Final Report)

The purpose of this project was to continue supporting customizations of algorithms and raw data file structures to enhance software workflows for liquid chromatography (LC), mass spectrometry (MS) and ion mobility mass spectrometry (IM-MS)-based protein and metabolite characterization. PNNL worked with Agilent to design, implement, evaluate, and demonstrate new algorithms and integrated them as functionalities into the PNNL-PreProcessor software. The project augmented PNNL’s capabilities to analyze complex proteomics and metabolomics samples. These capabilities are directly beneficial to DOE and PNNL efforts to characterize and analyze these compounds in microbial and plant communities. The project assisted Agilent in further developing improved instrument-software solutions combining liquid chromatography and ion mobility with mass spectrometry for widespread applications in life sciences and other fields.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Focused Metabolite Profiling for Dissecting Cellular and Molecular Processes of Living Organisms in Space Environments

Regulatory control in biological systems is exerted at all levels within the central dogma of biology. Metabolites are the end products of all cellular regulatory processes and reflect the ultimate outcome of potential changes suggested by genomics and proteomics caused by an environmental stimulus or genetic modification. Following on the heels of genomics, transcriptomics, and proteomics, metabolomics has become an inevitable part of complete-system biology because none of the lower "-omics" alone provide direct information about how changes in mRNA or protein are coupled to changes in biological function. The challenges are much greater than those encountered in genomics because of the greater number of metabolites and the greater diversity of their chemical structures and properties. To meet these challenges, much developmental work is needed, including (1) methodologies for unbiased extraction of metabolites and subsequent quantification, (2) algorithms for systematic identification of metabolites, (3) expertise and competency in handling a large amount of information (data set), and (4) integration of metabolomics with other "omics" and data mining (implication of the information). This article reviews the project accomplishments.

Source record↗

Hydrazinoacetic acid is a biosynthetic precursor of the bacterially produced nitramine, N -nitroglycine

Nitramines [R(R′)N–NO 2 ; R,R′=H or alkyl] are valuable synthetic products, but knowledge of the biosynthetic processes that generate these compounds is limited. This work sought to elucidate the biosynthesis of a nitramine natural product, N-nitroglycine (NNG) by Streptomyces noursei . Stable isotope studies showed that S. noursei cells supplemented with L-(ε- 15 N)lysine, ( 15 N)glycine, or ( 13 C)hydrazinoacetic acid (HAA) incorporated 67%, 88%, and 67% of the isotope label into NNG, respectively, indicating that these compounds are biosynthetic precursors of NNG. Liquid chromatography coupled tandem mass spectrometry (LC-MS/MS) of 15 N-Lys-labeled NNG confirmed that the nitro nitrogen of NNG originates from Lys. Bioinformatics analysis of the S. noursei genome showed evidence for a biosynthetic gene cluster (BGC) that contained machinery for HAA biosynthesis ( nngKLM ), consistent with the results of the isotope labeling. In vitro reconstitution of the gene products produced HAA. The borders of this BGC were defined by cross-referencing the predicted BGC with previously published differential proteomics data. Furthermore, we show that azaserine is produced alongside NNG in S. noursei cultures, linking the two biosynthetic pathways via a proposed nitrosamine biosynthetic intermediate. Finally, the oxygen balance for NNG is −20.2% for the formation of carbon dioxide (CO 2 ), which is comparable to that of hexahydro-1,3,5- trinitro-1,3,5-triazene (common name: RDX; −21.6%). Crystal structure data of NNG indicate that the unit crystalizes as a pure material, not a hydrate, suggesting a favorable energetic crystallization phase. The combined results suggest a route that, with further development, could lead to sustainable production of energetic nitramines via synthetic biology or biocatalytic approaches.

biosynthesis↗

The Thiamin Pyrophosphate-Motif

Using databases the authors have identified a common thiamin pyrophosphate (TPP)-motif in the family of functionally diverse TPP-dependent enzymes. This common motif consists of multimeric organization of subunits, two catalytic centers, common amino acid sequence, and specific contacts to provide a flip-flop, or alternate site, mechanism of action. Each catalytic center [PP:PYR] is formed at the interface of the PP-domain binding the magnesium ion, pyrophosphate and aminopyrimidine ring of TPP, and the PYR-domain binding the aminopyrimidine ring of that cofactor. A pair of these catalytic centers constitutes the catalytic core [PP:PYR]* within these enzymes. Analysis of the structural elements of this catalytic core reveals novel definition of the common amino acid sequences, which are GX@&(G)@XXGQ, and GDGX25-30 within the PP- domain, and the E&(G)@XXG@ within the PYR-domain, where Q, corresponds to a hydrophobic amino acid. This TPP-motif provides a novel tool for annotation of TPP-dependent enzymes useful in advancing functional proteomics.

Dominiak, Paulina M.↗

Data Summarization and Inference at Scale

This is the final report for the DOE ASCR grant SC-0022260, Data Summarization and Inference at Scale, PI: Alex Pothen, Purdue University. The goal of the project was to solve data-intensive and compute-intensive problems in the physical sciences, engineering, information science, data science, etc. by designing and implementing new algorithms that could work with a subset of the data. The four subgoals were: (a) The solution of problems where the data is too large to be stored in the memory of a computer. In this streaming model of computation, the data arrives as a stream of elements to the computer, each element is processed as it arrives, and a decision is made to discard the data or to store it; only a small subset of the data proportional to the size of the output solution is stored, and when all the data has been streamed, a solution to the problem is computed from the stored subset. (b) The use of machine learning methods to compute solutions to data-intensive problems. The use of GPUs is critical to obtain high performance on machine learning tasks, but their memory sizes are smaller relative to that of CPUs. For large-scale problems, the data is sampled many times, and small samples are used with repetition, for robustness, to compute solutions to inference tasks. This sampling reduces the memory required to solve the problem, but attention is needed to avoid slow convergence to the solutions, and reduced accuracy of inference. We propose submodular optimization, Large Language Models, and physics-informed neural networks to enable GPU computations here. (c) Modeling and visualization of high-dimensional data using interpretable features. Clinical proteomic data sets from immunology for the detection of cancer and other diseases are temporal and high-dimensional, and algorithms for visualizing these data sets using clinically interpretable features are lacking. We propose methods that compute distances based on the optimal transportation problem and graph edit distances to address this problem. We also propose the use of optimal transport-based distances, spatial statistics, and network structure to classify image data sets, We apply these algorithms to electron micrographs of the peripheral nervous system in the digestive tract. (d) The design of data-intensive algorithms on emerging architectures, specifically, noisy, intermediate-scale quantum (NISQ) devices. Quantum computers offer the possibility of exploring large solution spaces due to the principle of superposition, but current quantum computers are limited by few qubits, short coherence times due to noise, poor interconections among the qubits, etc. We propose the use of the divide and conquer paradigm to solve large-scale problems, wherein collections of small subproblems are solved on the quantum devices, and the solutions to the subproblems are integrated into a solution for the original problem on a classical computer.

97 MATHEMATICS AND COMPUTING↗

The Thiamin Pyrophosphate-Motif

Using databases the authors have identified a common thiamin pyrophosphate (TPP)-motif in the family of functionally diverse TPP-dependent enzymes. This common motif consists of multimeric organization of subunits and two catalytic centers. Each catalytic center (PP:PYR) is formed at the interface of the PP-domain binding the magnesium ion, pyrophosphate and amhopyrimidine ring of TPP, and the PYR-domain binding the aminopyrimidine ring of that cofactor. A pair of these catalytic centers constitutes the catalytic core (PP:PYR)(sub 2) within these enzymes. Analysis of the structural elements of this catalytic core reveals novel definition of the common amino acid sequences, which are GXPhiX(sub 4)(G)PhiXXGQ and GDGX(sub 25-30)NN in the PP-domain, and the EX(sub 4)(G)PhiXXGPhi in the PYR-domain, where Phi corresponds to a hydrophobic amino acid. This TPP-motif provides a novel tool for annotation of TPP-dependent enzymes useful in advancing functional proteomics.

Dominiak, P.↗

EMC3 regulates trafficking and pulmonary toxicity of the SFTPC I73T mutation associated with interstitial lung disease

The most common mutation in surfactant protein C gene (SFTPC), SFTPC I73T , causes interstitial lung disease with few therapeutic options. We previously demonstrated that EMC3, an important component of the multiprotein endoplasmic reticulum membrane complex (EMC), is required for surfactant homeostasis in alveolar type 2 epithelial (AT2) cells at birth. In the present study, we investigated the role of EMC3 in the control of SFTPC I73T metabolism and its associated alveolar dysfunction. Using a knock-in mouse model phenocopying the I73T mutation, we demonstrated that conditional deletion of Emc3 in AT2 cells rescued alveolar remodeling/simplification defects in neonatal and adult mice. Proteomic analysis revealed that Emc3 depletion reversed the disruption of vesicle trafficking pathways and rescued the mitochondrial dysfunction associated with I73T mutation. Affinity mass spectrometry analysis identified potential EMC3 interacting proteins in lung AT2 cells, including Valosin Containing Protein (VCP) and its interactors. Treatment of Sftpc I73T knock-in mice and SFTPC I73T expressing iAT2 cells derived from SFTPC I73T patient-specific iPSCs with the specific VCP inhibitor CB5083 restored alveolar structure and SFTPC I73T trafficking respectively. Taken together, the present work identifies the EMC complex and VCP in the metabolism of the disease-associated SFTPC I73T mutant, providing novel therapeutical targets for SFTPC I73T -associated interstitial lung disease.

60 APPLIED LIFE SCIENCES↗

Graph Identification of Proteins in Tomograms (GRIP-Tomo) 2.0: Topologically aware classification for proteins

Cryo-electron tomography (cryo-ET) enables structural characterization of biomolecules under near-native conditions. Existing approaches for interpreting the resulting three-dimensional volumes are computationally expensive and have difficulty interpreting density associated with small proteins/complexes. To explore alternate approaches for identifying proteins in cryo-ET data we pursued a Graph Network and topologically invariant approach. Here, we report on a fast algorithm that classifies particles by searching for nuances of evolutionarily conversed motifs and the geometrical characteristics of protein structure. GRIP-Tomo 2.0 is a machine-learning pipeline that extracts interpretable topological features of protein structures within noisy experimental backgrounds. Compared to version 1.0, the new pipeline includes three upgrades that significantly improve performance including synthetic tomogram generation simulating realistic noise, graph-based persistent feature extraction as protein fingerprints, and high-performance computing acceleration. GRIP-Tomo 2.0 achieves over 90% accuracy in classifying between proteins and noise using both real and synthetic datasets which represents a foundational step toward advancing cryo-ET workflows and empowering automated visual proteomics.

Li, Chengxuan↗

Blocking C-terminal processing of KRAS4b via a direct covalent attack on the CaaX-box cysteine

RAS is the most frequently mutated oncogene in cancer. RAS proteins show high sequence similarities in their G-domains but are significantly different in their C-terminal hypervariable regions (HVR). These regions interact with the cell membrane via lipid anchors that result from posttranslational modifications (PTM) of cysteine residues. KRAS4b is unique as it has only one cysteine that undergoes PTM, C185. Small molecule covalent modification of C185 would block any form of prenylation and subsequently inhibit attachment of KRAS4b to the cell membrane, blocking its biological activity. We translated this concept to the discovery and development of disulfide tethering screen hits into irreversible covalent modifiers of C185. These compounds inhibited proliferation of KRAS4b-driven mouse embryonic fibroblasts, but not cells driven by N-myristoylated KRAS4b that harbor a C185S mutation and are not dependent on C185 prenylation. Top–down proteomics was used to confirm target engagement in cells. These compounds bind in a pocket formed when the HVR folds back between helix 3 and 4 in the G-domain (HVR-α3-α4). This interaction can happen in the absence of small molecules as predicted by molecular dynamics simulations and is stabilized in the presence of C185 binders as confirmed by small-angle X-ray scattering and solution NMR. NOESY-HSQC, an NMR approach that measures internuclear distances of 6 Å or less, and structure analysis identified the critical residues and interactions that define the HVR-α3-α4 pocket. Further development of compounds that bind to this pocket could be the basis of a new approach to targeting KRAS cancers.

C185↗

Altering translation allows E. coli to overcome G-quadruplex stabilizers

G-quadruplex (G4) structures can form in guanine-rich DNA or RNA and have been found to modulate cellular processes, including replication, transcription, and translation. Many studies on the cellular roles of G4s have focused on eukaryotic systems, with far fewer probing bacterial G4s. Using a chemical-genetic approach, we identified genes in Escherichia coli that are important for growth in G4-stabilizing conditions. Reducing levels of translation elongation factor Tu or slowing translation initiation or elongation with kasugamycin, chloramphenicol, or spectinomycin suppress the effects of G4-stabilizing compounds. In contrast, reducing the expression of specific translation termination or ribosome recycling proteins is detrimental to growth in G4-stabilizing conditions. Proteomic and transcriptomic analyses reveal decreased protein and transcript levels, respectively, for ribosome assembly factors and proteins associated with translation in the presence of G4 stabilizer. Our results support a model in which reducing the rate of translation by altering translation initiation, translation elongation, or ribosome assembly can compensate for G4-related stress in E. coli.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative proteomics of a versatile, marine, iron-oxidizing chemolithoautotroph

This study conducted a comparative proteomic analysis to identify potential genetic markers for the biological function of chemolithoautotrophic iron oxidation in the marine bacterium Ghiorsea bivora. To date, this is the only characterized species in the class Zetaproteobacteria that is not an obligate iron-oxidizer, providing a unique opportunity to investigate differential protein expression to identify key genes involved in iron-oxidation at circumneutral pH. Over 1000 proteins were identified under both iron- and hydrogen-oxidizing conditions, with differentially expressed proteins found in both treatments. Notably, a gene cluster upregulated during iron oxidation was identified. This cluster contains genes encoding for cytochromes that share sequence similarity with the known iron-oxidase, Cyc2. Interestingly, these cytochromes, conserved in both Bacteria and Archaea, do not exhibit the typical β-barrel structure of Cyc2. This cluster potentially encodes a biological nanowire-like transmembrane complex containing multiple redox proteins spanning the inner membrane, periplasm, outer membrane, and extracellular space. The upregulation of key genes associated with this complex during iron-oxidizing conditions was confirmed by quantitative reverse transcription-PCR. These findings were further supported by electromicrobiological methods, which demonstrated negative current production by G. bivora in a three-electrode system poised at a cathodic potential. This research provides significant insights into the biological function of chemolithoautotrophic iron oxidation.

59 BASIC BIOLOGICAL SCIENCES↗

Nitrogen limitation causes a seismic shift in redox state and phosphorylation of proteins implicated in carbon flux and lipidome remodeling in Rhodotorula toruloides

Background: Oleaginous yeast are prodigious producers of oleochemicals, offering alternative and secure sources for applications in foodstuff, skincare, biofuels, and bioplastics. Nitrogen starvation is the primary strategy used to induce oil accumulation in oleaginous yeast as part of a global stress response. While research has demonstrated that post-translational modifications (PTMs), including phosphorylation and protein cysteine thiol oxidation (redox PTMs), are involved in signaling pathways that regulate stress responses in metazoa and algae, their role in oleaginous yeast remain understudied and unexplored. Results: Towards linking the yeast oleaginous phenotype to protein function, we integrated lipidomics, redox proteomics, and phosphoproteomics to investigate Rhodotorula toruloides under nitrogen-rich and starved conditions over time. Our lipidomics results unearthed interactions involving sphingolipids and cardiolipins with ER stress and mitophagy. Our redox and phosphoproteomics data highlighted the roles of the AMPK, TOR, and calcium signaling pathways in regulation of lipogenesis, autophagy, and oxidative stress response. As a first, we also demonstrated that lipogenic enzymes including fatty acid synthase are modified as a consequence of shifts in cellular redox states due to nutrient availability. Conclusions: We conclude that lipid accumulation is largely a consequence of carbon rerouting and autophagy governed by changes to PTMs, and not increases in the abundance of enzymes involved in central carbon metabolism and fatty acid biosynthesis. Our systems-level approach sets the stage for acquiring multidimensional data sets for protein structural modeling and predicting the functional relevance of PTMs using Artificial Intelligence/Machine Learning (AI/ML). Coupled to those bioinformatics approaches, the putative PTM switches that we delineate will enable advanced metabolic engineering strategies to decouple lipid accumulation from nitrogen limitation.

Lipid Signalling↗

Fast and deep phosphoproteome analysis with the Orbitrap Astral mass spectrometer

Owing to its roles in cellular signal transduction, protein phosphorylation plays critical roles in myriad cell processes. That said, detecting and quantifying protein phosphorylation has remained a challenge. We describe the use of a novel mass spectrometer (Orbitrap Astral) coupled with data-independent acquisition (DIA) to achieve rapid and deep analysis of human and mouse phosphoproteomes. With this method, we map approximately 30,000 unique human phosphorylation sites within a half-hour of data collection. The technology is benchmarked to other state-of-the-art MS platforms using both synthetic peptide standards and with EGF-stimulated HeLa cells. We apply this approach to generate a phosphoproteome multi-tissue atlas of the mouse. Altogether, we detect 81,120 unique phosphorylation sites within 12 hours of measurement. With this unique dataset, we examine the sequence, structural, and kinase specificity context of protein phosphorylation. Finally, we highlight the discovery potential of this resource with multiple examples of phosphorylation events relevant to mitochondrial and brain biology.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The Elements of Life, Photosynthesis and Genomics

I am a Professor of Biochemistry, Biophysics and Structural Biology and Plant and Microbial Biology at the University of California in Berkeley. I was born and raised in India, emigrated to the United States to attend university, earning a B.S. in Molecular Biology and a Ph.D. in Biochemistry at the University of Wisconsin in Madison. Following post-doctoral studies with Lawrence Bogorad at Harvard University where I became interested in genetic control of trace element quotas, I joined the department of Chemistry and Biochemistry at UCLA. One of the first to appreciate essential trace metals as potential regulators of gene expression, I articulated the details of the nutritional Cu regulon in Chlamydomonas. In parallel, I used genetic approaches to discover the genes governing missing steps in tetrapyrrole metabolism, including the attachment of heme to apocytochromes in the thylakoid lumen and the factors catalyzing the formation of ring V in chlorophyll. After biochemistry and classical genetics, I embraced genomics, taking a leadership role on the Joint Genome Institute’s efforts on the Chlamydomonas genome and more recently, contributing to high quality assemblies of several genomes in the green algal radiation, and large transcriptomic and proteomic datasets — focusing on the diel metabolic cycle in synchronized cultures and acclimation to key environmental and nutritional stressors — that are well-used and appreciated by the community. Finally, a new venture in Berkeley is the promotion of Auxenochlorella protothecoides as the true “green yeast” and as a platform for engineering algae to produce useful bioproducts.

59 BASIC BIOLOGICAL SCIENCES↗

Energy metric prediction for double insertion mutants via the RoseNet deep learning framework

Studying the structural and functional implications of protein mutations is an important task in computational biology and bioinformatics. We leverage our previously proposed RoseNet neural network architecture to predict energy metrics of proteins with double amino acid insertions or deletions (InDels). We train models on previously generated benchmark datasets containing the exhaustive double InDel mutations for three proteins, as well as an additional three proteins for which ∼145k random mutants, each with two InDels, have been generated. We expand on our previous work by evaluating three additional proteins and analyzing domain features that impact the prediction capabilities of RoseNet. These features include InDels into secondary structures and the solvent accessible surface area (SASA) scores of the residues. We uncover further evidence to support that RoseNet has a higher proficiency of generalizing to unseen residue combinations than unseen insertion positions. We also observe that RoseNet produces higher-quality predictions when inserting into a β-sheet over an α-helix. Additionally, when the insertions fall in an area of high SASA, RoseNet often displays better performance than inserting into areas of low SASA.

59 BASIC BIOLOGICAL SCIENCES↗

The NASA Open Science Data Repository: Biomedical Fair Data, Analysis Tools, User Communities, Publications, and Discoveries for Deep Space Missions

Increased biomedical risks and challenges associated with deep space missions require new knowledge discovery, new health countermeasures, and development of novel ecosystems, life support, crop production, and biomedical support capabilities. To meet NASA’s Moon to Mars strategic program goals for Human and Biological Sciences, findable, accessible, interoperable, reusable (FAIR), and maximally open-access data is going to be required to enable humanity to thrive in deep space. Indeed, this cornerstone perspective on FAIR and maximally open access data was also recommended in the recent 2023-2032 Decadal Survey from the National Academies of Sciences, Engineering, and Medicine. The NASA Open Science Data Repository (OSDR) is a maximally open access and FAIR database, and meets various scientific, technical, and operational spaceflight needs. It offers public users and submitters the ability to upload, download, search, share, analyze, and visualize data across ‘omics, physiological, phenotypic, behavioral, bioimaging, video, and environmental monitoring telemetry datasets. OSDR includes NASA GeneLab, NASA Ames Life Sciences Data Archive, and the NASA Biological Institutional Scientific Collection. OSDR has >455 studies with datasets from model organisms and non-NASA human astronauts. There are ~12 datasets from the Inspiration 4 (I4) mission, spanning metagenomics, comprehensive metabolic panels, clonal hematopoiesis, spatial transcriptomics, proteomics, and cytokine panels. In the interest of data privacy, two I4 datasets have raw FASTQ and FASTA files relating to the epitranscriptome, and a new request feature is live in OSDR (with a backend review process established) which was developed based on industry norms. OSDR also recently began a collaboration with the European Space Agency (ESA) to scientifically curate and make available >200 terabytes of human and model organism space-relevant data. The OSDR submission portal is designed to ingest and curate ~25 ‘omics assay data types, and ~50 physiological-phenotypic-imaging assay data types, spanning ultrasonography, micro-computed tomography, histology, morphometric photography, rebound tonometry, gait analysis, optical coherence tomography, novel object recognition, flow cytometry, and immunohistochemistry. A suite of analysis tools are available for OSDR users including: 1) an Environmental Data Application to compare radiation, CO2, relative humidity, temperature, and other telemetry across missions and subjects, 2) the RadLab database, a collaboration between NASA, ESA, the German and Italian Space Agencies, and the Bulgarian Academy of Sciences, which compiles radiation measurements relevant to human spaceflight and provides tools for accessing and manipulating the data, and 3) a Multi-study visualization tool which enables users to look across and combine GeneLab’s omics datasets across different experiments and missions. There are ~600 volunteer OSDR Analysis Working Group (AWG) members who: 1) provide feedback on scientific standards for reuse (subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability), and 2) collaborate to mine-reuse OSDR data conducting scientific analysis. OSDR has enabled 60 publications as of September 2023, many directly from AWG collaborations most notably the Cell Press package in 2020. Lastly, there are at least 15 articles which mine OSDR data part of a package of ~50 articles across Nature Portfolio with research stemming from I4, the Japan Aerospace Exploration Agency, NASA Space Biology, and the NASA Human Research Program.

space biology↗

NASA Open Science Data Repository: Biomedical FAIR Data, Analysis Tools, User Communities, and Discoveries for Deep Space Missions

Increased biomedical risks and challenges associated with deep space missions require new knowledge discovery, new health countermeasures, and development of novel ecosystems, life support, crop production, and biomedical support capabilities. To meet NASA’s Moon to Mars strategic program goals for Human and Biological Sciences, findable, accessible, interoperable, reusable (FAIR), and maximally open-access data is going to be required to enable humanity to thrive in deep space. Indeed, this cornerstone perspective on FAIR and maximally open access data was also recommended in the recent 2023-2032 Decadal Survey from the National Academies of Sciences, Engineering, and Medicine. The NASA Open Science Data Repository (OSDR) is a maximally open access and FAIR database, and meets various scientific, technical, and operational spaceflight needs. It offers public users and submitters the ability to upload, download, search, share, analyze, and visualize data across ‘omics, physiological, phenotypic, behavioral, bioimaging, video, and environmental monitoring telemetry datasets. OSDR includes NASA GeneLab, NASA Ames Life Sciences Data Archive, and the NASA Biological Institutional Scientific Collection. OSDR has >455 studies with datasets from model organisms and non-NASA human astronauts. There are ~12 datasets from the Inspiration 4 (I4) mission, spanning metagenomics, comprehensive metabolic panels, clonal hematopoiesis, spatial transcriptomics, proteomics, and cytokine panels. In the interest of data privacy, two I4 datasets have raw FASTQ and FASTA files relating to the epitranscriptome, and a new request feature is live in OSDR (with a backend review process established) which was developed based on industry norms. OSDR also recently began a collaboration with the European Space Agency (ESA) to scientifically curate and make available >200 terabytes of human and model organism space-relevant data. The OSDR submission portal is designed to ingest and curate ~25 ‘omics assay data types, and ~50 physiological-phenotypic-imaging assay data types, spanning ultrasonography, micro-computed tomography, histology, morphometric photography, rebound tonometry, gait analysis, optical coherence tomography, novel object recognition, flow cytometry, and immunohistochemistry. A suite of analysis tools are available for OSDR users including: 1) an Environmental Data Application to compare radiation, CO2, relative humidity, temperature, and other telemetry across missions and subjects, 2) the RadLab database, a collaboration between NASA, ESA, the German and Italian Space Agencies, and the Bulgarian Academy of Sciences, which compiles radiation measurements relevant to human spaceflight and provides tools for accessing and manipulating the data, and 3) a Multi-study visualization tool which enables users to look across and combine GeneLab’s omics datasets across different experiments and missions. There are ~600 volunteer OSDR Analysis Working Group (AWG) members who: 1) provide feedback on scientific standards for reuse (subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability), and 2) collaborate to mine-reuse OSDR data conducting scientific analysis. OSDR has enabled 60 publications as of September 2023, many directly from AWG collaborations most notably the Cell Press package in 2020. Lastly, there are at least 15 articles which mine OSDR data part of a package of ~50 articles across Nature Portfolio with research stemming from I4, the Japan Aerospace Exploration Agency, NASA Space Biology, and the NASA Human Research Program.

open access↗