Search NASA⌕ Search

SEARCH · Search NASA

Results for “proteomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Contrasting effects of glutamate and branched-chain amino acid metabolism on acid tolerance in a Castellaniella isolate from acidic groundwater

Groundwater acidification co-occurring with nitrate pollution is a common, global environmental health hazard. Denitrifying bacteria have been leveraged for the in situ removal of nitrate in groundwater. However, co-existing stressors—such as low pH—reduce the efficacy of biological removal processes. Castellaniella sp. str. MT123 is a complete denitrifier that was isolated from acidic, nitrate-contaminated groundwater. The strain grows robustly by nitrate respiration at pH < 6.0, completely reducing nitrate to dinitrogen gas. Genomic analyses of MT123 revealed few previously characterized acid tolerance genes. Thus, we utilized a combination of proteomics, metabolomics, and competitive mutant fitness to characterize the genetic mechanisms of MT123 acclimation to growth under mildly acidic conditions. We found that glutamate accumulation is critical in the acid acclimation of MT123, possibly through consumption of intracellular protons via glutamate decarboxylation to GABA. This is despite the fact that MT123 lacks the canonical glutamate decarboxylase-glutamate/GABA antiporter system implicated in acid tolerance in other bacteria. In contrast, branched-chain amino acid (BCAA) accumulation was detrimental to cell growth at lower pHs, possibly through indirect mechanisms impacting the cellular glutamate pool. Genetic analysis previously linked MT123 to a population of Castellaniella that bloomed—concurrent to nitrate removal—during a biostimulation effort to reduce groundwater nitrate concentrations at MT123’s location of origin. Thus, our analyses provide novel insight into mechanisms of acclimation to acidic conditions in a strain with significant potential for nitrate bioremediation.

59 BASIC BIOLOGICAL SCIENCES↗

A large-scale screening campaign of putative carbohydrate-active enzymes reveals a novel xylanase from anaerobic gut fungi

The genomes of anaerobic gut fungi (AGF) encode a diverse array of carbohydrate-active enzymes (CAZymes), yet exceedingly few of these enzymes have been experimentally validated or expressed in heterologous systems. Here, we developed a predictive bioinformatic pipeline to annotate novel putative CAZymes from anaerobic fungi and validate their activity through large-scale heterologous expression in Escherichia coli. A total of 173 fungal proteins from Piromyces finnis associated with biomass degradation were synthesized and expressed in E. coli, and 9.8% were soluble with expression levels exceeding 5% of the total proteome using high-throughput proteomic screening. Among these 17 heterologously expressed proteins, analysis with AlphaFold and FoldSeek predicted 13 multi-functional proteins containing catalytic domains fused with repetitive fungal dockerins, and half of the substrate predictions were experimentally validated. One promising enzyme, celsome_012, exhibited robust and specific activity against beechwood xylan at 37°C and pH 6.4, with titers that were also fivefold higher than those of other recombinant proteins screened here. Both Michaelis-Menten kinetics and the linearized Lineweaver-Burk equation yielded consistent values for K m , and its activation energy was estimated at 51.9 kJ/mol based on the Arrhenius model. This work supports the industrial translation of anaerobic fungal CAZymes due to their robust lignocellulolytic activity and provides a framework for prioritizing AGF proteins for efficient E. coli heterologous expression.

59 BASIC BIOLOGICAL SCIENCES↗

The anaerobic fungus Caecomyces churrovis produces H2 via a non-3 bifurcating NADH-dependent enzyme complex

Anaerobic fungi (AF) decompose lignocellulose-based biomass into fermentable sugars through the production of powerful biomass-degrading enzymes. AF are unusual among fungi in that they generate energy via hydrogenosomes, which are also associated with the release of H2 though yet unknown metabolic mechanisms. In particular, it remains unclear how NAD(P)+ is regenerated within hydrogenosomes and how H2 is formed. Here, we reveal the molecular mechanism for hydrogenosomal H2 production in the AF strain C. churrovis by combining genomic search, proteomic analysis, and enzymology. Our enzyme assays on the large organelle fraction of C. churrovis revealed the activity of H2:NAD+ oxidoreductase but not pyruvate:ferredoxin oxidoreductase activity. We identified genes encoding [FeFe] hydrogenase (Hyd) and NADH dehydrogenase subunits E and F (NuoE, NuoF) in C. churrovis, and confirmed their expression in the isolated hydrogenosomal fractions by proteomic analysis. Combining the individually purified proteins, we found that the assay system consisting of Hyd-Strep and NuoEF-Strep reduced NAD+ with H2. Furthermore, this system formed H2 directly from NADH independent of ferredoxin, functioning as a non-bifurcating NADH-dependent enzyme rather than an electron-bifurcating enzyme. We identified homologs of hydrogenosomal NuoE, NuoF, and Hyd in many other AF, indicating this pathway is widely conserved among the early-branching AF. This work demonstrates the existence of a non-bifurcating NADH-dependent enzyme complex in eukaryotes. Moreover, this complex could be a target for controlling AF H2 production and altering fungal metabolism.

fungi↗

Regulation of bacterial stringent response by an evolutionarily conserved ribosomal protein L11 methylation

Lysine and arginine methylation is an important regulator of enzyme activity and transcription in eukaryotes. However, little is known about this covalent modification in bacteria. In this work, we investigated the role of methylation in bacteria. By reanalyzing a large phyloproteomics data set from 48 bacterial strains representing six phyla, we found that almost a quarter of the bacterial proteome is methylated. Many of these methylated proteins are conserved across diverse bacterial lineages, including those involved in central carbon metabolism and translation. Among the proteins with the most conserved methylation sites is ribosomal protein L11 (bL11). bL11 methylation has been a mystery for five decades, as the deletion of its methyltransferase PrmA causes no cell growth defects. Comparative proteomics analysis combined with inorganic polyphosphate and guanosine tetra/pentaphosphate assays of the ΔprmA mutant in Escherichia coli revealed that bL11 methylation is important for stringent response signaling. In the stationary phase, we found that the ΔprmA mutant has impaired guanosine tetra/pentaphosphate production. This leads to a reduction in inorganic polyphosphate levels, accumulation of RNA and ribosomal proteins, and an abnormal polysome profile. Overall, our investigation demonstrates that the evolutionarily conserved bL11 methylation is important for stringent response signaling and ribosomal activity regulation and turnover.

59 BASIC BIOLOGICAL SCIENCES↗

EMSL-Computing/pspecterlib

Proteomics R package for matching peptide fragments for both digested and intact proteomics

Degnan, David [Pacific Northwest National Laborato↗

Kinetic Deep Learning v0.1

Here, we present a method that uses protein levels to predict times series of metabolite concentrations. Understanding this type of pathway dynamics is important in order to predict the behavior of the pathway and, more pragmatically, to be able to design biological systems (such as strains bioengineered to produce chemical products) reliably. Typically, for this purpose, kinetic models consisting of differential equations based on the Michaelis-Menten dynamics have been used in the past. However, these methods can rarely produce good fits to measured data time series. Possibly, this happens because the kinetic constants are unknown or are different from the ones measured in vivo, or perhaps because Michaelis-Menten dynamics is not a satisfactory description. In order to improve the predictive nature of these kinetic models we have eliminated the Michaelis-Menten description of pathway dynamics and we have substituted it by algorithms that automatically learn these dynamics from previously obtained metabolomics and proteomics data using machine learning approaches. Specifically, kinetic deep learning uses deep learning to map proteomics time series to metabolite concentration time series, instead of learning the first metabolite derivative and integrating in (as in the first version of kinetic learning). This approach is shown to provide good to excellent results with a data set specifically collected for this purpose.

Garcia Martin, Hector [Joint BioEnergy Institute (↗

Evaluating isoprenol production using the IPP-bypass pathway in the oleaginous yeast Rhodosporidium toruloides

Background To strengthen the national energy supply, there is an increasing demand for domestically generated aviation fuels. Bio-derived advanced aviation fuels offer the opportunity to meet this domestic need while presenting a unique opportunity to investigate the production of novel aviation fuels. Isoprenol, a chemical precursor to such novel fuels, has been shown to be a biologically producible compound in model organisms, but its bio-producibility needs to be further explored in organisms more compatible with industrial bioproduction. Results In this work, we evaluate isoprenol production using the promising bioproduction yeast, Rhodosporidium toruloides. First, we show successful isoprenol production using the IPP-bypass pathways most successful in laboratory strains of E. coli and S. cerevisiae. Next, we demonstrate that increased flux through the mevalonate pathway only modestly increases isoprenol titers. Using proteomics, we identified a potential bottleneck in production at the final step in the IPP-bypass pathway and explored alternative enzymes for this step. Finally, the top three strains of R. toruloides were evaluated in sorghum hydrolysates generated using cholinium lysinate. Through this work, 93.1 mg/L of isoprenol was produced in mock medium and 27.3 mg/L in sorghum hydrolysates. Conclusion Together these results lay the foundation for future work for the production of isoprenol from bioproduction crops.

Advanced aviation fuel↗

Rhodotorula toruloides Nitrogen Limitation PTM Profiling Multi-Omics (TZ-DP1)

The purpose of this experiment was to evaluate the regulatory stress response of Oleaginous yeast species Rhodotorula toruloides NBRC 0880 (JGI strain IFO0880 v4.0) under nitrogen-rich and nitrogen-limited conditions over time. Time course experimental samples (0, 24, 48, and 72 hours after inoculation) were prepared using a semi-automated multi-PTM proteomic approach, using tandem mass tag 18-plex (TMT18), and lipidome remodeling for downstream multi-omics analysis. Processed datasets are openly accessible from PNNL DataHub and contain secondary processed proteomic (redox, phospho, and global TMT) and lipidomic (positive and negative ion mode) results files and experimental design metadata.

59 BASIC BIOLOGICAL SCIENCES↗

Multi-omics data resource: Data package 25 (Pck025)

This data package comprises omics datasets from human pancreatic islets treated with IL-1β + IFNγ or with estrogen (E2) for 18 h. Two RNA-seq datasets are available: the first is a discovery dataset involving human islets treated with or without IL-1β + IFNγ for 18 hours; the second is a validation dataset, where human islets are treated with or without IL-1β + IFNγ or E2 for 18 hours. DIA proteomic analysis was performed on the same validation dataset samples. Data contributors: Kiersten L. Webster, Sarah Tersey & Raghavendra G. Mirmir: Kovler Diabetes Center and Department of Medicine, The University of Chicago, Chicago, IL, 60637, USA. Soumyadeep Sarkar, Raghavendra Mirmira, Ernesto S. Nakayasu: Biological Sciences Division, Pacific Northwest National Laboratory, Richland, WA, 99354, USA. Data repository: RNA-seq: GSE310965 Proteomics: MSV000101892 Publication: PMID 41279069

Sarkar, Soumyadeep [Pacific Northwest National Lab↗

Elevated Temperature Effects on Protein Turnover Dynamics in Arabidopsis thaliana Seedlings Revealed by 15 N-Stable Isotope Labeling and ProteinTurnover Algorithm

Global warming poses a threat to plant survival, impacting growth and agricultural yield. Protein turnover, a critical regulatory mechanism balancing protein synthesis and degradation, is crucial for the cellular response to environmental changes. We investigated the effects of elevated temperature on proteome dynamics in Arabidopsis thaliana seedlings using 15 N-stable isotope labeling and ultra-performance liquid chromatography-high resolution mass spectrometry, coupled with the ProteinTurnover algorithm. Analyzing different cellular fractions from plants grown under 22 °C and 30 °C growth conditions, we found significant changes in the turnover rates of 571 proteins, with a median 1.4-fold increase, indicating accelerated protein dynamics under thermal stress. Notably, soluble root fraction proteins exhibited smaller turnover changes, suggesting tissue-specific adaptations. Significant turnover alterations occurred with redox signaling, stress response, protein folding, secondary metabolism, and photorespiration, indicating complex responses enhancing plant thermal resilience. Conversely, proteins involved in carbohydrate metabolism and mitochondrial ATP synthesis showed minimal changes, highlighting their stability. This analysis highlights the intricate balance between proteome stability and adaptability, advancing our understanding of plant responses to heat stress and supporting the development of improved thermotolerant crops.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Metalloproteomics Reveals Multi-Level Stress Response in Escherichia coli When Exposed to Arsenite

The arsRBC operon encodes a three-protein arsenic resistance system. ArsR regulates the transcription of the operon, while ArsB and ArsC are involved in exporting trivalent arsenic and reducing pentavalent arsenic, respectively. Previous research into Agrobacterium tumefaciens 5A has demonstrated that ArsR has regulatory control over a wide range of metal-related proteins and metabolic pathways. We hypothesized that ArsR has broad regulatory control in other Gram-negative bacteria and set out to test this. Here, we use differential proteomics to investigate changes caused by the presence of the arsR gene in human microbiome-relevant Escherichia coli during arsenite (AsIII) exposure. We show that ArsR has broad-ranging impacts such as the expression of TCA cycle enzymes during AsIII stress. Additionally, we found that the Isc [Fe-S] cluster and molybdenum cofactor assembly proteins are upregulated regardless of the presence of ArsR under these same conditions. An important finding from this differential proteomics analysis was the identification of response mechanisms that were strain-, ArsR-, and arsenic-specific, providing new clarity to this complex regulon. Given the widespread occurrence of the arsRBC operon, these findings should have broad applicability across microbial genera, including sensitive environments such as the human gastrointestinal tract.

Biochemistry & Molecular Biology↗

Current State, Challenges, and Opportunities in Genome-Scale Resource Allocation Models: A Mathematical Perspective

Stoichiometric genome-scale metabolic models (generally abbreviated GSM, GSMM, or GEM) have had many applications in exploring phenotypes and guiding metabolic engineering interventions. Nevertheless, these models and predictions thereof can become limited as they do not directly account for protein cost, enzyme kinetics, and cell surface or volume proteome limitations. Lack of such mechanistic detail could lead to overly optimistic predictions and engineered strains. Initial efforts to correct these deficiencies were by the application of precursor tools for GSMs, such as flux balance analysis with molecular crowding. In the past decade, several frameworks have been introduced to incorporate proteome-related limitations using a genome-scale stoichiometric model as the reconstruction basis, which herein are called resource allocation models (RAMs). This review provides a broad overview of representative or commonly used existing RAM frameworks. This review discusses increasingly complex models, beginning with stoichiometric models to precursor to RAM frameworks to existing RAM frameworks. RAM frameworks are broadly divided into two categories: coarse-grained and fine-grained, with different strengths and challenges. Discussion includes pinpointing their utility, data needs, highlighting framework strengths and limitations, and appropriateness to various research endeavors, largely through contrasting their mathematical frameworks. Finally, promising future applications of RAMs are discussed.

59 BASIC BIOLOGICAL SCIENCES↗

Untargeted GC-MS Metabolic Profiling of Anaerobic Gut Fungi Reveals Putative Terpenoids and Strain-Specific Metabolites

Background/Objectives: Anaerobic gut fungi (Neocallimastigomycota) are biotechnologically relevant, lignocellulose-degrading microbes with under-explored biosynthetic potential for secondary metabolites. Untargeted metabolomic profiling with gas chromatography–mass spectrometry (GC-MS) was applied to two gut fungal strains, Anaeromyces robustus and Caecomyces churrovis, to establish a foundational metabolomic dataset to identify metabolites and provide insights into gut fungal metabolic capabilities. Methods: Gut fungi were cultured anaerobically in rumen-fluid-based media with a soluble substrate (cellobiose), and metabolites were extracted using the Metabolite, Protein, and Lipid Extraction (MPLEx) method, enabling metabolomic and proteomic analysis from the same cell samples. Samples were derivatized and analyzed via GC-MS, followed by compound identification by spectral matching to reference databases, molecular networking, and statistical analyses. Results: Distinct metabolites were identified between A. robustus and C. churrovis, including 2,3-dihydroxyisovaleric acid produced by A. robustus and maltotriitol, maltotriose, and melibiose produced by C. churrovis. C. churrovis may polymerize maltotriose to form an extracellular polysaccharide, like pullulan. GC-MS profiling potentially captured sufficiently volatile products of proteomically detected, putative non-ribosomal peptide synthetases and polyketide synthases of A. robustus and C. churrovis. The triterpene squalene and triterpenoid tetrahymanol were putatively identified in A. robustus and C. churrovis. Their conserved, predicted biosynthetic genes—squalene synthase and squalene tetrahymanol cyclase—were identified in A. robustus, C. churrovis, and other anaerobic gut fungal genera. Conclusions: This study provides a foundational, untargeted metabolomic dataset to unmask gut fungal metabolic pathways and biosynthetic potential and to prioritize future efforts for compound isolation and identification.

Biochemistry & Molecular Biology↗

Lipids accumulation in nitrogen and phosphorus-limited yeast is caused by less growth-related dilution

Oleaginous yeasts are used commercially to produce oleochemicals and hold potential also for biodiesel production. In response to nitrogen or phosphorous limitation, oleaginous yeast accumulate triacylglycerol. Previous work has investigated potential mechanisms by which nutrient limitation induces lipid biosynthesis without verifying whether lipid biosynthesis flux is actually enhanced. Here we show, using 13C-glucose tracing, that in nitrogen or phosphorous limitation, lipid accumulation occurs without consistent increases in biosynthetic flux. Instead, the main driver of increased lipid pools is decreased growth-related dilution. This conclusion holds across two divergent oleaginous yeasts: Rhodotorula toruloides and Yarrowia lipolytica. Quantitative proteomics shows a substantial proteome reallocation in response to nitrogen and phosphorous limitation, with ribosomal proteins strongly down regulated, while lipid biosynthetic enzymes are preserved but not consistently upregulated in absolute quantity. Thus, nutrient limitation, rather than triggering greatly enhanced lipid synthesis, results in roughly sustained lipid enzyme levels and biosynthetic flux. Due to slower lipid dilution by cell division, this suffices to drive marked lipid accumulation.

Keber, Felix [Lewis Sigler Institute for Integrati↗

The anaerobic fungus Caecomyces churrovis produces H 2 via a non-bifurcating NADH-dependent enzyme complex

ABSTRACT Hydrogenosomes are mitochondria-derived organelles that produce ATP and H 2 to support energy metabolism in anaerobic eukaryotes. H 2 production allows reoxidation of reduced cofactors generated during fermentative metabolism; however, the metabolic mechanisms for H 2 production in anaerobic eukaryotes remains incompletely understood. In particular, it remains unclear whether anaerobic fungi (AF) hydrogenosomes use a ferredoxin-dependent pathway or a distinct mechanism to regenerate NAD(P) + and link electron transfer to H 2 formation. Here, by combining genomic search, proteomic analysis, and enzymology, we reveal the molecular mechanism for H 2 production in the AF strain Caecomyces churrovis . Our enzyme assays on the organelle fraction of C. churrovis revealed the activity of H 2 :NAD + oxidoreductase but not pyruvate:ferredoxin oxidoreductase, which is usually linked to H 2 formation. We identified genes encoding [FeFe] hydrogenase (Hyd) and NADH dehydrogenase subunits E and F (NuoE, NuoF) in C. churrovis , and confirmed their expression in the isolated hydrogenosomal fractions by proteomic analysis. Combining the individually purified enzymes, we found Hyd and NuoEF proteins formed H 2 directly from NADH independently of ferredoxin, functioning as a non-bifurcating NADH-dependent enzyme rather than an electron-bifurcating enzyme. We identified homologs of hydrogenosomal NuoE, NuoF, and Hyd in many other AF, indicating this pathway is commonly shared among the AF. This work demonstrates the existence of a non-bifurcating NADH-dependent enzyme complex in eukaryotes. Moreover, this complex could potentially be exploited as a target for controlling AF H 2 production and altering fungal metabolism. IMPORTANCE H 2 production is a prominent feature of anaerobic energy metabolism, yet our understanding of eukaryotic mechanisms remains limited. Anaerobic fungi (AF) are key decomposers of lignocellulose and contribute to hydrogen flux in anaerobic environments. Although it has been more than 40 years since the H 2 production from AF was first reported, the molecular mechanism for hydrogenosomal H 2 production and redox balance remains unclear. We demonstrate that AF produce H 2 from NADH utilizing a non-bifurcating NADH-dependent enzyme complex rather than an electron-bifurcating, ferredoxin-dependent variant. We show that this enzyme complex is conserved across multiple AF lineages and thus demonstrate the occurrence of a non-bifurcating NADH-dependent enzyme in eukaryotes. This discovery expands our understanding of eukaryotic hydrogenosomal metabolism, reveals a previously unknown strategy for redox balancing, and highlights potential targets for manipulating H 2 production. These insights have broad implications for microbial energy metabolism, anaerobic ecosystems, and bioengineering of H 2 -producing systems.

Zhang, Bo [Department of Chemical Engineering, Uni↗

Multiparametric Determination of Radiation Risk

Predicting risk of human cancer following exposure to ionizing space radiation is challenging in part because of uncertainties of low-dose distribution amongst cells, of unknown potentially synergistic effects of microgravity upon cellular protein-expression, and of processing dose-related damage within cells to produce rare and late-appearing malignant transformation, degrade the confidence of cancer risk-estimates. The NASA- specific responsibility to estimate the risks of radiogenic cancer in a limited number of astronauts is not amenable to epidemiologic study, thereby increasing this challenge. Developing adequately sensitive cellular biodosimeters that simultaneously report 1) the quantity of absorbed close after exposure to ionizing radiation, 2) the quality of radiation delivering that dose, and 3) the risk of developing malignant transformation by the cells absorbing that dose could be useful for resolving these challenges. Use of a multiparametric cellular biodosimeter is suggested using analyses of gene-expression and protein-expression whereby large datasets of cellular response to radiation-induced damage are obtained and analyzed for expression-profiles correlated with established end points and molecular markers predictive for cancer-risk. Analytical techniques of genomics and proteomics may be used to establish dose-dependency of multiple gene- and protein- expressions resulting from radiation-induced cellular damage. Furthermore, gene- and protein-expression from cells in microgravity are known to be altered relative to cells grown on the ground at 1g. Therefore, hypotheses are proposed that 1) macromolecular expression caused by radiation-induced damage in cells in microgravity may be different than on the ground, and 2) different patterns of macromolecular expression in microgravity may alter human radiogenic cancer risk relative to radiation exposure on Earth. A new paradigm is accordingly suggested as a national database wherein genomic and proteomic datasets are registered and interrogated in order to provide statistically significant dose-dependent risk estimation of radiogenic cancer in astronauts.

Richmond, Robert C.↗

Focused Metabolite Profiling for Dissecting Cellular and Molecular Processes of Living Organisms in Space Environments

Regulatory control in biological systems is exerted at all levels within the central dogma of biology. Metabolites are the end products of all cellular regulatory processes and reflect the ultimate outcome of potential changes suggested by genomics and proteomics caused by an environmental stimulus or genetic modification. Following on the heels of genomics, transcriptomics, and proteomics, metabolomics has become an inevitable part of complete-system biology because none of the lower "-omics" alone provide direct information about how changes in mRNA or protein are coupled to changes in biological function. The challenges are much greater than those encountered in genomics because of the greater number of metabolites and the greater diversity of their chemical structures and properties. To meet these challenges, much developmental work is needed, including (1) methodologies for unbiased extraction of metabolites and subsequent quantification, (2) algorithms for systematic identification of metabolites, (3) expertise and competency in handling a large amount of information (data set), and (4) integration of metabolomics with other "omics" and data mining (implication of the information). This article reviews the project accomplishments.

Source record↗

GeneLab Phase 2: Integrated Search Data Federation of Space Biology Experimental Data

The GeneLab project is a science initiative to maximize the scientific return of omics data collected from spaceflight and from ground simulations of microgravity and radiation experiments, supported by a data system for a public bioinformatics repository and collaborative analysis tools for these data. The mission of GeneLab is to maximize the utilization of the valuable biological research resources aboard the ISS by collecting genomic, transcriptomic, proteomic and metabolomic (so-called omics) data to enable the exploration of the molecular network responses of terrestrial biology to space environments using a systems biology approach. All GeneLab data are made available to a worldwide network of researchers through its open-access data system. GeneLab is currently being developed by NASA to support Open Science biomedical research in order to enable the human exploration of space and improve life on earth. Open access to Phase 1 of the GeneLab Data Systems (GLDS) was implemented in April 2015. Download volumes have grown steadily, mirroring the growth in curated space biology research data sets (61 as of June 2016), now exceeding 10 TB/month, with over 10,000 file downloads since the start of Phase 1. For the period April 2015 to May 2016, most frequently downloaded were data from studies of Mus musculus (39) followed closely by Arabidopsis thaliana (30), with the remaining downloads roughly equally split across 12 other organisms (each 10 of total downloads). GLDS Phase 2 is focusing on interoperability, supporting data federation, including integrated search capabilities, of GLDS-housed data sets with external data sources, such as gene expression data from NIHNCBIs Gene Expression Omnibus (GEO), proteomic data from EBIs PRIDE system, and metagenomic data from Argonne National Laboratory's MG-RAST. GEO and MG-RAST employ specifications for investigation metadata that are different from those used by the GLDS and PRIDE (e.g., ISA-Tab). The GLDS Phase 2 system will implement a Google-like, full-text search engine using a Service-Oriented Architecture by utilizing publicly available RESTful web services Application Programming Interfaces (e.g., GEO Entrez Programming Utilities) and a Common Metadata Model (CMM) in order to accommodate the different metadata formats between the heterogeneous bioinformatics databases. GLDS Phase 2 completion with fully implemented capabilities will be made available to the general public in September 2017.

Space Biology↗