Search NASA⌕ Search

SEARCH · Search NASA

Results for “Protein databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Predicting protein functions from redundancies in large-scale protein interaction networks

Interpreting data from large-scale protein interaction experiments has been a challenging task because of the widespread presence of random false positives. Here, we present a network-based statistical algorithm that overcomes this difficulty and allows us to derive functions of unannotated proteins from large-scale interaction data. Our algorithm uses the insight that if two proteins share significantly larger number of common interaction partners than random, they have close functional associations. Analysis of publicly available data from Saccharomyces cerevisiae reveals >2,800 reliable functional associations, 29% of which involve at least one unannotated protein. By further analyzing these associations, we derive tentative functions for 81 unannotated proteins with high certainty. Our method is not overly sensitive to the false positives present in the data. Even after adding 50% randomly generated interactions to the measured data set, we are able to recover almost all (approximately 89%) of the original associations.

Proteins/chemistry/metabolism↗

RECOVIR Software for Identifying Viruses

Most single-stranded RNA (ssRNA) viruses mutate rapidly to generate a large number of strains with highly divergent capsid sequences. Determining the capsid residues or nucleotides that uniquely characterize these strains is critical in understanding the strain diversity of these viruses. RECOVIR (an acronym for "recognize viruses") software predicts the strains of some ssRNA viruses from their limited sequence data. Novel phylogenetic-tree-based databases of protein or nucleic acid residues that uniquely characterize these virus strains are created. Strains of input virus sequences (partial or complete) are predicted through residue-wise comparisons with the databases. RECOVIR uses unique characterizing residues to identify automatically strains of partial or complete capsid sequences of picorna and caliciviruses, two of the most highly diverse ssRNA virus families. Partition-wise comparisons of the database residues with the corresponding residues of more than 300 complete and partial sequences of these viruses resulted in correct strain identification for all of these sequences. This study shows the feasibility of creating databases of hitherto unknown residues uniquely characterizing the capsid sequences of two of the most highly divergent ssRNA virus families. These databases enable automated strain identification from partial or complete capsid sequences of these human and animal pathogens.

Chakravarty, Sugoto↗

Protein crystal growth - Growth kinetics for tetragonal lysozyme crystals

Results are reported from theoretical and experimental studies of the growth rate of lysozyme as a function of diffusion in earth-gravity conditions. The investigations were carried out to form a comparison database for future studies of protein crystal growth in the microgravity environment of space. A diffusion-convection model is presented for predicting crystal growth rates in the presence of solutal concentration gradients. Techniques used to grow and monitor the growth of hen egg white lysozyme are detailed. The model calculations and experiment data are employed to discuss the effects of transport and interfacial kinetics in the growth of the crystals, which gradually diminished the free energy in the growth solution. Density gradient-driven convection, caused by presence of the gravity field, was a limiting factor in the growth rate.

Pusey, M. L.↗

NASA GeneLab: Open Science for Life in Space

The NASA GeneLab project (genelab.nasa.gov) seeks to get the most from space-relevant biology experiments by providing and maintaining a public database consisting of DNA, RNA, protein, and metabolite data from spaceflight experiments. Since these types of data, referred to as omics data, are difficult to understand for non-bioinformaticians, the GeneLab data processing team works with the scientific community to develop methods to process these data. The processed data found on GeneLab reveals information about which genes are turned on and turned off in the space environment, which helps us understand how space changes our biology and how we can best mitigate these effects to travel deeper into space.

Jonathan Oribello↗

BIGEL analysis of gene expression in HL60 cells exposed to X rays or 60 Hz magnetic fields

We screened a panel of 1,920 randomly selected cDNAs to discover genes that are differentially expressed in HL60 cells exposed to 60 Hz magnetic fields (2 mT) or X rays (5 Gy) compared to unexposed cells. Identification of these clones was accomplished using our two-gel cDNA library screening method (BIGEL). Eighteen cDNAs differentially expressed in X-irradiated compared to control HL60 cells were recovered from a panel of 1,920 clones. Differential expression in experimental compared to control cells was confirmed independently by Northern blotting of paired total RNA samples hybridized to each of the 18 clone-specific cDNA probes. DNA sequencing revealed that 15 of the 18 cDNA clones produced matches with the database for genes related to cell growth, protein synthesis, energy metabolism, oxidative stress or apoptosis (including MYC, neuroleukin, copper zinc-dependent superoxide dismutase, TC4 RAS-like protein, peptide elongation factor 1alpha, BNIP3, GATA3, NF45, cytochrome c oxidase II and triosephosphate isomerase mRNAs). In contrast, BIGEL analysis of the same 1,920 cDNAs revealed no differences greater than 1.5-fold in expression levels in magnetic-field compared to sham-exposed cells. Magnetic-field-exposed and control samples were analyzed further for the presence of mRNA encoding X-ray-responsive genes by hybridization of the 18 specific cDNA probes to RNA from exposed and control HL60 cells. Our results suggest that differential gene expression is induced in approximately 1% of a random pool of cDNAs by ionizing radiation but not by 60 Hz magnetic fields under the present experimental conditions.

Non-NASA Center↗

Oxygen Deficiency in Spaceflight & its Impact on Plants’ Adaptive Changes

The goal of this study was to investigate the effects of hypoxic conditions in spaceflight. The distribution of genes involved with hypoxia in Arabidopsis thaliana and Brassica rapa were analyzed with the results from past spaceflight experiments to evaluate genes for future studies. Transcriptomes data of two different spaceflight studies of Arabidopsis thaliana from the NASA GeneLab database, GLDS-7 and GLDS-17, were compared. DNA microarrays were utilized for transcription profiling to conduct these studies. For GLDS-7, the response in spaceflight was studied with approaches that collected gene expression data. Leaves, hypocotyls, and root tissues were compared to the whole plant. For GLDS-17, seedlings and undifferentiated cultured cells were placed in the Biological Research in Canisters (BRIC), specifically BRIC-16. The genes related to hypoxia in Arabidopsis thaliana from these two studies were compared to genes in Brassica rapa with the TOAST database to evaluate similarities. When transcriptomes were analyzed for GLDS-7 and 17, genes that were considered significant had p-values ≤ 0.05 and log fold change values ≤ -1 or ≥1. Sixteen genes fulfilled the criteria. The genes related to hypoxia were alcohol dehydrogenase, elongation factor, ethylene-responsive factor, GUS, heat-shock proteins, NAP, RAP2.12, and RD20. The genes most impacted by spaceflight were heat-shock proteins. These genes were compared with Brassica rapa through Arabidopsis Ensemble Orthology from the TOAST Database. Similarities were seen in alcohol dehydrogenase, elongation factor, ethylene-responsive factor, heat-shock proteins, NAP, and RAP2.12. Overall, transcription profiling indicates that plants’ survival in spaceflight is dependent on adaptive changes with gene expression. This study also indicates that there are similarities in gene expression between Arabidopsis thaliana and Brassica rapa with comparable gene expression. Future studies could include analyzing additional species to understand which genes could be modified to ensure better yield of space crops amid hypoxic conditions.

hypoxia↗

GeneLab

GeneLab collects and enables analysis of spaceflight and ground-based spaceflight simulation genomic data, RNA and protein expression, and metabolic profiles. It interfaces with other existing databases containing spaceflight omic data. The 2011 National Research Council (NRC) Decadal Survey on NASA Life and Physical Sciences called for increased opportunities for multi-investigator spaceflight opportunities and greater use of genomic approaches to meet the needs of NASA researchers. To address these recommendations of the NRC Decadal Survey, the Space Life and Physical Sciences Research and Applications Division of NASA's Human Exploration and Operations Mission Directorate has initiated a transition to an Open Science architecture to increase research opportunities, and has developed the GeneLab Platform based on highly leveraged and integrated bioinformatics analytics. GeneLab is an interactive, open-access resource where scientists can upload, download, store, search, share, transfer, and analyze omics data from spaceflight and corresponding analogue experiments. Users can explore GeneLab datasets in the Data Repository, analyze data using the Analysis Platform, visualize high-order data and create collaborative projects using the Collaborative Workspace. Our primary goal is to maximize the utilization of the valuable biological research conducted aboard the International Space Station (ISS) by collecting genomic, transcriptomic, proteomic, and metabolomics data known as “omics”. By providing a portal linking processed data to flight parameters, GeneLab enables exploration of the molecular network responses of terrestrial biology to the space environment. This allows researchers to understand the complex responses of biological systems to the space environment. This technology development activity was transferred from the Human Exploration and Operations Mission Directorate to the Science Mission Directorate Division of Biological and Physical Sciences (BPS) in October 2020.

GeneLab↗

Finding the global minimum: a fuzzy end elimination implementation

The 'fuzzy end elimination theorem' (FEE) is a mathematically proven theorem that identifies rotameric states in proteins which are incompatible with the global minimum energy conformation. While implementing the FEE we noticed two different aspects that directly affected the final results at convergence. First, the identification of a single dead-ending rotameric state can trigger a 'domino effect' that initiates the identification of additional rotameric states which become dead-ending. A recursive check for dead-ending rotameric states is therefore necessary every time a dead-ending rotameric state is identified. It is shown that, if the recursive check is omitted, it is possible to miss the identification of some dead-ending rotameric states causing a premature termination of the elimination process. Second, we examined the effects of removing dead-ending rotameric states from further considerations at different moments of time. Two different methods of rotameric state removal were examined for an order dependence. In one case, each rotamer found to be incompatible with the global minimum energy conformation was removed immediately following its identification. In the other, dead-ending rotamers were marked for deletion but retained during the search, so that they influenced the evaluation of other rotameric states. When the search was completed, all marked rotamers were removed simultaneously. In addition, to expand further the usefulness of the FEE, a novel method is presented that allows for further reduction in the remaining set of conformations at the FEE convergence. In this method, called a tree-based search, each dead-ending pair of rotamers which does not lead to the direct removal of either rotameric state is used to reduce significantly the number of remaining conformations. In the future this method can also be expanded to triplet and quadruplet sets of rotameric states. We tested our implementation of the FEE by exhaustively searching ten protein segments and found that the FEE identified the global minimum every time. For each segment, the global minimum was exhaustively searched in two different environments: (i) the segments were extracted from the protein and exhaustively searched in the absence of the surrounding residues; (ii) the segments were exhaustively searched in the presence of the remaining residues fixed at crystal structure conformations. We also evaluated the performance of the method for accurately predicting side chain conformations. We examined the influence of factors such as type and accuracy of backbone template used, and the restrictions imposed by the choice of potential function, parameterization and rotamer database. Conclusions are drawn on these results and future prospects are given.

NASA Program Exobiology↗

(abstract) Modeling Protein Families and Human Genes: Hidden Markov Models and a Little Beyond

We will first give a brief overview of Hidden Markov Models (HMMs) and their use in Computational Molecular Biology. In particular, we will describe a detailed application of HMMs to the G-Protein-Coupled-Receptor Superfamily. We will also describe a number of analytical results on HMMs that can be used in discrimination tests and database mining. We will then discuss the limitations of HMMs and some new directions of research. We will conclude with some recent results on the application of HMMs to human gene modeling and parsing.

Hidden Markov Models HMMs proteins computational m↗

Analysis of xylem formation in pine by cDNA sequencing

Secondary xylem (wood) formation is likely to involve some genes expressed rarely or not at all in herbaceous plants. Moreover, environmental and developmental stimuli influence secondary xylem differentiation, producing morphological and chemical changes in wood. To increase our understanding of xylem formation, and to provide material for comparative analysis of gymnosperm and angiosperm sequences, ESTs were obtained from immature xylem of loblolly pine (Pinus taeda L.). A total of 1,097 single-pass sequences were obtained from 5' ends of cDNAs made from gravistimulated tissue from bent trees. Cluster analysis detected 107 groups of similar sequences, ranging in size from 2 to 20 sequences. A total of 361 sequences fell into these groups, whereas 736 sequences were unique. About 55% of the pine EST sequences show similarity to previously described sequences in public databases. About 10% of the recognized genes encode factors involved in cell wall formation. Sequences similar to cell wall proteins, most known lignin biosynthetic enzymes, and several enzymes of carbohydrate metabolism were found. A number of putative regulatory proteins also are represented. Expression patterns of several of these genes were studied in various tissues and organs of pine. Sequencing novel genes expressed during xylem formation will provide a powerful means of identifying mechanisms controlling this important differentiation pathway.

Non-NASA Center↗

A calmodulin binding protein from Arabidopsis is induced by ethylene and contains a DNA-binding motif

Calmodulin (CaM), a key calcium sensor in all eukaryotes, regulates diverse cellular processes by interacting with other proteins. To isolate CaM binding proteins involved in ethylene signal transduction, we screened an expression library prepared from ethylene-treated Arabidopsis seedlings with 35S-labeled CaM. A cDNA clone, EICBP (Ethylene-Induced CaM Binding Protein), encoding a protein that interacts with activated CaM was isolated in this screening. The CaM binding domain in EICBP was mapped to the C-terminus of the protein. These results indicate that calcium, through CaM, could regulate the activity of EICBP. The EICBP is expressed in different tissues and its expression in seedlings is induced by ethylene. The EICBP contains, in addition to a CaM binding domain, several features that are typical of transcription factors. These include a DNA-binding domain at the N terminus, an acidic region at the C terminus, and nuclear localization signals. In database searches a partial cDNA (CG-1) encoding a DNA-binding motif from parsley and an ethylene up-regulated partial cDNA from tomato (ER66) showed significant similarity to EICBP. In addition, five hypothetical proteins in the Arabidopsis genome also showed a very high sequence similarity with EICBP, indicating that there are several EICBP-related proteins in Arabidopsis. The structural features of EICBP are conserved in all EICBP-related proteins in Arabidopsis, suggesting that they may constitute a new family of DNA binding proteins and are likely to be involved in modulating gene expression in the presence of ethylene.

Non-NASA Center↗

Materials processing in space tasks. WBS task 5.4: Generic tasks

This task encompassed a wide range of activities related to materials processing in space. For example, all aspects of the space station's flight and ground based systems design were assessed for the Office of Advanced Concepts and Technology (OACT) Space Processing Division Office. Activities for that organization also included the consolidation of space processing payload requirements for the space station and the development of an OACT payload operations plan. Similar duties were performed for the MSFC Payload Project Office. The SPACECOM database was used to conduct preliminary design studies for microgravity payload carriers and to conduct assessments of materials processing technology. Concepts for the Advanced Protein Crystal Growth Facility (APCGF) were developed. Materials processing vent products were analyzed and a furnace facility filter concept was developed using those studies. A preliminary design for a space station aluminum payload rack was developed. Analysis was conducted to characterize the acceleration environment onboard the space shuttle. Equipment for two fluid experiment apparatus was designed and manufactured for the Space Science Laboratory. The Fluids and Materials Experiments (FAME) data base was expanded. Also, Mir payload integration, technology transfer, and spacelab-to-space station transition studies were conducted.

Crull, Robert↗

Shedding Light on Microbial Dark Matter with A Universal Language of Life

The majority of microbial genomes have yet to be cultured, and most proteins predicted from microbial genomes or sequenced from the environment cannot be functionally annotated. As a result, current computational approaches to describe microbial systems rely on incomplete reference databases that cannot adequately capture the full functional diversity of the microbial tree of life, limiting our ability to model high-level features of biological sequences. The scientific community needs a means to capture the functionally and evolutionarily relevant features underlying biology, independent of our incomplete reference databases. Such a model can form the basis for transfer learning tasks, enabling downstream applications in environmental microbiology, medicine, and bioengineering. Here we present LookingGlass, a deep learning model capturing a “universal language of life”. LookingGlass encodes contextually-aware, functionally and evolutionarily relevant representations of short DNA reads, distinguishing reads of disparate function, homology, and environmental origin. We demonstrate the ability of LookingGlass to be fine-tuned to perform a range of diverse tasks: to identify novel oxidoreductases, to predict enzyme optimal temperature, and to recognize the reading frames of DNA sequence fragments. LookingGlass is the first contextually-aware, general purpose pre-trained “biological language” representation model for short-read DNA sequences. LookingGlass enables functionally relevant representations of otherwise unknown and unannotated sequences, shedding light on the microbial dark matter that dominates life on Earth.

A Hoarfrost↗

The growing world of expansins

Expansins are cell wall proteins that induce pH-dependent wall extension and stress relaxation in a characteristic and unique manner. Two families of expansins are known, named alpha- and beta-expansins, and they comprise large multigene families whose members show diverse organ-, tissue- and cell-specific expression patterns. Other genes that bear distant sequence similarity to expansins are also represented in the sequence databases, but their biological and biochemical functions have not yet been uncovered. Expansin appears to weaken glucan-glucan binding, but its detailed mechanism of action is not well established. The biological roles of expansins are diverse, but can be related to the action of expansins to loosen cell walls, for example during cell enlargement, fruit softening, pollen tube and root hair growth, and abscission. Expansin-like proteins have also been identified in bacteria and fungi, where they may aid microbial invasion of the plant body.

NASA Discipline Plant Biology↗

CRITICA: coding region identification tool invoking comparative analysis

Gene recognition is essential to understanding existing and future DNA sequence data. CRITICA (Coding Region Identification Tool Invoking Comparative Analysis) is a suite of programs for identifying likely protein-coding sequences in DNA by combining comparative analysis of DNA sequences with more common noncomparative methods. In the comparative component of the analysis, regions of DNA are aligned with related sequences from the DNA databases; if the translation of the aligned sequences has greater amino acid identity than expected for the observed percentage nucleotide identity, this is interpreted as evidence for coding. CRITICA also incorporates noncomparative information derived from the relative frequencies of hexanucleotides in coding frames versus other contexts (i.e., dicodon bias). The dicodon usage information is derived by iterative analysis of the data, such that CRITICA is not dependent on the existence or accuracy of coding sequence annotations in the databases. This independence makes the method particularly well suited for the analysis of novel genomes. CRITICA was tested by analyzing the available Salmonella typhimurium DNA sequences. Its predictions were compared with the DNA sequence annotations and with the predictions of GenMark. CRITICA proved to be more accurate than GenMark, and moreover, many of its predictions that would seem to be errors instead reflect problems in the sequence databases. The source code of CRITICA is freely available by anonymous FTP (rdp.life.uiuc.edu in/pub/critica) and on the World Wide Web (http:/(/)rdpwww.life.uiuc.edu).

Non-NASA Center↗

Systemic Microgravity Response: Utilizing GeneLab to Develop Hypotheses for Spaceflight Risks

Biological risks associated with microgravity is a major concern for space travel. Although determination of risk has been a focus for NASA research, data examining systemic (i.e., multi- or pan-tissue) responses to space flight are sparse. The overall goal of our work is to identify potential master regulators responsible for such responses to microgravity conditions. To do this we utilized the NASA GeneLab database which contains a wide array of omics experiments, including data from: 1) different flight conditions (space shuttle (STS) missions vs. International Space Station (ISS); 2) different tissues; and 3) different types of assays that measure epigenetic, transcriptional, and protein expression changes. We have performed meta-analysis identifying potential master regulators involved with systemic responses to microgravity. The analysis used 7 different murine and rat data sets, examining the following tissues: liver, kidney, adrenal gland, thymus, mammary gland, skin, and skeletal muscle (soleus, extensor digitorum longus, tibialis anterior, quadriceps, and gastrocnemius). Using a systems biology approach, we were able to determine that p53 and immune related pathways appear central to pan-tissue microgravity responses. Evidence for a universal response in the form of consistency of change across tissues in regulatory pathways was observed in both STS and ISS experiments with varying durations; while degree of change in expression of these master regulators varied across species and strain, some change in these master regulators was universally observed. Interestingly, certain skeletal muscle (gastrocnemius and soleus) show an overall down-regulation in these genes, while in other types (extensor digitorum longus, tibialis anterior and quadriceps) they are up-regulated, suggesting certain muscle tissues may be compensating for atrophy responses caused by microgravity. Studying these organtissue-specific perturbations in molecular signaling networks, we demonstrate the value of GeneLab in characterizing potential master regulators associated with biological risks for spaceflight.

Microgravity↗

Materials dispersion and biodynamics project research

The Materials Dispersion and Biodynamics Project (MDBP) focuses on dispersion and mixing of various biological materials and the dynamics of cell-to-cell communication and intracellular molecular trafficking in microgravity. Research activities encompass biomedical applications, basic cell biology, biotechnology (products from cells), protein crystal development, ecological life support systems (involving algae and bacteria), drug delivery (microencapsulation), biofilm deposition by living organisms, and hardware development to support living cells on Space Station Freedom (SSF). Project goals are to expand the existing microgravity science database through experiments on sounding rockets, the Shuttle, and COMET program orbiters and to evolve,through current database acquisition and feasibility testing, to more mature and larger-scale commercial operations on SSF. Maximized utilization of SSF for these science applications will mean that service companies will have a role in providing equipment for use by a number of different customers. An example of a potential forerunner of such a service for SSF is the Materials Dispersion Apparatus (MDA) 'mini lab' of Instrumentation Technology Associates, Inc. (ITA) in use on the Shuttle for the Commercial MDAITA Experiments (CMIX) Project. The MDA wells provide the capability for a number of investigators to perform mixing and bioprocessing experiments in space. In the area of human adaptation to microgravity, a significant database has been obtained over the past three decades. Some low-g effects are similar to Earth-based disorders (anemia, osteoporosis, neuromuscular diseases, and immune system disorders). As new information targets potential profit-making processes, services and products from microgravity, commercial space ventures are expected to expand accordingly. Cooperative CCDS research in the above mentioned areas is essential for maturing SSF biotechnology and to ensure U.S. leadership in space technology. Currently, the MDBP conducts collaborative research with investigators at the Rockefeller University, National Cancer Institute, and the Universities of California, Arizona, and Alabama in Birmingham. The growing database from these collaborations provides fundamental information applicable to development of cell products, manipulation of immune cell response, bone cell growth and mineralization, and other processes altered by low-gravity. Contacts with biotechnology and biopharmaceutical companies are being increased to reach uninformed potential SSF users, provide access through the CMDS to interested users for feasibility studies, and to continue active involvement of current participants. We encourage and actively seek participation of private sector companies, and university and government researchers interested in biopharmaceuticals, hardware development and fundamental research in microgravity.

Lewis, Marian L.↗