Search NASASearch

SEARCH · Search NASA

Results for “Biological databases”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Database of Nonaqueous Proton-Conducting Materials

This work presents the assembly of 48 papers, representing 74 different compounds and blends, into a machine-readable database of nonaqueous proton-conducting materials. SMILES was used to encode the chemical structures of the molecules, and we tabulated the reported proton conductivity, proton diffusion coefficient, and material composition for a total of 3152 data points. The data spans a broad range of temperatures ranging from -70 to 260 °C. To explore this landscape of nonaqueous proton conductors, DFT was used to calculate the proton affinity of 18 unique proton carriers. The results were then compared to the activation energy derived from fitting experimental data to the Arrhenius equation. It was found that while the widely recognized positive correlation between the activation energy and proton affinity may hold among closely related molecules, this correlation does not necessarily apply across a broader range of molecules. This work serves as an example of the potential analyses that can be conducted using literature data combined with emerging research tools in computation and data science to address specific materials design problems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Harvest Initiated Volatile Organic Compound Emissions from In-Field Tall Wheatgrass

While crop and grassland usage continues to increase, the full diversity of plant-specific volatile organic compounds (VOCs) emitted from these ecosystems, including their implications for atmospheric chemistry and carbon cycling, remains poorly understood. It is particularly important to investigate VOCs in the context of potential biofuels: aside from the implications of largescale land use, harvest may shift both the flux and speciation of emitted VOCs. To this point, we evaluate the diversity of VOCs emitted both pre and postharvest from “Alkar” tall wheatgrass (Thinopyrum ponticum), a candidate biofuel that exhibits greater tolerance to frost and saline land compared to other grass varieties. Mature plants grown under field conditions (n = 6) were sampled for VOCs both pre- and postharvest (October 2022). Via hierarchical clustering of emitted VOCs from each plant, we observe distinct “volatilomes” (diversity of VOCs) specific to the pre- and postharvest conditions despite plant-to-plant variability. In total, 50 VOCs were found to be unique to the postharvest tall wheatgrass volatilome, and these unique VOCs constituted a significant portion (26%) of total postharvest signal. While green leaf volatiles (GLVs) dominate the speciation of postharvest emissions (e.g., 54% of unique postharvest VOC signal was due to 1-penten-3-ol), we demonstrate novel postharvest VOCs from tall wheatgrass that are under characterized in the context of carbon cycling and atmospheric chemistry (e.g., 3-octanone). Continuing evaluations will quantitatively investigate tall wheatgrass VOC fluxes, better informing the feasibility and environmental impact of tall wheatgrass as a biofuel.

09 BIOMASS FUELS

Computationally evaluating high-yield metabolites for sustainable aviation fuel (SAF) using machine learning

The computational tool described in this report helps identify promising biological pathways that produce SAF platform molecules (either a drop-in SAF, or a precursor that can be easily converted to a drop-in SAF). The workflow the computational tool follows first identifies possible biological pathways from a user-defined metabolite. These pathways may, or may not lead to a SAF platform molecule, thus the second step involves insilico testing of the end product of each pathway to assess whether it is, or is not, a SAF platform molecule. The identification of biological pathways performed in the first step is facilitated by linking the metabolite to a biological reaction database. Pathways are found by identifying pathways in the reaction database that include the metabolite. The computational tool includes an alternative way to find pathways. The alternative way develops a Flux Balanced Analysis (FBA), and modifying the FBA to include reactions that transform the metabolite. These modifications serve as a basis for understanding, in a semi-quantitative way, if there is an increase in the flux to desirable products. The second step, in silico testing of the end-products, is accomplished by estimating key physical properties relevant to SAF. When good models are available, we have integrated those models into the computational tool. In a few instances, we have developed our own models. In all instances, we have validated the models against available measured data. Finally, we have evaluated the effectiveness of our computational tool by genetically engineering Rhodosporidium toruloides. Validation occurred without the use of a FBA, and further validation is required.

09 BIOMASS FUELS

Association of orogenic activity with the Ordovician radiation of marine life

The Ordovician radiation of marine life was among the most substantial pulses of diversification in Earth history and coincided in time with a major increase in the global level of orogenic activity. To investigate a possible causal link between these two patterns, the geographic distributions of 6576 individual appearances of Ordovician vician genera around the world were evaluated with respect to their proximity to probable centers of orogeny (foreland basins). Results indicate that these genera, which belonged to an array of higher taxa that diversified in the Middle and Late Ordovician (trilobites, brachiopods, bivalves, gastropods, monoplacophorans), were far more diverse in, and adjacent to, foreland basins than they were in areas farther removed from orogenic activity (carbonate platforms). This suggests an association of orogeny with diversification at that time.

NASA Discipline Exobiology

Open Science for Life in Space: Bioimaging, Data Sharing, and Tools for Knowledge Discovery

Precious space-flown biological experiments have both multi-omic and phenotypic data which NASA strives to make maximally open access for reuse. Currently a number of these space-relevant bioimaging datasets are being reused for AI/ML approaches. NASA Ames Life Science Data Archive and NASA GeneLab are working to make all current and future bioimaging data even more accessible and reusable. Standards for collection and curation are being implemented to enable scientists worldwide access to these data for further discovery and use.

data science

RadLab: A Comprehensive Database and Graphical and Programming Interfaces for Biologically Relevant Space Radiation Data

RadLab, a new component of the NASA Open Science Data Repository (OSDR), is a platform built upon a database of radiation data relevant to space biology. RadLab provides visual and programmatic interfaces for interrogation of its database, as well as a submission process for inclusion of data from investigators. The RadLab application programming interface (API) implements a request syntax enabling users to retrieve data filtered by various combinations of parameters (detector type, location, direction, timespan, etc), which are delivered in machine-readable text formats, ready to be ingested by downstream analysis pipelines; while the graphical user interface (GUI) provides easy means to iteratively modify query parameters and incorporates a number of standard analyses and visualizations (time series plots, geospatial visualizations, detector comparison). Investigators from many countries, including US, Russia, Japan, Canada, the Czech Republic, Germany, Hungary, and Italy, have committed to provide data from their instruments located on the ISS; RadLab will also include data from other spacecraft in LEO (e.g., the Space Shuttle, the Mir space station), BLEO (e. g. BioSentinel, Mars Orbiter, among others), and on other celestial bodies (e. g. Chang’e 4, Curiosity). The first release of RadLab has been made available to the public. Once fully operational, RadLab will provide a comprehensive and ever-growing compendium of space radiation data, facilitating straightforward access to multiple types of readings and enabling space biology researchers to perform intercomparisons of detectors and to determine the radiation environment of research missions, both via programmatic retrieval of these data and via the graphical analysis toolkit; as well as a user-friendly submission portal for ingesting data from space agencies and research institutions. Radiation scientists will be able to use RadLab to gain a deeper understanding of the space radiation environment for future human space exploration. The RadLab Working Group has been formed to foster close collaborations among data contributors and users, to identify data sources, to put in place standards for data normalization, to guide the development of features of the analysis toolkit, to establish the use of RadLab in space radiation biology research, and eventually to provide a forum for discussing relevant research issues that can take advantage of RadLab's capabilities.

radiation

RadLab: A Comprehensive Database and Graphical and Programming Interfaces for Biologically Relevant Space Radiation Data

RadLab, a new component of the NASA Open Science Data Repository (OSDR), comprises a database of radiation measurements relevant to space biology, and visual and programmatic interfaces for interrogation and retrieval of these data. The attributes of data available through RadLab include spacecraft, types of radiation sensing instruments, locations within the spacecraft (e.g. modules of the ISS), associated celestial bodies, trajectories, and spacecraft coordinates. The application programming interface (API) implements a request syntax for retrieval of timestamped data filtered by various combinations of such attributes; the graphical user interface (GUI) extends this functionality with visualizations, such as spacecraft schematics, time series plots, geospatial visualizations, and provides easy means to iteratively refine search parameters, inspect the data on the fly, and download target subsets. The release of RadLab currently available to the public contains datasets provided by US and international collaborators and focuses on data recorded on the ISS. Investigators from multiple countries, including the US, Canada, Germany, Bulgaria, Hungary, Italy, Japan, Russia and the Czech Republic, have committed to provide data from their instruments in and beyond low Earth orbit; RadLab will also soon expand to include past (e.g. Shuttle and Mir) and future (e.g. Artemis) data. RadLab will provide a comprehensive, dynamic compendium of space radiation data, enabling the scientific community to perform analyses of data from multiple detectors and to determine the radiation environment of research missions and experiments. The RadLab Working Group has been formed to foster collaborations among data contributors and users, to identify data sources, to put in place standards for data harmonization, and to guide the development of the platform, with the goal to establish the use of RadLab in space radiation research and to advance our understanding of the space radiation environment in human habitats.

database

RadLab: A Comprehensive Database and Graphical and Programming Interfaces for Biologically Relevant Space Radiation Data

RadLab, a new component of the NASA Open Science Data Repository (OSDR), comprises a database of radiation measurements relevant to space biology, and visual and programmatic interfaces for interrogation and retrieval of these data. The attributes of data available through RadLab include spacecraft, types of radiation sensing instruments, locations within the spacecraft (e.g. modules of the ISS), associated celestial bodies, trajectories, and spacecraft coordinates. The application programming interface (API) implements a request syntax for retrieval of timestamped data filtered by various combinations of such attributes; the graphical user interface (GUI) extends this functionality with visualizations, such as spacecraft schematics, time series plots, geospatial visualizations, and provides easy means to iteratively refine search parameters, inspect the data on the fly, and download target subsets of these data. The release of RadLab currently available to the public contains datasets provided by US and international collaborators and focuses on data recorded on the ISS. Investigators from multiple countries, including the US, Canada, Germany, Bulgaria, Hungary, Italy, Japan, Russia and the Czech Republic, have committed to provide data from their instruments in and beyond low Earth orbit; RadLab will also soon expand to include past (e.g. Shuttle and Mir) and future (e.g. Artemis) data. RadLab will provide a comprehensive, dynamic compendium of space radiation data, enabling the scientific community to perform analyses of data from multiple detectors and to determine the radiation environment of research missions and experiments, both via programmatic retrieval of these data and through the graphical analysis toolkit. The RadLab Working Group has been formed to foster collaborations among data contributors and users, to identify data sources, to put in place standards for data harmonization, and to guide the development of the platform, with the goal to establish the use of RadLab in space radiation research and to advance our understanding of the space radiation environment in human habitats.

radiation

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, whole organism, behavior; tabular, imagery). Open Science is the concept that the more people have access to scientifically curated data, the more knowledge will be gained. This led NASA to start the development of GeneLab in 2015. GeneLab houses spaceflight and space-analog multi-omics datasets from plant, rodent, small animal, and microbial experiments. The success and knowledge gained from GeneLab led to a new alliance of NASA “Open Science Data Repositories” (OSDR), which include the Ames Life Sciences Data Archive (ALSDA) and the NASA Biological Institutional Scientific Collection (NBISC). Both are adopting the GeneLab data system, so data are more findable, accessible, interoperable, and reusable (FAIR). OSDR systems provide users the ability to upload, download, search, share, analyze, and visualize. Open Science also needs strong confidence in the data, which is gained through building science communities. With ~400 current members, GeneLab and ALSDA formed Analysis Working Groups (AWGs) to provide feedback on processing pipelines, metadata curation standards (for ‘omics and phenotypic-physiological-behavioral assays), and to collaborate in effectively reusing data. The AWG also led to the development of the Radiation Biology Ontology (RBO), ensuring radiation metadata are efficiently captured, connected, and interoperable. Feedback from the AWG provided design input toward the new single point-of-entry data submission portal for all investigators to submit, curate, and share their research data. Space biological data is now maximally open access, collected-curated with rich metadata, and formatted for interoperability to enable systems biology, meta-analysis, knowledge graphs, machine learning, modeling, and other reuse approaches. With potential for further federation of OSDR for data mining with traditional biological and medical databases (NIH, NCI, EBI, etc.), a new era for space biology has begun to support the knowledge discovery necessary for Lunar and Martian missions.

Ryan T Scott

Maximizing Spaceflight Biological Data with Omics Analytics: The NASA GeneLab Database

NASA’s GeneLab includes an open-access repository of some 250+ omics datasets generated by biological experiments relevant to spaceflight including simulated cosmic radiation and microgravity. In order to maximize the intelligibility of these data, particularly for users with limited bioinformatics background, GeneLab has become a knowledgebase platform converting raw genetic and proteomic signatures found in flight samples into biological and physiological meanings. A large community of more than 100 scientists has rallied behind GeneLab and organized into four Analysis Working Groups (AWGs: Animal, Plant, Microbe, and Multi-Omics). Together, the AWGs have gained scientific recognition worldwide by establishing a consortium in charge of adopting new complex standards for data analysis workflows and omics sample processing in a rapidly evolving field. We will demonstrate the usage of the repository with smart search capability, an online controlled-access toolshed "Galaxy" to process user data with vetted standard workflows, a workspace for data sharing and a data submission portal with ontology control for better metadata curation. The GeneLab visualization portal will also be demonstrated, showing how anyone without formal training in bioinformatics can now browse the space biology omics data to discover new biology and potential solutions to improve life in space.

Sylvain Vincent Costes

Chemical classification program synthesis using generative artificial intelligence

Accurately classifying chemical structures is essential for cheminformatics and bioinformatics, including tasks such as identifying bioactive compounds of interest, screening molecules for toxicity to humans, finding non-organic compounds with desirable material properties, or organizing large chemical libraries for drug discovery or environmental monitoring. However, manual classification is labor-intensive and difficult to scale to large chemical databases. Existing automated approaches either rely on manually constructed classification rules, or are deep learning methods that lack explainability. This work presents an approach that uses generative artificial intelligence to automatically write chemical classifier programs for classes in the Chemical Entities of Biological Interest (ChEBI) database. These programs can be used for efficient deterministic run-time classification of SMILES structures, with natural language explanations. The programs themselves constitute an explainable computable ontological model of chemical class nomenclature, which we call the ChEBI Chemical Class Program Ontology (C3PO). We validated our approach against the ChEBI database, and compared our results against deep learning models and a naive SMARTS pattern based classifier. C3PO outperforms the naive classifier, but does not reach the performance of state of the art deep learning methods. However, C3PO has a number of strengths that complement deep learning methods, including explainability and reduced data dependence. C3PO can be used alongside deep learning classifiers to provide an explanation of the classification, where both methods agree. The programs can be used as part of the ontology development process, and iteratively refined by expert human curators.

Artificial Intelligence

Decades of Data: Extracting Trends from Microgravity Crystallization History

The reduced acceleration environment of an orbiting spacecraft has been posited as an ideal environment for biological crystal growth since buoyancy driven convection and sedimentation are greatly reduced. Since the first sounding rocket flight in 1981 many crystallization experiments have flown with some showing improvement and others not. To further explore macromolecule crystal improvement in microgravity we have accumulated data from published reports and reports submitted by individual investigators to NASA, forming a database called BIOSEArCH (Biological Space Experiment Archive of Crystallization History). To date it contains information from 63 missions including, the Space Shuttle program, unmanned satellites, the Russian Space Station MIR and sounding rocket experiments, containing reports for more than 736 macromolecule experiments. While it is not at this point in time a comprehensive record of all flight crystallization experimental results, there is however sufficient information for emerging trends to be identified. These trends will be highlighted.

Judge, Russell A.

The growing world of expansins

Expansins are cell wall proteins that induce pH-dependent wall extension and stress relaxation in a characteristic and unique manner. Two families of expansins are known, named alpha- and beta-expansins, and they comprise large multigene families whose members show diverse organ-, tissue- and cell-specific expression patterns. Other genes that bear distant sequence similarity to expansins are also represented in the sequence databases, but their biological and biochemical functions have not yet been uncovered. Expansin appears to weaken glucan-glucan binding, but its detailed mechanism of action is not well established. The biological roles of expansins are diverse, but can be related to the action of expansins to loosen cell walls, for example during cell enlargement, fruit softening, pollen tube and root hair growth, and abscission. Expansin-like proteins have also been identified in bacteria and fungi, where they may aid microbial invasion of the plant body.

NASA Discipline Plant Biology

Genelab: Scientific Partnerships and an Open-Access Database to Maximize Usage of Omics Data from Space Biology Experiments

NASA's mission includes expanding our understanding of biological systems to improve life on Earth and to enable long-duration human exploration of space. The GeneLab Data System (GLDS) is NASA's premier open-access omics data platform for biological experiments. GLDS houses standards-compliant, high-throughput sequencing and other omics data from spaceflight-relevant experiments. The GeneLab project at NASA-Ames Research Center is developing the database, and also partnering with spaceflight projects through sharing or augmentation of experiment samples to expand omics analyses on precious spaceflight samples. The partnerships ensure that the maximum amount of data is garnered from spaceflight experiments and made publically available as rapidly as possible via the GLDS. GLDS Version 1.0, went online in April 2015. Software updates and new data releases occur at least quarterly. As of October 2016, the GLDS contains 80 datasets and has search and download capabilities. Version 2.0 is slated for release in September of 2017 and will have expanded, integrated search capabilities leveraging other public omics databases (NCBI GEO, PRIDE, MG-RAST). Future versions in this multi-phase project will provide a collaborative platform for omics data analysis. Data from experiments that explore the biological effects of the spaceflight environment on a wide variety of model organisms are housed in the GLDS including data from rodents, invertebrates, plants and microbes. Human datasets are currently limited to those with anonymized data (e.g., from cultured cell lines). GeneLab ensures prompt release and open access to high-throughput genomics, transcriptomics, proteomics, and metabolomics data from spaceflight and ground-based simulations of microgravity, radiation or other space environment factors. The data are meticulously curated to assure that accurate experimental and sample processing metadata are included with each data set. GLDS download volumes indicate strong interest of the scientific community in these data. To date GeneLab has partnered with multiple experiments including two plant (Arabidopsis thaliana) experiments, two mice experiments, and several microbe experiments. GeneLab optimized protocols in the rodent partnerships for maximum yield of RNA, DNA and protein from tissues harvested and preserved during the SpaceX-4 mission, as well as from tissues from mice that were frozen intact during spaceflight and later dissected on the ground. Analysis of GeneLab data will contribute fundamental knowledge of how the space environment affects biological systems, and as well as yield terrestrial benefits resulting from mitigation strategies to prevent effects observed during exposure to space environments.

bioinformatics

GeneLab: Scientific Partnerships and an Open-Access Database to Maximize Usage of Omics Data from Space Biology Experiments

NASA's mission includes expanding our understanding of biological systems to improve life on Earth and to enable long-duration human exploration of space. The GeneLab Data System (GLDS) is NASAs premier open-access omics data platform for biological experiments. GLDS houses standards-compliant, high-throughput sequencing and other omics data from spaceflight-relevant experiments. The GeneLab project at NASA-Ames Research Center is developing the database, and also partnering with spaceflight projects through sharing or augmentation of experiment samples to expand omics analyses on precious spaceflight samples. The partnerships ensure that the maximum amount of data is garnered from spaceflight experiments and made publically available as rapidly as possible via the GLDS. GLDS Version 1.0, went online in April 2015. Software updates and new data releases occur at least quarterly. As of October 2016, the GLDS contains 80 datasets and has search and download capabilities. Version 2.0 is slated for release in September of 2017 and will have expanded, integrated search capabilities leveraging other public omics databases (NCBI GEO, PRIDE, MG-RAST). Future versions in this multi-phase project will provide a collaborative platform for omics data analysis. Data from experiments that explore the biological effects of the spaceflight environment on a wide variety of model organisms are housed in the GLDS including data from rodents, invertebrates, plants and microbes. Human datasets are currently limited to those with anonymized data (e.g., from cultured cell lines). GeneLab ensures prompt release and open access to high-throughput genomics, transcriptomics, proteomics, and metabolomics data from spaceflight and ground-based simulations of microgravity, radiation or other space environment factors. The data are meticulously curated to assure that accurate experimental and sample processing metadata are included with each data set. GLDS download volumes indicate strong interest of the scientific community in these data. To date GeneLab has partnered with multiple experiments including two plant (Arabidopsis thaliana) experiments, two mice experiments, and several microbe experiments. GeneLab optimized protocols in the rodent partnerships for maximum yield of RNA, DNA and protein from tissues harvested and preserved during the SpaceX-4 mission, as well as from tissues from mice that were frozen intact during spaceflight and later dissected on the ground. Analysis of GeneLab data will contribute fundamental knowledge of how the space environment affects biological systems, and as well as yield terrestrial benefits resulting from mitigation strategies to prevent effects observed during exposure to space environments.

spaceflight

An Open-Science Approach to Address Individual Response to Simulated GCR In Genetically Diverse Populations of Mice and Humans

This project addresses the challenge of understanding and predicting individual radiation sensitivity by integrating genetics, demographics and biomarker characteristics across species (mice and humans). We hypothesize that ex vivo DNA repair response to GCR components is a central determinant of cancer risk from space radiation and can serve as a biomarker of radiation risk in combination with genetics. Automated image quantification of 53BP1+ radiation-induced foci (RIF) during the first 4-48 h post-irradiation was performed as a function of dose and LET in non-immortalized primary skin fibroblasts derived from 76 mice across 15 strains (5 inbred reference strains and 10 collaborative-cross strains) exposed to X rays (0.1, 1 and 4 Gy), 350 MeV/n 40Ar and 600 MeV/n 56Fe (1.1 and 3 particles/100sq. μm), as well as in peripheral blood mononuclear cells (PBMCs) from 768 healthy donors (matched ethnicity, 50/50 male/female, 18-70 years old) exposed to gamma rays (0.1 and 1 Gy), 350 MeV/n 28Si, 350 MeV/n 40Ar and 600 MeV/n 56Fe (1.1 and 3 particles/100sq. μm). A genome-wide association study (GWAS) was performed on the mouse strains between DNA damage responses to space radiation and single nucleotide polymorphisms (SNPs). We found SNPs, which were significantly associated to the RIF phenotype, mapped to genes and pathways that are functionally linked to health hazards for deep space exploration (e.g. carcinogenesis, nervous system damage and immune dysfunction). Some of these SNPs were located within protein coding regions, potentially interfering with protein functions and providing promising genetic targets for countermeasures. We also found correlations between both spontaneous and radiation-induced DNA damage and SNPs mapped to pathways associated with cellular metabolism. GWAS is undergoing for the human data. All data have been made available via the NASA Space Biology Open-Science database (genelab.nasa.gov) and we will discuss how various genomic and transcriptomic datasets can be accessed for modeling and integrated using machine learning methods for discovering new radiation biology.

Sylvain V Costes