Search NASA⌕ Search

SEARCH · Search NASA

Results for “Integrative bioinformatics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

SyPro Poplar: Improving Poplar Biomass Production under Abiotic Stress Conditions: an Integrated Omics, Bioinformatics, Synthetic Biology and Genetic Engineering Approach (Final Report)

SyPro Poplar: Improving Poplar Biomass Production under Abiotic Stress Conditions: an Integrated Omics, Bioinformatics, Synthetic Biology and Genetic Engineering Approach In the SyPro Poplar project, we aimed to integrate omics, bioinformatics, synthetic biology, and genetic engineering approaches to develop transgenic poplar trees with sustained photosynthetic activity and increased biomass production under individual and the simultaneous occurrence of water deficit, increased soil salinity, and elevated temperatures. Specifically, we intended to (i) study the functions of selected stress-responsive genes at tissue and cell type-specific levels, (ii) discover novel motifs and construct stress-responsive synthetic promoters, and (iii) use these promoters to drive the expression of genes shown to confer abiotic stress tolerance in a variety of crops and develop abiotic stress-tolerant poplars in a coordinated fashion.

09 BIOMASS FUELS↗

GeneLab

GeneLab collects and enables analysis of spaceflight and ground-based spaceflight simulation genomic data, RNA and protein expression, and metabolic profiles. It interfaces with other existing databases containing spaceflight omic data. The 2011 National Research Council (NRC) Decadal Survey on NASA Life and Physical Sciences called for increased opportunities for multi-investigator spaceflight opportunities and greater use of genomic approaches to meet the needs of NASA researchers. To address these recommendations of the NRC Decadal Survey, the Space Life and Physical Sciences Research and Applications Division of NASA's Human Exploration and Operations Mission Directorate has initiated a transition to an Open Science architecture to increase research opportunities, and has developed the GeneLab Platform based on highly leveraged and integrated bioinformatics analytics. GeneLab is an interactive, open-access resource where scientists can upload, download, store, search, share, transfer, and analyze omics data from spaceflight and corresponding analogue experiments. Users can explore GeneLab datasets in the Data Repository, analyze data using the Analysis Platform, visualize high-order data and create collaborative projects using the Collaborative Workspace. Our primary goal is to maximize the utilization of the valuable biological research conducted aboard the International Space Station (ISS) by collecting genomic, transcriptomic, proteomic, and metabolomics data known as “omics”. By providing a portal linking processed data to flight parameters, GeneLab enables exploration of the molecular network responses of terrestrial biology to the space environment. This allows researchers to understand the complex responses of biological systems to the space environment. This technology development activity was transferred from the Human Exploration and Operations Mission Directorate to the Science Mission Directorate Division of Biological and Physical Sciences (BPS) in October 2020.

GeneLab↗

Developing a Hybrid Spacesuit Simulator as a Research Tool for Assessing Extravehicular Activity Relevant Workload

Conducting human tests in a pressurized spacesuit is limited by availability, cost, and manpower; however, pressurized spacesuits are not always needed depending on the objectives of testing, including the development and testing of new informatics capabilities. The Human Physiology, Performance, Protection & Operations Laboratory (H-3PO) at NASA is developing a Hybrid Spacesuit Simulator (HS3) to support testing and characterization of human performance during analog planetary exploration extravehicular activities (EVAs). The goal of HS3 is to create a low-cost, modular, and unpressurized spacesuit simulator as a research tool that provides relevant physical and cognitive workload approximations with EVA-like immersion. HS3 consists of a soft outer suit, thermal control, gloves, boots, helmet, and integrated bioinformatics and communications. Baseline HS3 assessments were performed during 3-hour EVA simulations in two different subjects (DEMO1 and DEMO2) that included traverses at variable resistances and geological sampling activities. Liquid cooling garment (LCG) temperature, mean skin temperature, heart rate, motion capture, and metabolic rate were collected during each 3-hour simulated EVA. During DEMO1 and DEMO2, baseline metabolic rates at rest were 836 ± 327 BTU/hr and 869 ± 207 BTU/hr and increased to 2124 ± 548 BTU/hr and 2269 ± 559 BTU/hr, respectively, during 500m traverse. Average inlet LCG temperatures were 29.57 ± 6.62 °C and 25.63 ± 6.48 °C for DEMO1 and DEMO2 with increased outlet LCG temperatures of 33.53 ± 6.62 °C and 29.21 ± 4.79 °C, respectively. Overall, HS3 will enable future studies to characterize EVA tasks, human performance, and test future EVA capabilities in analog test environments without the need for pressurized suited environments.

Monica Hew↗

Developing A Hybrid Spacesuit Simulator as A Research Tool for Assessing Extravehicular Activity Relevant Workload

Conducting human tests in a pressurized spacesuit is limited by availability, cost, and manpower; however, pressurized spacesuits are not always needed depending on the objectives of testing, including the development and testing of new informatics capabilities. The Human Physiology, Performance, Protection & Operations Laboratory (H-3PO) at NASA is developing a Hybrid Spacesuit Simulator (HS3) to support testing and characterization of human performance during analog planetary exploration extravehicular activities (EVAs). The goal of HS3 is to create a low-cost, modular, and unpressurized spacesuit simulator as a research tool that provides relevant physical and cognitive workload approximations with EVA-like immersion. HS3 consists of a soft outer suit, thermal control, gloves, boots, helmet, and integrated bioinformatics and communications. Baseline HS3 assessments were performed during 3-hour EVA simulations in two different subjects (DEMO1 and DEMO2) that included traverses at variable resistances and geological sampling activities. Liquid cooling garment (LCG) temperature, mean skin temperature, heart rate, motion capture, and metabolic rate were collected during each 3-hour simulated EVA. During DEMO1 and DEMO2, baseline metabolic rates at rest were 836 ± 327 BTU/hr and 869 ± 207 BTU/hr and increased to 2124 ± 548 BTU/hr and 2269 ± 559 BTU/hr, respectively, during 500m traverse. Average inlet LCG temperatures were 29.57 ± 6.62 °C and 25.63 ± 6.48 °C for DEMO1 and DEMO2 with increased outlet LCG temperatures of 33.53 ± 6.62 °C and 29.21 ± 4.79 °C, respectively. Overall, HS3 will enable future studies to characterize EVA tasks, human performance, and test future EVA capabilities in analog test environments without the need for pressurized suited environments.

Suit simulator↗

Strategies for community-sourced biocuration in bioinformatics: a case study on MIBiG 4.0

Biocuration is essential to transform molecular sequence data into standardized, machine-readable resources. Such curated datasets enable comparative analysis, predictive modeling, and data integration across bioinformatics platforms. While professional biocuration is resource-intensive and usually limited to institutional settings, community-driven approaches can mobilize large-scale annotation of specialized datasets and are more resilient to disruptions in scientific funding. Here, we present a model for community-powered curation applied to the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) repository. Through a framework of workflows for metadata capture, annotation validation, and contributor coordination, the MIBiG 4.0 initiative recruited 267 scientists across 178 institutions from 33 countries, volunteering an estimated 4000 h of work. These efforts expanded the MIBiG repository by 22% and enhanced its usability in downstream molecular data analyses in comparative genomic analyses, natural product discovery, and machine learning applications. We provide strategies and actionable lessons for adopting this model, supporting the sustainability of curated bioinformatics resources central to nucleic acid research and related fields.

biocuration↗

Challenges and Opportunities in State‐of‐the‐Art Proteomics Analysis for Biomarker Development From Plasma Extracellular Vesicles

Extracellular vesicles (EVs) are membrane-bound particles secreted by cells, playing crucial roles in intercellular communication. The composition of EVs can undergo changes in response to stress and disease conditions, making them excellent biomarker candidates. However, extracting protein information from EVs can be challenging due to their low abundance in complex biofluids and copurification with contaminant proteins and particles. Techniques to enrich EVs have their strengths and limitations, without one being able to purify EVs to complete homogeneity. This can lead to compromised recovery rates and increased complexity, making data interpretation difficult. In this viewpoint article, we explore the concept that better characterization of EV composition, followed by quantification of EV proteins in complex samples, might be a more viable route for biomarker development. Mass spectrometers can provide reproducible deep coverage of the EV proteome, despite sample impurities. This paradigm shift presents opportunities to integrate advanced bioinformatics tools to refine the EV proteome landscape, identify novel biomarkers, and streamline validation processes in biomarker development. By focusing on leveraging technology rather than achieving absolute purity, this approach can transform current practices and open opportunities for robust biomarker discovery. Herein, we highlight not only such opportunities but also challenges to implement this concept.

Dakup, Panshak P. [Pacific Northwest National Labo↗

The endohyphal microbiome: current progress and challenges for scaling down integrative multi-omic microbiome research

Abstract As microbiome research has progressed, it has become clear that most, if not all, eukaryotic organisms are hosts to microbiomes composed of prokaryotes, other eukaryotes, and viruses. Fungi have only recently been considered holobionts with their own microbiomes, as filamentous fungi have been found to harbor bacteria (including cyanobacteria), mycoviruses, other fungi, and whole algal cells within their hyphae. Constituents of this complex endohyphal microbiome have been interrogated using multi-omic approaches. However, a lack of tools, techniques, and standardization for integrative multi-omics for small-scale microbiomes (e.g., intracellular microbiomes) has limited progress towards investigating and understanding the total diversity of the endohyphal microbiome and its functional impacts on fungal hosts. Understanding microbiome impacts on fungal hosts will advance explorations of how “microbiomes within microbiomes” affect broader microbial community dynamics and ecological functions. Progress to date as well as ongoing challenges of performing integrative multi-omics on the endohyphal microbiome is discussed herein. Addressing the challenges associated with the sample extraction, sample preparation, multi-omic data generation, and multi-omic data analysis and integration will help advance current knowledge of the endohyphal microbiome and provide a road map for shrinking microbiome investigations to smaller scales.

59 BASIC BIOLOGICAL SCIENCES↗

Visualizing and analyzing 3D biomolecular structures using Mol* at RCSB.org: Influenza A H5N1 virus proteome case study

The easiest and often most useful way to work with experimentally determined or computationally predicted structures of biomolecules is by viewing their three-dimensional (3D) shapes using a molecular visualization tool. Mol* was collaboratively developed by RCSB Protein Data Bank (RCSB PDB, RCSB.org) and Protein Data Bank in Europe (PDBe, PDBe.org) as an open-source, web-based, 3D visualization software suite for examination and analyses of biostructures. It is capable of displaying atomic coordinates and related experimental data of biomolecular structures together with a variety of annotations, facilitating basic and applied research, training, education, and information dissemination. Across RCSB.org, the RCSB PDB research-focused web portal, Mol* has been implemented to support single-mouse-click atomic-level visualization of biomolecules (e.g., proteins, nucleic acids, carbohydrates) with bound cofactors, small-molecule ligands, ions, water molecules, or other macromolecules. RCSB.org Mol* can seamlessly display 3D structures from various sources, allowing structure interrogation, superimposition, and comparison. Using influenza A H5N1 virus as a topical case study of an important pathogen, we exemplify how Mol* has been embedded within various RCSB.org tools—allowing users to view polymer sequence and structure-based annotations integrated from trusted bioinformatics data resources, assess patterns and trends in groups of structures, and view structures of any size and compositional complexity. In addition to being linked to every experimentally determined biostructure and Computed Structure Model made available at RCSB.org, Standalone Mol* is freely available for visualizing any atomic-level or multi-scale biostructure at rcsb.org/3d-view.

3D biostructure↗

Automating methods for estimating metabolite volatility

The volatility of metabolites can influence their biological roles and inform optimal methods for their detection. Yet, volatility information is not readily available for the large number of described metabolites, limiting the exploration of volatility as a fundamental trait of metabolites. Here, we adapted methods to estimate vapor pressure from the functional group composition of individual molecules (SIMPOL.1) to predict the gas-phase partitioning of compounds in different environments. We implemented these methods in a new open pipeline called volcalc that uses chemoinformatic tools to automate these volatility estimates for all metabolites in an extensive and continuously updated pathway database: the Kyoto Encyclopedia of Genes and Genomes (KEGG) that connects metabolites, organisms, and reactions. We first benchmark the automated pipeline against a manually curated data set and show that the same category of volatility (e.g., nonvolatile, low, moderate, high) is predicted for 93% of compounds. We then demonstrate how volcalc might be used to generate and test hypotheses about the role of volatility in biological systems and organisms. Specifically, we estimate that 3.4 and 26.6% of compounds in KEGG have high volatility depending on the environment (soil vs. clean atmosphere, respectively) and that a core set of volatiles is shared among all domains of life (30%) with the largest proportion of kingdom-specific volatiles identified in bacteria. With volcalc , we lay a foundation for uncovering the role of the volatilome using an approach that is easily integrated with other bioinformatic pipelines and can be continually refined to consider additional dimensions to volatility. The volcalc package is an accessible tool to help design and test hypotheses on volatile metabolites and their unique roles in biological systems.

59 BASIC BIOLOGICAL SCIENCES↗

Inference of cell type-specific gene regulatory networks on cell lineages from single cell omic datasets

Abstract Cell type-specific gene expression patterns are outputs of transcriptional gene regulatory networks (GRNs) that connect transcription factors and signaling proteins to target genes. Single-cell technologies such as single cell RNA-sequencing (scRNA-seq) and single cell Assay for Transposase-Accessible Chromatin using sequencing (scATAC-seq), can examine cell-type specific gene regulation at unprecedented detail. However, current approaches to infer cell type-specific GRNs are limited in their ability to integrate scRNA-seq and scATAC-seq measurements and to model network dynamics on a cell lineage. To address this challenge, we have developed single-cell Multi-Task Network Inference (scMTNI), a multi-task learning framework to infer the GRN for each cell type on a lineage from scRNA-seq and scATAC-seq data. Using simulated and real datasets, we show that scMTNI is a broadly applicable framework for linear and branching lineages that accurately infers GRN dynamics and identifies key regulators of fate transitions for diverse processes such as cellular reprogramming and differentiation.

59 BASIC BIOLOGICAL SCIENCES↗

VLSI Microsystem for Rapid Bioinformatic Pattern Recognition

A system comprising very-large-scale integrated (VLSI) circuits is being developed as a means of bioinformatics-oriented analysis and recognition of patterns of fluorescence generated in a microarray in an advanced, highly miniaturized, portable genetic-expression-assay instrument. Such an instrument implements an on-chip combination of polymerase chain reactions and electrochemical transduction for amplification and detection of deoxyribonucleic acid (DNA).

Fang, Wai-Chi↗

ZMPY3D: accelerating protein structure volume analysis through vectorized 3D Zernike moments and Python-based GPU integration

Abstract Motivation Volumetric 3D object analyses are being applied in research fields such as structural bioinformatics, biophysics, and structural biology, with potential integration of artificial intelligence/machine learning (AI/ML) techniques. One such method, 3D Zernike moments, has proven valuable in analyzing protein structures (e.g., protein fold classification, protein–protein interaction analysis, and molecular dynamics simulations). Their compactness and efficiency make them amenable to large-scale analyses. Established methods for deriving 3D Zernike moments, however, can be inefficient, particularly when higher order terms are required, hindering broader applications. As the volume of experimental and computationally-predicted protein structure information continues to increase, structural biology has become a “big data” science requiring more efficient analysis tools. Results This application note presents a Python-based software package, ZMPY3D, to accelerate computation of 3D Zernike moments by vectorizing the mathematical formulae and using graphical processing units (GPUs). The package offers popular GPU-supported libraries such as CuPy and TensorFlow together with NumPy implementations, aiming to improve computational efficiency, adaptability, and flexibility in future algorithm development. The ZMPY3D package can be installed via PyPI, and the source code is available from GitHub. Volumetric-based protein 3D structural similarity scores and transform matrix of superposition functionalities have both been implemented, creating a powerful computational tool that will allow the research community to amalgamate 3D Zernike moments with existing AI/ML tools, to advance research and education in protein structure bioinformatics. Availability and implementation ZMPY3D, implemented in Python, is available on GitHub (https://github.com/tawssie/ZMPY3D) and PyPI, released under the GPL License.

Lai, Jhih-Siang (ORCID:0000000156775890)↗

PYK-SubstitutionOME: an integrated database containing allosteric coupling, ligand affinity and mutational, structural, pathological, bioinformatic and computational information about pyruvate kinase isozymes

Interpreting changes in patient genomes, understanding how viruses evolve and engineering novel protein function all depend on accurately predicting the functional outcomes that arise from amino acid substitutions. To that end, the development of first-generation prediction algorithms was guided by historic experimental datasets. However, these datasets were heavily biased toward substitutions at positions that have not changed much throughout evolution (i.e. conserved). Although newer datasets include substitutions at positions that span a range of evolutionary conservation scores, these data are largely derived from assays that agglomerate multiple aspects of function. To facilitate predictions from the foundational chemical properties of proteins, large substitution databases with biochemical characterizations of function are needed. We report here a database derived from mutational, biochemical, bioinformatic, structural, pathological and computational studies of a highly studied protein family—pyruvate kinase (PYK). A centerpiece of this database is the biochemical characterization—including quantitative evaluation of allosteric regulation—of the changes that accompany substitutions at positions that sample the full conservation range observed in the PYK family. We have used these data to facilitate critical advances in the foundational studies of allosteric regulation and protein evolution and as rigorous benchmarks for testing protein predictions. We trust that the collected dataset will be useful for the broader scientific community in the further development of prediction algorithms.

59 BASIC BIOLOGICAL SCIENCES↗

Cohort-based pan-cancer analysis and experimental studies reveal ISG15 gene as a novel biomarker for prognosis and immunotherapy efficacy prediction

Abstract ISG15, an interferon-stimulated ubiquitin-like protein, plays a multifaceted role in tumorigenesis and immune regulation. This study comprehensively evaluates ISG15 as a prognostic biomarker and predictor of immunotherapy response through pan-cancer bioinformatics analysis and experimental validation. By integrating multiomics data from TCGA, GEO, and clinical cohorts, we found that ISG15 is significantly overexpressed in multiple cancers and generally correlates with poor prognosis. Elevated ISG15 expression is associated with increased immune checkpoint gene expression, particularly PD-L1, and immune infiltration, notably M2-like tumor-associated macrophages. Immunohistochemistry and multiplexed immunofluorescence confirmed a strong positive correlation between ISG15, PD-L1, and M2-TAM infiltration in lung and gastric cancer samples. Functional analysis at the single-cell level revealed significant associations between ISG15 and tumor proliferation, angiogenesis, and immune suppression. Immunotherapy cohort analysis demonstrated that tumors with high ISG15 expression responded favorably to PD-L1 inhibitors but exhibited resistance to CTLA-4 blockade, findings further validated in lung cancer patients receiving anti-PD-1 therapy. These results suggest that ISG15 is a promising biomarker for prognosis and immunotherapy response prediction across cancers. Its integration into clinical decision-making may enhance personalized treatment strategies, improve immunotherapy outcomes, and provide new insights into the tumor immune microenvironment, cancer progression, and potential therapeutic targets for future drug development.

Immunology↗

The need for standardization and improved open (meta)data practices in metaproteomics

Metaproteomics enables functional insight into microbial communities by identifying and quantifying proteins in complex samples. Yet, heterogeneous analytical workflows and the lack of standardization across experimental and bioinformatics stages hinder reproducibility and comparability, limiting integration with other omics data. We here present a community-developed reporting checklist tailored to the specific needs of metaproteomics. We also outline current efforts to enable structured and interoperable metadata capture, drawing on standards from proteomics and microbiome research wherever possible. By promoting transparent reporting and advancing metadata practices, our recommendations aim to align metaproteomics more closely with FAIR principles and support reproducible and interoperable research practices.

Armengaud, Jean [Universite Paris-Saclay, France]↗

rcsb-api : Python Toolkit for Streamlining Access to RCSB Protein Data Bank APIs

The Protein Data Bank (PDB) was founded in 1971 as the first open-access digital data resource in biology to serve as the single global archive for three-dimensional (3D) macromolecular structure data. Current PDB holdings exceed 230,000 experimentally determined structures of proteins, nucleic acids, viruses, and macromolecular machines. The RCSB Protein Data Bank RCSB.org research-focused web portal facilitates search, analyses, and visualization of every PDB structure along with more than one million Computed Structure Models from AlphaFold DB and the ModelArchive. It is powered by a set of publicly available Application Programming Interfaces (APIs) that both support RCSB.org users and provide programmatic access to PDB data. Given the breadth and levels of granularity encompassed in this rich data collection, efficiently accessing the information programmatically may be challenging for new users. RCSB PDB has developed a Python software package, rcsb-api , that facilitates easy and efficient use of RCSB PDB APIs within a Python environment. This software tool is designed to streamline access to the extensive corpus of data housed within the PDB, enabling researchers to search, retrieve, and analyze 3D biostructure data seamlessly. Its use will accelerate research in structural biology, molecular biology and biochemistry, drug discovery, and bioinformatics by providing more efficient tools for data integration and analysis. The new toolkit is available on GitHub (github.com/rcsb/py-rcsb-api) and published to the public Python package repository (PyPI) to foster wider usage and support basic and applied research in fundamental biology, biomedicine, and the energy sciences.

FAIR principles↗

The Factors Governing Metal Dependence of an Emergent Superfamily of Bimetallic Oxygenases

Metalloenzyme superfamilies are typically defined by their protein scaffolds and active sites. Owing to the high tunability of protein structures, members of a single superfamily can catalyze diverse reactions with the same metallocofactor. Some superfamilies, such as amidohydrolase-related dinuclear oxygenases (AROs), display further versatility by utilizing multiple metallocofactors. We have shown that certain AROs catalyze monooxygenation reactions with diiron, dimanganese, and/or mixed manganese−iron cofactors, but the molecular factors governing the selection of a particular cofactor remain unknown, and the extent of this superfamily in biology is unclear. Here, we report bioinformatic analyses that expand the ARO superfamily to approximately 17,000 unique UniProt sequences, far exceeding the number of previously characterized enzymes. Through the integration of structural, spectroscopic, and thermodynamic analyses of representative proteins with a bioinformatic pipeline that identifies key secondary- and tertiary-sphere residues, we can predict in silico the metal preference for the majority of reported ARO sequences. These annotations were validated via the characterization of multiple new AROs, including ones implicated in key oxidative steps of natural product biosyntheses. This study establishes the key structure−function relationships governing metal preferences in AROs and highlights their vastly underappreciated role in myriad biological processes.

Liu, Chang [University of California, Berkeley, CA↗

2024 International Conference on Microbiome Engineering (ICME)

The 2024 International Conference on Microbiome Engineering (ICME) took place November 12-14 at Tufts University in Medford, MA. ICME connects experts from academia and industry to share the most recent developments in the field of microbiome engineering. This includes genetically engineered organisms that function within microbiomes, control of microbiomes through environmental/nutrient modifications, and inference of engineering principles from analysis of synthetic and natural microbiomes. The conference is unique and distinct from other microbiome conferences in that it specifically highlights the integration of engineering design principles with microbiome research (others are more focused on basic biological principles). The conference thus integrates synthetic biology, systems biology, microbial ecology, and bioinformatics across a range of application spaces from the environment to manufacturing, food, and human health. This project utilized support from the Department of Energy’s (DOE) Office of Biological and Environmental Research (BER) to help trainees and early career faculty attend ICME.

60 APPLIED LIFE SCIENCES↗