Search NASA⌕ Search

SEARCH · Search NASA

Results for “Bioinformatics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

mzPeak: Designing a Scalable, Interoperable, and Future-Ready Mass Spectrometry Data Format

Advances in mass spectrometry (MS) instrumentation, such as higher resolution, faster scan speeds, and improved sensitivity, have significantly increased the volume and complexity of data. The growing adoption of imaging and ion mobility further amplifies these challenges across MS-based omics fields, including proteomics, metabolomics, and lipidomics. While these technologies unlock new possibilities, they also present significant challenges in data management, storage, and accessibility. Existing open formats, such as the XML-based community standards mzML and imzML, struggle to meet the demands of modern MS workflows due to their large file sizes, slow data access, and limited metadata support. Vendor-specific formats, while optimized for proprietary instruments, lack interoperability, comprehensive metadata support and long-term archival reliability. This white paper lays the groundwork for mzPeak, a next-generation community data format designed to address these challenges and support high-throughput, multi-dimensional MS workflows. By adopting a hybrid model that combines efficient binary storage for numerical data and both human and machine-readable metadata storage, mzPeak will reduce file sizes, accelerate data access, and offer a scalable, adaptable solution for evolving MS technologies. For researchers, mzPeak will enable enhanced interoperability across platforms, seamless support for complex workflows including ion mobility and MS imaging, and faster data access compared to existing community formats such as mzML. Its design will ensure data is managed in compliance with regulatory standards, essential for applications such as precision medicine and chemical safety, where long-term data integrity and accessibility are critical. For vendors, mzPeak provides a streamlined, open alternative to proprietary formats, reducing the burden of regulatory compliance while aligning with the industry's push for transparency and standardization. By offering a high-performance, interoperable solution, mzPeak positions vendors to meet customer demands for sustainable data management tools which will be able to handle emerging and future data types and workflows. mzPeak aspires to become the cornerstone of MS data management, empowering researchers, vendors, and developers to innovate and collaborate more effectively.

data formats↗

Single-cell chromatin accessibility and cis -regulatory element analyses in plants using the scPlantReg platform

Understanding gene regulation is fundamental to plant improvement, but the lack of plant-specific single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) frameworks and cross-species databases has limited insights into cell-type-specific cellular regulation. Here we present ‘scPlantReg’, an integrated framework and database for plant scATAC-seq data. scPlantReg supports end-to-end analyses from raw data processing to biological interpretation and features ‘scATACtor’, a supervised machine-learning approach that outperforms existing tools for cell-type annotation. We applied scPlantReg to pearl millet to characterize cell-type-specific chromatin accessibility and identify validated activating and repressing accessible chromatin regions (ACRs), revealing WRKY transcription factors as potential regulators of xylem development. Furthermore, we reanalysed scATAC-seq datasets from 8 plant species, spanning 11 tissues and multiple developmental stages, enabling cross-species comparisons. Furthermore, these analyses uncovered conserved regulatory programmes, including AP2/EREBP-associated ACRs linked to cell wall development and cell-type-conserved TFs across grasses. Collectively, scPlantReg provides a general framework and resource for comparative regulatory analysis in plants.

Epigenomics↗

Q -score as a reliability measure for protein, nucleic acid and small-molecule atomic coordinate models derived from 3DEM maps

Atomic coordinate models are important for the interpretation of 3D maps produced with cryoEM and cryoET (3D electron microscopy; 3DEM). In addition to visual inspection of such maps and models, quantitative metrics can inform about the reliability of the atomic coordinates, in particular how well the model is supported by the experimentally determined 3DEM map. A recently introduced metric, Q-score, was shown to correlate well with the reported resolution of the map for well fitted models. Here, we present new statistical analyses of Q-score based on its application to ∼10 000 maps and models archived in the EMDB (Electron Microscopy Data Bank) and PDB (Protein Data Bank). Further, we introduce two new metrics based on Q-score to represent each map and model relative to all entries in the EMDB and those with similar resolution. We explore through illustrative examples of proteins, nucleic acids and small molecules how Q-scores can indicate whether the atomic coordinates are well fitted to 3DEM maps and also whether some parts of a map may be poorly resolved due to factors such as molecular flexibility, radiation damage and/or conformational heterogeneity. These examples and statistical analyses provide a basis for how Q-scores can be interpreted effectively in order to evaluate 3DEM maps and atomic coordinate models prior to publication and archiving.

B factors↗

The microbiologist's guide to metaproteomics

Metaproteomics is an emerging approach for studying microbiomes, offering the ability to characterize proteins that underpin microbial functionality within diverse ecosystems. As the primary catalytic and structural components of microbiomes, proteins provide unique insights into the active processes and ecological roles of microbial communities. By integrating metaproteomics with other omics disciplines, researchers can gain a comprehensive understanding of microbial ecology, interactions, and functional dynamics. This review, developed by the Metaproteomics Initiative (www.metaproteomics.org), serves as a practical guide for both microbiome and proteomics researchers, presenting key principles, state-of-the-art methodologies, and analytical workflows essential to metaproteomics. Topics covered include experimental design, sample preparation, mass spectrometry techniques, data analysis strategies, and statistical approaches.

bioinformatics↗

Modeling Protein–Protein and Protein–Ligand Interactions by the ClusPro Team in CASP16

ABSTRACT In the CASP16 experiment, our team employed hybrid computational strategies to predict both protein–protein and protein–ligand complex structures. For protein–protein docking, we combined physics‐based sampling—using ClusPro FFT docking and molecular dynamics—with AlphaFold (AF)‐based sampling, followed by AF‐based refinement. Our method produced numerous high‐accuracy complex models, including cases where AF alone failed, underscoring the critical role of physics‐based sampling alongside deep learning‐based refinement. For protein–ligand docking, we integrated the ClusPro LigTBM template‐based approach with a machine learning‐based confidence model for rescoring. The method preserves conserved interaction fragments derived from homologous complexes, followed by local resampling using physics‐based sampling and a diffusion model. Our template‐based strategy achieved a mean lDDT‐PLI of 0.69 across 233 targets, which was highly competitive. These results demonstrate that combining physics‐based modeling with AI‐driven refinement can significantly enhance the accuracy of both protein–protein and protein–ligand structure predictions.

Ashizawa, Ryota [Department of Applied Mathematics↗

PERCEPTIVE: an R shiny $\underline{p}$ipelin$\underline{e}$ for the p$\underline{r}$edi$\underline{c}$tion of $\underline{ep}$igenetic modula$\underline{t}$ors $\underline{i}$n no$\underline{v}$el sp$\underline{e}$cies

Epigenetic processes are central to regulating gene expression, genome stability, and metabolic function across the tree of life; yet, their roles remain underexplored in microalgae, especially as new species continue to be identified and characterized. This is likely due to the cumbersome nature and species-dependent attributes of epigenetic wet-lab methodologies, which preclude the rapid identification of epigenetic modifications and modulators. However, there is high conservation of epigenetic processes from budding yeast to humans; in many cases, one may infer how behavior and function are epigenetically regulated in novel species by identifying epigenetic modulators, or the proteins responsible for conferring epigenetic modifications. Here, to this end, we have developed a graphical software package, titled PERCEPTIVE (pipeline for the prediction of epigenetic modulators in novel species). This platform solely uses the genomic sequence of an algal species, and preexisting information from other model organisms, to predict the epigenetic modulators and associated modifications in algae. Predictions are presented to the user in a graphical interface, which provides literature-based interpretation of results, enabling users to quickly understand potential epigenetic processes in their algal species of interest and plan follow-up experiments. To test PERCEPTIVE, we predicted epigenetic modulators in several feedstock candidate algae species. To validate these predictions, wet-lab studies were performed, including mass spectrometry; these results underscore the high accuracy of PERCEPTIVE predictions. Overall, PERCEPTIVE represents a powerful in silico tool for the research and manipulation of algal species, which does not require a priori knowledge of epigenetics and is accessible to a broad set of investigators.

59 BASIC BIOLOGICAL SCIENCES↗

Multiomic Network Analysis Identifies Dysregulated Neurobiological Pathways in Opioid Addiction

BACKGROUND: Opioid addiction is a worldwide public health crisis. In the United States, for example, opioids cause more drug overdose deaths than any other substance. However, opioid addiction treatments have limited efficacy, meaning that additional treatments are needed. METHODS: To help address this problem, we used network-based machine learning techniques to integrate results from genome-wide association studies of opioid use disorder and problematic prescription opioid misuse with transcriptomic, proteomic, and epigenetic data from the dorsolateral prefrontal cortex of people who died of opioid overdose and control individuals. RESULTS: Here we identified 211 highly interrelated genes identified by genome-wide association studies or dysregulation in the dorsolateral prefrontal cortex of people who died of opioid overdose that implicated the Akt, BDNF (brain-derived neurotrophic factor), and ERK (extracellular signal-regulated kinase) pathways, identifying 414 drugs targeting 48 of these opioid addiction–associated genes. Some of the identified drugs are approved to treat other substance use disorders or depression. CONCLUSIONS: Our synthesis of multiomics using a systems biology approach revealed key gene targets that could contribute to drug repurposing, genetics-informed addiction treatment, and future discovery.

60 APPLIED LIFE SCIENCES↗

Bioelectrocatalytic conversion of CO₂ to PHA bioplastics using engineered methylotrophs

The sustainable generation of biodegradable plastics represents an opportunity to capture atmospheric CO 2 while reducing plastic waste accumulation in the environment. This study implements an integrated platform for bioelectrocatalytic CO 2 conversion to medium-chain-length polyhydroxyalkanoates (mcl-PHAs). Immobilizing cobalt phthalocyanine electrocatalysts on a covalent-organic framework in a gas recirculation electrolyzer enabled CO 2 -to-methanol conversion with a carbon conversion efficiency of 98%. Integration of polymer biosynthesis pathways enabled Methylotuvimicrobium alcaliphilum 20Z R to produce ~20% mcl-PHA of the dry cell weight with a CO 2 -to-bioproducts carbon conversion efficiency of 50%. This cell line was adapted to high sodium bicarbonate media, eliminating costly intermediate separation steps while improving economic potential. Transcriptomic analysis revealed sulfate transporters and peptidoglycan biosynthesis as key pathways involved in sodium bicarbonate halotolerance. Altogether, this research presents a foundation for integrating divergent chemical and biological processes into a transformative electrobiomanufacturing platform, addressing the need for alternative pipelines for generating valuable plastics and chemicals.

CO2 utilization↗

scPlantAnnotate: an accurate and robust transformer-based model for plant cell type annotation

Accurate cell type annotation remains a major bottleneck in plant single-cell RNA sequencing (scRNA-seq), where existing tools are often adapted from animal studies and perform sub-optimally on plant data. The lack of plant-specific computational frameworks limits the construction of plant cell atlases and downstream biological discovery. We develop and evaluate scPlantAnnotate, a Transformer-based reference annotation framework tailored for plant scRNA-seq data, and benchmark it against state-of-the-art deep learning and conventional methods across multiple plant species. Species-specific scPlantAnnotate models were trained using curated datasets from Arabidopsis thaliana, Zea mays, Oryza sativa, and Glycine max. We compared scPlantAnnotate with leading baselines under both standard random-split evaluation and a more stringent leave-one-dataset-out setting, which tests robustness to completely unseen datasets and tissue types. scPlantAnnotate consistently outperforms existing approaches across all four species under random-split evaluation. In the leave-one-dataset-out setting for A. thaliana, where performance drops markedly for all methods due to strong batch effects and dataset heterogeneity, scPlantAnnotate nonetheless achieves the highest Accuracy, Macro-F1, Balanced Accuracy, and Macro-AUROC on average and ranks first on most held-out datasets. These results demonstrate improved robustness to dataset shifts, a critical yet underexplored challenge in plant scRNA-seq analysis. A freely accessible web server enables users to annotate their own datasets using pretrained models. scPlantAnnotate provides a plant-specific, Transformer-based framework for single-cell annotation that delivers state-of-the-art performance and enhanced robustness to unseen datasets. By addressing limitations of existing tools and enabling scalable reference-based annotation, scPlantAnnotate supports the development of comprehensive plant cell atlases and facilitates broader use of single-cell genomics in plant biology.

Bioinformatics↗

Protocol to detect dilution cycles in chemostat experiments and estimate growth rate slopes with linear modeling with R software chemostat_regression

Chemostat growth chambers measure optical density over time and require manual calculation of growth rates. Here, we present chemostat_regression, R software that enables users to automatically identify chemostat cycles and estimate growth rate using a linear regression approach. We describe steps for creating requisite software environment(s), formatting input data, executing the software via command line/RStudio/R-Shiny, interpreting results, assessing the validity of results, and modifying input parameters.

59 BASIC BIOLOGICAL SCIENCES↗

Discovery, characterization, and application of chromosomal integration sites for stable heterologous gene expression in Rhodotorula toruloides

Rhodotorula toruloides is a non-model, oleaginous yeast uniquely suited to produce acetyl-CoA-derived chemicals. However, the lack of well-characterized genomic integration sites has impeded the metabolic engineering of this organism. Here we report a set of computationally predicted and experimentally validated chromosomal integration sites in R. toruloides. We first implemented an in silico platform by integrating essential gene information and transcriptomic data to identify candidate sites that meet stringent criteria. We then conducted a full experimental characterization of these sites, assessing integration efficiency, gene expression levels, impact on cell growth, and long-term expression stability. Among the identified sites, 12 exhibited integration efficiencies of 50% or higher, making them sufficient for most metabolic engineering applications. Using selected high-efficiency sites, we achieved simultaneous double and triple integrations and efficiently integrated long functional pathways (up to 14.7 kb). Additionally, we developed a new inducible marker recycling system that allows multiple rounds of integration at our characterized sites. Here, we validated this system by performing five sequential rounds of GFP integration and three sequential rounds of MaFAR integration for fatty alcohol production, demonstrating, for the first time, precise gene copy number tuning in R. toruloides. These characterized integration sites should significantly advance metabolic engineering efforts and future genetic tool development in R. toruloides.

59 BASIC BIOLOGICAL SCIENCES↗

Protein Data Bank (PDB): Fifty-three years young and having a transformative impact on science and society

This review article describes the co-evolution of structural biology as a discipline and the Protein Data Bank (PDB), established in 1971 as the first open-access data resource in biology by like-minded structural scientists. As the PDB archive grew in size and scope to encompass macromolecular crystallography, NMR spectroscopy, and cryo-electron microscopy, new technologies were developed to ingest, validate, curate, store, and distribute the information. Community engagement ensured that the needs of structural biologists (data depositors) and data consumers were met. Today, the archive houses more than 230,000 experimentally determined structures of proteins, nucleic acids, and macromolecular machines and their complexes with one another and small-molecule ligands. Aggregate costs of PDB data preservation are ~1% of the cost of structure determination. The enormous impact of PDB data on basic and applied research and education across the natural and medical sciences is presented and highlighted with illustrative examples. Enablement of de novo protein structure prediction (AlphaFold2, RoseTTAfold, OpenFold, etc.) is the most widely appreciated benefit of having a corpus of rigorously validated, expertly curated 3D biostructure data.

bioinformatics↗

Open-Source and FAIR Research Software for Proteomics

Scientific discovery relies on innovative software as much as experimental methods, especially in proteomics, where computational tools are essential for mass spectrometer setup, data analysis, and interpretation. Since the introduction of SEQUEST, proteomics software has grown into a complex ecosystem of algorithms, predictive models, and workflows, but the field faces challenges, including the increasing complexity of mass spectrometry data, limited reproducibility due to proprietary software, and difficulties integrating with other omics disciplines. Closed-source, platform-specific tools exacerbate these issues by restricting innovation, creating inefficiencies, and imposing hidden costs on the community. Open-source software (OSS), aligned with the FAIR Principles (Findable, Accessible, Interoperable, Reusable), offers a solution by promoting transparency, reproducibility, and community-driven development, which fosters collaboration and continuous improvement. In this manuscript, we explore the role of OSS in computational proteomics, its alignment with FAIR principles, and its potential to address challenges related to licensing, distribution, and standardization. Drawing on lessons from other omics fields, we present a vision for a future where OSS and FAIR principles underpin a transparent, accessible, and innovative proteomics community.

97 MATHEMATICS AND COMPUTING↗

Applying the FAIR Principles to computational workflows

Recent trends within computational and data sciences show an increasing recognition and adoption of computational workflows as tools for productivity and reproducibility that also democratize access to platforms and processing know-how. As digital objects to be shared, discovered, and reused, computational workflows benefit from the FAIR principles, which stand for Findable, Accessible, Interoperable, and Reusable. The Workflows Community Initiative’s FAIR Workflows Working Group (WCI-FW), a global and open community of researchers and developers working with computational workflows across disciplines and domains, has systematically addressed the application of both FAIR data and software principles to computational workflows. We present recommendations with commentary that reflects our discussions and justifies our choices and adaptations. These are offered to workflow users and authors, workflow management system developers, and providers of workflow services as guidelines for adoption and fodder for discussion. The FAIR recommendations for workflows that we propose in this paper will maximize their value as research assets and facilitate their adoption by the wider community.

97 MATHEMATICS AND COMPUTING↗