Search NASASearch

SEARCH · Search NASA

Results for “bioinformatics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

BOSC 2025, the 26th Bioinformatics Open Source Conference

The 26th annual Bioinformatics Open Source Conference (BOSC 2025, open-bio.org/events/bosc-2025) brought its community-driven focus on open-source bioinformatics and open science to the 2025 conference on Intelligent Systems for Molecular Biology and the European Conference on Computational Biology (ISMB/ECCB 2025). Since its launch in 2000, BOSC has been the premier annual meeting covering open-source bioinformatics and open science. Framed by two keynote addresses and a thought-provoking panel discussion, the two-day conference included sessions dedicated to open data, analytic tools and pipelines, workflow platforms, knowledge representation, and the application of AI/ML. The first keynote talk was delivered by Christine Orengo: “Working together to develop, promote and protect our data resources: Lessons learnt developing CATH and TED.” A joint session with the Bio-Ontologies and Knowledge Representation (BOKR) track the second day of BOSC started with a keynote talk by Chris Mungall entitled “Open Knowledge Bases in the Age of Generative AI”. A closing panel on Data Sustainability, moderated by Mónica Muñoz Torres, featured panelists Scott Edmunds, Varsha Khodiyar, Tony Burdett, Nicky Mulder, and Chris Mungall. This year, the CollaborationFest collaborative work event that typically precedes or follows ISMB was incorporated as part of the main conference and organized by BOSC with help from the Function and 3D-SIG tracks.

bioinformatics

AlgaeOrtho, a bioinformatics tool for processing ortholog inference results in algae

Introduction: Microalgae constitute a prominent feedstock for producing biofuels and biochemicals by virtue of their prolific reproduction, high bioproduct accumulation, and the ability to grow in brackish and saline water. However, naturally occurring wild type algal strains are rarely optimal for industrial use; therefore, bioengineering of algae is necessary to generate superior performing strains that can address production challenges in industrial settings, particularly the bioenergy and bioproduct sectors. One of the crucial steps in this process is deciding on a bioengineering target: namely, which gene/protein to differentially express. These targets are often orthologs which are defined as genes/proteins originating from a common ancestor in divergent species. Although bioinformatics tools for the identification of protein orthologs already exist, processing the output from such tools is nontrivial, especially for a researcher with little or no bioinformatics experience. Methods: The present study introduces AlgaeOrtho, a user-friendly tool that builds upon the SonicParanoid orthology inference tool (based on an algorithm that identifies potential protein orthologs based on amino acid sequences) and the PhycoCosm database from JGI (Joint Genome Institute) to help researchers identify orthologs of their proteins of interest in multiple diverse algal species. Results: The output of this application includes a table of the putative orthologs of their protein of interest, a heatmap showing sequence similarity (%), and an unrooted tree of the putative protein orthologs. Notably, the tool would be instrumental in identifying novel bioengineering targets in different algal strains, including targets in not-fully annotated algal species, since it does not depend on existing protein annotations. We tested AlgaeOrtho using three case studies, for which orthologs of proteins relevant to bioengineering targets, were identified from diverse algal species, demonstrating its ease of use and utility for bioengineering researchers. Discussion: This tool is unique in the protein ortholog identification space as it can visualize putative orthologs, as desired by the user, across several algal species.

09 BIOMASS FUELS

Strategies for community-sourced biocuration in bioinformatics: a case study on MIBiG 4.0

Biocuration is essential to transform molecular sequence data into standardized, machine-readable resources. Such curated datasets enable comparative analysis, predictive modeling, and data integration across bioinformatics platforms. While professional biocuration is resource-intensive and usually limited to institutional settings, community-driven approaches can mobilize large-scale annotation of specialized datasets and are more resilient to disruptions in scientific funding. Here, we present a model for community-powered curation applied to the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) repository. Through a framework of workflows for metadata capture, annotation validation, and contributor coordination, the MIBiG 4.0 initiative recruited 267 scientists across 178 institutions from 33 countries, volunteering an estimated 4000 h of work. These efforts expanded the MIBiG repository by 22% and enhanced its usability in downstream molecular data analyses in comparative genomic analyses, natural product discovery, and machine learning applications. We provide strategies and actionable lessons for adopting this model, supporting the sustainability of curated bioinformatics resources central to nucleic acid research and related fields.

biocuration

VirJenDB: a FAIR (meta)data and bioinformatics platform for all viruses

High-throughput sequencing has generated an unprecedented volume of data. However, researcher-submitted data in repositories requires extensive curation and quality control for reuse. These tasks are hindered by the multiplicity of repositories, the sheer volume of the data, and the complexity of virus (meta)data curation. To address these challenges, VirJenDB offers a user-friendly platform to facilitate versioned, community-driven curation, and ontology development. Virus sequences were ingested from 16 sources, including ~200 fields of metadata or standards, covering taxonomy, sample, and host information. Up to 85 metadata fields have undergone at least one round of curation, and are linked to 15.4 million virus sequences, with 88 % from those infecting eukaryotes and the remaining infecting prokaryotes. Subsets were created, including a novel collection of 0.91 million viral operational taxonomic unit (vOTU) sequences across all viruses, while keeping the original sequences from each vOTU to facilitate downstream analyses, e.g. sequence variation. The VirJenDB web portal (https://www.virjendb.org) provides HTTPS and Application Programming Interface (API) access to the sequence datasets and metadata, offering a search engine, filtering, download, visualizations, and documentation. VirJenDB aims to connect the phage and eukaryotic virus research communities by supporting webtool integration, meta-analyses, and metadata schema extensions.

Saghaei, Shahram

A Tale from the Trenches: Applying Metamorphic and Differential Testing to Bioinformatics Software

Metamorphic and differential testing have been proposed as best practices for testing software that is difficult to test, such as for programs in scientific domains. An assumption is that these approaches can be easily customized and applied to almost any domain. However, scientific software is often data-driven, and metamorphic relations may require significant domain knowledge to develop. In addition, tools are often written for ad-hoc experimentation by the scientists and often embed many assumptions about the importance and representation of different natural phenomena. In this paper, we present our experience applying both metamorphic and differential testing to a set of four computational biology tools that predict the growth of an organism. While our original goal was to evaluate these techniques to improve our system-level testing, we encountered multiple roadblocks along the way. Although we did find faults (some confirmed by developers), we also uncovered a set of challenges, including the considerable manual effort required for (a) defining domain-specific tests, (b) validating correctness, and (c) distinguishing between issues stemming from poor data and those arising from incorrect software.

Marsh, Alexis L [Iowa State University/Ames Labora

A cost and community perspective on the barriers to microbiome data reuse

Microbiome research is becoming a mature field with a wealth of data amassed from diverse ecosystems, yet the ability to fully leverage multi-omics data for reuse remains challenging. To provide a view into researchers’ behavior and attitudes towards data reuse, we surveyed over 700 microbiome researchers to evaluate data sharing and reuse challenges. We found that many researchers are impeded by difficulties with metadata records, challenges with processing and bioinformatics, and problems with data repository submissions. We also explored the cost constraints of data reuse at each step of the data reuse process to better understand “pain points” and to provide a more quantitative perspective from sixteen active researchers. The bioinformatics and data processing step was estimated to be the most time consuming, which aligns with some of the most frequently reported challenges from the community survey. From these two approaches, we present evidence-based recommendations for how to address data sharing and reuse challenges with concrete actions for future work.

59 BASIC BIOLOGICAL SCIENCES

Beneath the surface: Unsolved questions in soil virus ecology

Soil virus ecology is an exciting but still nascent field of research in soil microbiology. While there has been a recent surge in soil virus research studies, many fundamental questions remain unanswered, and a range of technical and bioinformatic challenges need to be overcome. In this perspective article, we present a series of key questions that highlight fruitful research areas for ongoing and future efforts. These include describing the challenges involved in understanding soil viral abundance and activity, spatiotemporal dynamics, life strategy prevalence, virus-mediated biogeochemical impacts, viral protein function, host prediction, and soil RNA virus discovery. In the near term, combining approaches (e.g., cultivation-based, meta-omics, biogeochemical, experimental, and bioinformatic) will be key to assessing the ecological and biogeochemical impacts of soil viruses from the microscopic to the field and global scales. Still, we stress that results must be tempered by current methodological limitations and highlight knowledge gaps that are most pressing to fill via new methods or measurements, such as the prevalence of different viral replication strategies across soils, the fate of microbial necromass carbon after viral lysis, the frequency of virus-host encounters that do not lead to successful infections yet could be bioinformatically mistaken as infections, and the diversity and ecological impacts of RNA viruses in soil.

59 BASIC BIOLOGICAL SCIENCES

The Factors Governing Metal Dependence of an Emergent Superfamily of Bimetallic Oxygenases

Metalloenzyme superfamilies are typically defined by their protein scaffolds and active sites. Owing to the high tunability of protein structures, members of a single superfamily can catalyze diverse reactions with the same metallocofactor. Some superfamilies, such as amidohydrolase-related dinuclear oxygenases (AROs), display further versatility by utilizing multiple metallocofactors. We have shown that certain AROs catalyze monooxygenation reactions with diiron, dimanganese, and/or mixed manganese−iron cofactors, but the molecular factors governing the selection of a particular cofactor remain unknown, and the extent of this superfamily in biology is unclear. Here, we report bioinformatic analyses that expand the ARO superfamily to approximately 17,000 unique UniProt sequences, far exceeding the number of previously characterized enzymes. Through the integration of structural, spectroscopic, and thermodynamic analyses of representative proteins with a bioinformatic pipeline that identifies key secondary- and tertiary-sphere residues, we can predict in silico the metal preference for the majority of reported ARO sequences. These annotations were validated via the characterization of multiple new AROs, including ones implicated in key oxidative steps of natural product biosyntheses. This study establishes the key structure−function relationships governing metal preferences in AROs and highlights their vastly underappreciated role in myriad biological processes.

Liu, Chang [University of California, Berkeley, CA

Genomic reconstruction of Bacillus anthracis from complex environmental samples enables high-throughput identification and lineage assignment in Pakistan

Bacillus anthracis, the causative agent of anthrax, is a highly virulent zoonotic pathogen primarily affecting domesticated and wild herbivores. Human exposure to B. anthracis is primarily through contact with infected animals or contaminated animal products. In Pakistan, where livestock vaccines are largely unavailable and infected carcasses are often disposed of improperly, the risk to humans, wildlife and livestock is significant. Currently, the diagnosis of anthrax infections and outbreak tracing necessitates the isolation and culturing of B. anthracis, a process that requires BSL-3 facilities. In this study, we show that positive identification, genome reconstruction and lineage assignment can be accomplished using bioinformatic analysis of DNA extracted directly from environmental samples that would otherwise provide the starting material for isolation and culturing. This approach does not require laboratory target enrichment as is necessary for other pathogens, due in part to the extremely high bacterial load in the bloodstream in the deceased animals. Using these methods, we greatly expand the knowledge of endemic B. anthracis in Pakistan. We provide the first reference B. anthracis genomes from Pakistan since the 1970s and identify A.Br.014 Aust94 as a minor circulating sublineage alongside the dominant A.Br.047 Vollum. Future work will focus on the limits of detection and will determine if this bioinformatic method can be expanded more broadly for B. anthracis or other pathogens to replace typical culture-based methods.

A.Br.047 Vollum

Xylanolytic metabolism is regulated by coordination of transcription factors XynR and XylR in extremely thermophilic Caldicellulosiruptorales

ABSTRACT Global transcription factors (TFs) control metabolic processes in bacteria to efficiently utilize available carbon. The orderCaldicellulosiruptoraleshas drawn interest due to the ability of its members to degrade components of lignocellulosic biomass. Regulatory reconstruction ofAnaerocellum (f. Caldicellulosiruptor) besciiidentified two major global transcription factors for xylan utilization, XynR and XylR, and the corresponding putative transcription factor binding sites. Recombinant versions of XynR (LacI family) and XylR (ROK family) were subjected to fluorescence polarization (FP) and biolayer interferometry (BLI) analysis to confirm the predicted binding sites. Four XynR sites and two XylR sites were validated, accounting for 20 of 26 genes regulated by XynR and six of seven genes regulated by XylR. Bioinformatic analysis of the individual genes controlled by the two regulators showed an inter-dependent scheme for xylan conversion; the transport of xylooligosaccharides (XOS) is dependent on XylR, while enzymes responsible for hydrolysis are controlled by both regulators. For xylose catabolism by the xylose isomerase-xylulose kinase pathway, regulation is also split, with XylR controlling xylose isomerase and XynR controlling xylokinase. The XynR/XylR regulator pair withinA. besciiis conserved in all sequenced species ofCaldicellulosiruptorales, suggesting similarities in regulating linear xylan conversion. In other xylanolytic thermophiles, XylR homologs control xylan degradation, compared to just 6 out of 26 genes forA. bescii. These results show that two separate regulatory schemes (dual repression) are coordinated byA. besciito effectively regulate the hemicellulose inventory and xylan catabolism. IMPORTANCE To take full advantage of extreme thermophiles as platform metabolic engineering microorganisms, the tools for genetic manipulation must be further developed, and strategies that exploit a better understanding of metabolic regulation need to be discerned.Anaerocellum bescii, the most studied of the extremely thermophilic fermentative anaerobic bacteria that can utilize microcrystalline cellulose, can degrade microcrystalline cellulose and hemicellulose and has been metabolically engineered to convert the resulting sugars to products such as ethanol and acetone. For xylan, in particular, two major global transcription factors (TFs), XynR and XylR, play a role in sugar metabolism, although their predicted regulatory interdependence from bioinformatics analysis has not been elucidated experimentally. Here, fluorescence polarization (FP) and biolayer interferometry (BLI) were used to explore this issue to support metabolic engineering efforts aimed at improving carbohydrate processing to industrial chemicals.

Biotechnology & Applied Microbiology

Results from a multi-laboratory ocean metaproteomic intercomparison: effects of LC-MS acquisition and data analysis procedures

Metaproteomics is an increasingly popular methodology that provides information regarding the metabolic functions of specific microbial taxa and has potential for contributing to ocean ecology and biogeochemical studies. A blinded multi-laboratory intercomparison was conducted to assess comparability and reproducibility of taxonomic and functional results and their sensitivity to methodological variables. Euphotic zone samples from the Bermuda Atlantic Time-series Study (BATS) in the North Atlantic Ocean collected by in situ pumps and the autonomous underwater vehicle (AUV) Clio were distributed with a paired metagenome, and one-dimensional (1D) liquid chromatographic data-dependent acquisition mass spectrometry analysis was stipulated. Analysis of mass spectra from seven laboratories through a common bioinformatic pipeline identified a shared set of 1056 proteins from 1395 shared peptide constituents. Quantitative analyses showed good reproducibility: pairwise regressions of spectral counts between laboratories yielded R 2 values averaged 0.62±0.11, and a Sørensen similarity analysis of the top 1000 proteins revealed 70 %–80 % similarity between laboratory groups. Taxonomic and functional assignments showed good coherence between technical replicates and different laboratories. A bioinformatic intercomparison study, involving 10 laboratories using eight software packages, successfully identified thousands of peptides within the complex metaproteomic datasets, demonstrating the utility of these software tools for ocean metaproteomic research. Lessons learned and potential improvements in methods were described. Future efforts could examine reproducibility in deeper metaproteomes, examine accuracy in targeted absolute quantitation analyses, and develop standards for data output formats to improve data interoperability. Together, these results demonstrate the reproducibility of metaproteomic analyses and their suitability for microbial oceanography research, including integration into global-scale ocean surveys and ocean biogeochemical models.

59 BASIC BIOLOGICAL SCIENCES

The need for standardization and improved open (meta)data practices in metaproteomics

Metaproteomics enables functional insight into microbial communities by identifying and quantifying proteins in complex samples. Yet, heterogeneous analytical workflows and the lack of standardization across experimental and bioinformatics stages hinder reproducibility and comparability, limiting integration with other omics data. We here present a community-developed reporting checklist tailored to the specific needs of metaproteomics. We also outline current efforts to enable structured and interoperable metadata capture, drawing on standards from proteomics and microbiome research wherever possible. By promoting transparent reporting and advancing metadata practices, our recommendations aim to align metaproteomics more closely with FAIR principles and support reproducible and interoperable research practices.

Armengaud, Jean [Universite Paris-Saclay, France]

rcsb-api : Python Toolkit for Streamlining Access to RCSB Protein Data Bank APIs

The Protein Data Bank (PDB) was founded in 1971 as the first open-access digital data resource in biology to serve as the single global archive for three-dimensional (3D) macromolecular structure data. Current PDB holdings exceed 230,000 experimentally determined structures of proteins, nucleic acids, viruses, and macromolecular machines. The RCSB Protein Data Bank RCSB.org research-focused web portal facilitates search, analyses, and visualization of every PDB structure along with more than one million Computed Structure Models from AlphaFold DB and the ModelArchive. It is powered by a set of publicly available Application Programming Interfaces (APIs) that both support RCSB.org users and provide programmatic access to PDB data. Given the breadth and levels of granularity encompassed in this rich data collection, efficiently accessing the information programmatically may be challenging for new users. RCSB PDB has developed a Python software package, rcsb-api , that facilitates easy and efficient use of RCSB PDB APIs within a Python environment. This software tool is designed to streamline access to the extensive corpus of data housed within the PDB, enabling researchers to search, retrieve, and analyze 3D biostructure data seamlessly. Its use will accelerate research in structural biology, molecular biology and biochemistry, drug discovery, and bioinformatics by providing more efficient tools for data integration and analysis. The new toolkit is available on GitHub (github.com/rcsb/py-rcsb-api) and published to the public Python package repository (PyPI) to foster wider usage and support basic and applied research in fundamental biology, biomedicine, and the energy sciences.

FAIR principles

Emerging protein sequencing technologies: proteomics without mass spectrometry?

Liquid chromatography-tandem mass spectrometry (LC-MS/MS) has been a leading method for proteomics for 30 years. Advantages provided by LC-MS/MS are offset by significant disadvantages, including cost. Recently, several non-mass spectrometric methods have emerged, but little information is available about their capacity to analyze the complex mixtures routine for mass spectrometry. Areas Covered: We review recent non-mass-spectrometric methods for sequencing proteins and peptides, including those using nanopores, sequencing by degradation, reverse translation, and short-epitope mapping, with comments on bioinformatics challenges, fundamental limitations, and areas where new technologies will be more or less competitive with LC-MS/MS. In addition to conventional literature searches, instrument vendor websites, patents, webinars, and preprints were also consulted to give a more up-to-date picture. Expert Opinion: Many new technologies are promising. However, demonstrations that they outperform mass spectrometry in terms of peptides and proteins identified have not yet been published, and astute observers note important disadvantages, especially relating to the dynamic range of single-molecule measurements of complex mixtures. Still, even if the performance of emerging methods proves inferior to LC-MS/MS, their low cost could create a different kind of revolution: a dramatic increase in the number of biology laboratories engaging in new forms of proteomics research.

59 BASIC BIOLOGICAL SCIENCES

Energy metric prediction for double insertion mutants via the RoseNet deep learning framework

Studying the structural and functional implications of protein mutations is an important task in computational biology and bioinformatics. We leverage our previously proposed RoseNet neural network architecture to predict energy metrics of proteins with double amino acid insertions or deletions (InDels). We train models on previously generated benchmark datasets containing the exhaustive double InDel mutations for three proteins, as well as an additional three proteins for which ∼145k random mutants, each with two InDels, have been generated. We expand on our previous work by evaluating three additional proteins and analyzing domain features that impact the prediction capabilities of RoseNet. These features include InDels into secondary structures and the solvent accessible surface area (SASA) scores of the residues. We uncover further evidence to support that RoseNet has a higher proficiency of generalizing to unseen residue combinations than unseen insertion positions. We also observe that RoseNet produces higher-quality predictions when inserting into a β-sheet over an α-helix. Additionally, when the insertions fall in an area of high SASA, RoseNet often displays better performance than inserting into areas of low SASA.

59 BASIC BIOLOGICAL SCIENCES

Mining Thermophile Photosynthesis Genes: A Synthetic Operon Expressing Chloroflexota Species Reaction Center Genes in Rhodobacter sphaeroides

Photosynthesis is the foundation of the vast majority of life systems, and is therefore the most important bioenergetic process on earth. The greatest diversity of photosynthetic systems is found in microorganisms. However, our understanding of the biophysical and biochemical processes that transduce light into chemical energy is derived from a relatively small subset of proteins from microbes that are amenable to cultivation, in contrast to the huge number of predicted proteins that catalyze the initial photochemical reactions deposited in databases, such as from metagenomics. We describe the use of a Rhodobacter sphaeroides laboratory strain for the expression of heterologous photosynthesis genes to demonstrate the feasibility of mining this resource, focusing on hot spring Chloroflexota gene sequences. Using a synthetic operon of genes, we produced a photochemically active complex of reaction center proteins in our biological system. We also present bioinformatic analyses of anoxygenic type II reaction center sequences from metagenomic samples collected from hot (42–90 °C) springs available through the JGI IMG database, to generate a resource of diverse sequences that are potentially adapted to photosynthesis at such temperatures. These data provide a view into the natural diversity of anoxygenic photosynthesis, through a lens focused on high-temperature environments. The approach we took to express such genes can be applied for potential biotechnology purposes as well as for studies of fundamental catalytic properties of these heretofore inaccessible protein complexes.

Chloroflexota

Challenges and Opportunities in State‐of‐the‐Art Proteomics Analysis for Biomarker Development From Plasma Extracellular Vesicles

Extracellular vesicles (EVs) are membrane-bound particles secreted by cells, playing crucial roles in intercellular communication. The composition of EVs can undergo changes in response to stress and disease conditions, making them excellent biomarker candidates. However, extracting protein information from EVs can be challenging due to their low abundance in complex biofluids and copurification with contaminant proteins and particles. Techniques to enrich EVs have their strengths and limitations, without one being able to purify EVs to complete homogeneity. This can lead to compromised recovery rates and increased complexity, making data interpretation difficult. In this viewpoint article, we explore the concept that better characterization of EV composition, followed by quantification of EV proteins in complex samples, might be a more viable route for biomarker development. Mass spectrometers can provide reproducible deep coverage of the EV proteome, despite sample impurities. This paradigm shift presents opportunities to integrate advanced bioinformatics tools to refine the EV proteome landscape, identify novel biomarkers, and streamline validation processes in biomarker development. By focusing on leveraging technology rather than achieving absolute purity, this approach can transform current practices and open opportunities for robust biomarker discovery. Herein, we highlight not only such opportunities but also challenges to implement this concept.

Dakup, Panshak P. [Pacific Northwest National Labo