Search NASA⌕ Search

DOE OSTI · 2427476

AbBERT: Learning Antibody Humanness via Masked Language Modeling

Abstract

Understanding the degree of humanness of antibody sequences is critical to the therapeutic antibody development process to reduce the risk of failure modes like immunogenicity or poor manufacturability. We introduce AbBERT, a transformer-based language model trained on up to 20 million unpaired heavy/light chain sequences from the Observed Antibody Space database. We first validate AbBERT using a novel “multi-mask” scoring procedure to demonstrate high accuracy in predicting complementary determining regions—including the challenging hypervariable H3 region. We then demonstrate several uses of AbBERT at various points along the antibody design process. AbBERT enhances in silico antibody optimization via deep reinforcement learning by utilizing its learned embeddings as additional observations during optimization. Within a larger computational antibody design platform, AbBERT has been successfully applied as an additional design objective, where it displays strong correlations with computational tools predicting antibody structural stability. Finally, mutant antibody sequences that have been scored as unfavorable by AbBERT have shown corresponding low yields when expressed in cells. These use cases demonstrate the power of language modeling within computational antibody design.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Vashchenko, Denis, Nguyen, Sam, Goncalves, Andre, da Silva, Felipe Leno, Petersen, Brenden, Desautels, Thomas, Faissol, Daniel. 2022-08-04. AbBERT: Learning Antibody Humanness via Masked Language Modeling. https://doi.org/10.1101/2022.08.02.502236

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

scPlantAnnotate: an accurate and robust transformer-based model for plant cell type annotation

Accurate cell type annotation remains a major bottleneck in plant single-cell RNA sequencing (scRNA-seq), where existing tools are often adapted from animal studies and perform sub-optimally on plant data. The lack of plant-specific computational frameworks limits the construction of plant cell atlases and downstream biological discovery. We develop and evaluate scPlantAnnotate, a Transformer-based reference annotation framework tailored for plant scRNA-seq data, and benchmark it against state-of-the-art deep learning and conventional methods across multiple plant species. Species-specific scPlantAnnotate models were trained using curated datasets from Arabidopsis thaliana, Zea mays, Oryza sativa, and Glycine max. We compared scPlantAnnotate with leading baselines under both standard random-split evaluation and a more stringent leave-one-dataset-out setting, which tests robustness to completely unseen datasets and tissue types. scPlantAnnotate consistently outperforms existing approaches across all four species under random-split evaluation. In the leave-one-dataset-out setting for A. thaliana, where performance drops markedly for all methods due to strong batch effects and dataset heterogeneity, scPlantAnnotate nonetheless achieves the highest Accuracy, Macro-F1, Balanced Accuracy, and Macro-AUROC on average and ranks first on most held-out datasets. These results demonstrate improved robustness to dataset shifts, a critical yet underexplored challenge in plant scRNA-seq analysis. A freely accessible web server enables users to annotate their own datasets using pretrained models. scPlantAnnotate provides a plant-specific, Transformer-based framework for single-cell annotation that delivers state-of-the-art performance and enhanced robustness to unseen datasets. By addressing limitations of existing tools and enabling scalable reference-based annotation, scPlantAnnotate supports the development of comprehensive plant cell atlases and facilitates broader use of single-cell genomics in plant biology.

Bioinformatics↗

CRISPR-COPIES: Web Tool

CRISPR/Cas system has emerged as a powerful genome-editing tool for metabolic engineering and human gene therapy. However, the conundrum of where to integrate heterologous genes on the chromosome using the CRISPR/Cas system remains an open question. Selecting a site for gene integration requires incorporation of complex criteria such as factors involved in CRISPR/Cas-mediated integration, genetic stability, and gene expression and therefore, usually requires strenuous characterization of sites on particular or different chromosomal locations. To address these issues, we developed CRISPR-COPIES, a COmputational Pipeline for the Identification of CRISPR/Cas-facilitated intEgration Sites. The tool applies ScaNN, a state-of-the-art model on the embedding-based nearest neighbor search for fast and accurate off-target search and can identify genome-wide intergenic sites for most bacterial and fungal genomes within minutes. This submission contains the code we developed to create a user-friendly web interface for CRISPR-COPIES (https://biofoundry.web.illinois.edu/copies/). We anticipate CRISPR-COPIES will serve as a useful tool for targeted DNA integration and aid in the characterization of synthetic biology toolkits, rapid strain construction to produce valuable biochemicals, and human gene and cell therapy.

Bioinformatics↗