Search NASA⌕ Search

DOE OSTI · 1765039

A pre-training and self-training approach for biomedical named entity recognition

Abstract

Named entity recognition (NER) is a key component of many scientific literature mining tasks, such as information retrieval, information extraction, and question answering; however, many modern approaches require large amounts of labeled training data in order to be effective. This severely limits the effectiveness of NER models in applications where expert annotations are difficult and expensive to obtain. In this work, we explore the effectiveness of transfer learning and semi-supervised self-training to improve the performance of NER models in biomedical settings with very limited labeled data (250-2000 labeled samples). We first pre-train a BiLSTM-CRF and a BERT model on a very large general biomedical NER corpus such as MedMentions or Semantic Medline, and then we fine-tune the model on a more specific target NER task that has very limited training data; finally, we apply semi-supervised self-training using unlabeled data to further boost model performance. We show that in NER tasks that focus on common biomedical entity types such as those in the Unified Medical Language System (UMLS), combining transfer learning with self-training enables a NER model such as a BiLSTM-CRF or BERT to obtain similar performance with the same model trained on 3x-8x the amount of labeled data. We further show that our approach can also boost performance in a low-resource application where entities types are more rare and not specifically covered in UMLS.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Gao, Shang (ORCID:0000000318031457), Kotevska, Olivera, Sorokine, Alexandre (ORCID:0000000179931534), Christian, J. Blair (ORCID:0000000246581635), Fiorini, ed., Nicolas. 2021-02-09. A pre-training and self-training approach for biomedical named entity recognition. https://doi.org/10.1371/journal.pone.0246310

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Ontologies for Intelligent Data Science

As anyone even vaguely aware of current technology can tell you, machine learning (ML) and artificial intelligence (AI) have made exceptional breakthroughs in recent years. Generative artificial intelligence (GAI) emerged circa 2022 dominated by Large Language Models (LLMs) and generative tools for images emerged at about the same time.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Data Cards for Standardized Metadata Across DOE-Aligned Data Initiatives: Toward Transparent, Interoperable, and Governed Dataset Documentation

As data-intensive research, advanced computing, and artificial intelligence become increasingly central to scientific and operational workflows, the need for consistent, transparent, and machine-actionable documentation has grown correspondingly. Multiple DOE-aligned communities—including Office of Science, Genesis Mission, American Science Cloud (AmSC), National Nuclear Security Administration (NNSA) stewardship and governance, and related cross-laboratory collaborations—have independently developed metadata practices to support discovery, access, reuse, repository deposit, and compliance.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Verification and Validation of Performance with Dissemination of Best Practices in District Energy and CHP for Enhanced Resiliency, Energy Efficiency, and Cybersecurity

This report contains the results of the International District Energy Association’s work to analyze, validate, and verify performance data of existing district energy systems and identify industy best practices for the purpose of improving system reliability, resiliency, and efficiency, and to accellerate decarbonization. In addition to a technical evaluation of the surveyed systems and identification of a series of technical performance metrics, the report illustrates the accompanying operations and financial best practices employed by surveyed systems to fully serve their customer base. Additionally, the third chapter of the report describes the current landscape of cybersecurity threats and counteracting measures, and recommends a series of steps for effectively guarding highly networked district energy systems against cybersecurity attacks.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗