Search NASA⌕ Search

DOE OSTI · 2433806

Structured information extraction from scientific text with large language models

Abstract

Extracting structured knowledge from scientific text remains a challenging task for machine learning models. Here, we present a simple approach to joint named entity recognition and relation extraction and demonstrate how pretrained large language models (GPT-3, Llama-2) can be fine-tuned to extract useful records of complex scientific knowledge. We test three representative tasks in materials chemistry: linking dopants and host materials, cataloging metal-organic frameworks, and general composition/phase/morphology/application information extraction. Records are extracted from single sentences or entire paragraphs, and the output can be returned as simple English sentences or a more structured format such as a list of JSON objects. This approach represents a simple, accessible, and highly flexible route to obtaining large databases of structured specialized scientific knowledge extracted from research papers.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Dagdelen, John, Dunn, Alexander, Lee, Sanghoon, Walker, Nicholas, Rosen, Andrew S., Ceder, Gerbrand, Persson, Kristin A., Jain, Anubhav. 2024-02-15. Structured information extraction from scientific text with large language models. https://doi.org/10.1038/s41467-024-45563-x

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Ontologies for Intelligent Data Science

As anyone even vaguely aware of current technology can tell you, machine learning (ML) and artificial intelligence (AI) have made exceptional breakthroughs in recent years. Generative artificial intelligence (GAI) emerged circa 2022 dominated by Large Language Models (LLMs) and generative tools for images emerged at about the same time.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Data Cards for Standardized Metadata Across DOE-Aligned Data Initiatives: Toward Transparent, Interoperable, and Governed Dataset Documentation

As data-intensive research, advanced computing, and artificial intelligence become increasingly central to scientific and operational workflows, the need for consistent, transparent, and machine-actionable documentation has grown correspondingly. Multiple DOE-aligned communities—including Office of Science, Genesis Mission, American Science Cloud (AmSC), National Nuclear Security Administration (NNSA) stewardship and governance, and related cross-laboratory collaborations—have independently developed metadata practices to support discovery, access, reuse, repository deposit, and compliance.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Verification and Validation of Performance with Dissemination of Best Practices in District Energy and CHP for Enhanced Resiliency, Energy Efficiency, and Cybersecurity

This report contains the results of the International District Energy Association’s work to analyze, validate, and verify performance data of existing district energy systems and identify industy best practices for the purpose of improving system reliability, resiliency, and efficiency, and to accellerate decarbonization. In addition to a technical evaluation of the surveyed systems and identification of a series of technical performance metrics, the report illustrates the accompanying operations and financial best practices employed by surveyed systems to fully serve their customer base. Additionally, the third chapter of the report describes the current landscape of cybersecurity threats and counteracting measures, and recommends a series of steps for effectively guarding highly networked district energy systems against cybersecurity attacks.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗