DOE OSTI · code-72658
EXSCLAIM!
Abstract
Due to recent improvements in image resolution and acquisition speed, materials microscopy is experiencing an explosion of published imaging data. The standard publication format, while sufficient for traditional data ingestion scenarios where a select number of images can be critically examined and curated manually, is not conducive tolarge-scale data aggregation or analysis, hindering data sharing and reuse. Most images in publications are presented as components of a larger figure with their explicit context buried in the main body or caption text, so even if aggregated, collections of images with weak or no digitized contextual labels have limited value. To solve the problem of curating labeled microscopy data from literature, we introduce the EXSCLAIM! Python toolkit for the automatic EXtraction, Separation, and Caption-based natural Language Annotation of IMages from scientific literature. The software is implemented through a three part pipeline: the JournalScraper, which searches the web and downloads figures and captions based on a user provided query, the CaptionDistributor, which separates caption text based on the subfigure each portion of the caption refers to, and the FigueSeparator, which separates figures into component subfigures and extracts other visual information. Also included is a Django user interface for exploring the resulting dataset.
Keep this discovery
Explore connections, maps & timelines
CHAN, MARIA, SPREADBURY, TREVORJOSEPH, SCHWENKER, ERICS, JIANG, WEIXIN. 2022-04-08. EXSCLAIM!. https://doi.org/10.11578/dc.20220408.1
Cite the original work for its findings. Save a collection to share your selection of sources.