DOE OSTI · code-164482
AdaParse
Abstract
SF-25-127AdaParse (Adaptive Parallel PDF Parsing and Resource Scaling Engine) enables scalable, high-accuracy PDF parsing. AdaParse is a data-driven strategy that assigns an appropriate parser to each document, offering high accuracy for any computational budget. Moreover, it offers a workflow of various PDF parsing software that includes extraction tools: PyMuPDF, pypdf traditional OCR: Tesseract, modern OCR (e.g., Vision Transformers): Nougat and Marker
Keep this discovery
Explore connections, maps & timelines
Siebenschuh, Carlo, Hippe, Kyle, Gokdemir, Ozan, Brace, Alexander, Khan, Arham, Hosssain, MD Khalid, Babuji, YaduNand, Chia, Nicholas, Vishwanath, Venkatram, Stevens, RickL, Foster, IanT, Underwood, Robert. 2025-09-19. AdaParse. https://doi.org/10.11578/dc.20250919.2
Cite the original work for its findings. Save a collection to share your selection of sources.