3 papers
cs.CV2026
POTATR: A Lightweight Image-to-Graph Model for Page-Level Table Extraction
Brandon Smock, Libin Liang, Max Sokolov +4
Large-scale document processing requires contextually aware table extraction (TE) that is both accurate and efficient. Yet current approaches require billions of parameters, hundre…
cs.CV2025
PubTables-v2: A new large-scale dataset for full-page and multi-page table extraction
Brandon Smock, Valerie Faucon-Morin, Max Sokolov +4
Table extraction (TE) is a key challenge in document understanding. Traditional approaches detect tables first, then recognize their structure. Recently, interest has surged in dev…
cs.LG2023
A Graphical Approach to Document Layout Analysis
Jilin Wang, Michael Krumdick, Baojia Tong +5
Document layout analysis (DLA) is the task of detecting the distinct, semantic content within a document and correctly classifying these items into an appropriate category (e.g., t…