4 papers
DenTab: A Dataset for Table Recognition and Visual QA on Real-World Dental Estimates
Laziz Hamdi, Amine Tamasna, Thierry Paquet
Tables condense key transactional and administrative information into compact layouts, but practical extraction requires more than text recognition: systems must also recover struc…
TableSeq: Unified Generation of Structure, Content, and Layout
Laziz Hamdi, Amine Tamasna, Pascal Boisson +1
We present TableSeq, an image-only, end-to-end framework for joint table structure recognition, content recognition, and cell localization. The model formulates these tasks as a si…
Classifying the Unknown: In-Context Learning for Open-Vocabulary Text and Symbol Recognition
Tom Simon, William Mocaer, Pierrick Tranouez +2
We introduce Rosetta, a multimodal model that leverages Multimodal In-Context Learning (MICL) to classify sequences of novel script patterns in documents by leveraging minimal exam…
PILOT: A Promptable Interleaved Layout-aware OCR Transformer
Laziz Hamdi, Amine Tamasna, Pascal Boisson +1
Classical OCR pipelines decompose document reading into detection, segmentation, and recognition stages, which makes them sensitive to localization errors and difficult to extend t…