5 papers
Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images
A. Said Gurbuz, Ahmed Nassar, Christoph Auer +8
Document processing pipelines traditionally cascade optical character recognition (OCR) engines with downstream models for structured information extraction, leading to multi-stage…
Structured Layout Priors for Robust Out-of-Distribution Visual Document Understanding
Peter El Hachem, Ahmed Nassar, A. Said Gurbuz +2
Vision-Language Models (VLMs) parse documents end-to-end but frequently break down on layouts unlike those seen in training. We attribute this to a two-hop bottleneck: before the d…
ScreenParse: Moving Beyond Sparse Grounding with Complete Screen Parsing Supervision
A. Said Gurbuz, Sunghwan Hong, Ahmed Nassar +2
Modern computer-use agents (CUA) must perceive a screen as a structured state, what elements are visible, where they are, and what text they contain, before they can reliably groun…
Advanced Layout Analysis Models for Docling
Nikolaos Livathinos, Christoph Auer, Ahmed Nassar +16
This technical report documents the development of novel Layout Analysis models integrated into the Docling document-conversion pipeline. We trained several state-of-the-art object…
SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
Ahmed Nassar, Andres Marafioti, Matteo Omenetti +10
We introduce SmolDocling, an ultra-compact vision-language model targeting end-to-end document conversion. Our model comprehensively processes entire pages by generating DocTags, a…