activity
20182026
most citedDocling Technical Report

12 citations · 23 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images

A. Said Gurbuz, Ahmed Nassar, Christoph Auer +8

Document processing pipelines traditionally cascade optical character recognition (OCR) engines with downstream models for structured information extraction, leading to multi-stage…

cs.CV2026

Structured Layout Priors for Robust Out-of-Distribution Visual Document Understanding

Peter El Hachem, Ahmed Nassar, A. Said Gurbuz +2

Vision-Language Models (VLMs) parse documents end-to-end but frequently break down on layouts unlike those seen in training. We attribute this to a two-hop bottleneck: before the d…

cs.CV2025

Advanced Layout Analysis Models for Docling

Nikolaos Livathinos, Christoph Auer, Ahmed Nassar +16

This technical report documents the development of novel Layout Analysis models integrated into the Docling document-conversion pipeline. We trained several state-of-the-art object…

cs.CV2025

SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Ahmed Nassar, Andres Marafioti, Matteo Omenetti +10

We introduce SmolDocling, an ultra-compact vision-language model targeting end-to-end document conversion. Our model comprehensively processes entire pages by generating DocTags, a…

cs.CV20251 cited

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence

Granite Vision Team, Leonid Karlinsky, Assaf Arbelle +60

We introduce Granite Vision, a lightweight large language model with vision capabilities, specifically designed to excel in enterprise use cases, particularly in visual document un…

cs.CV2023

ICDAR 2023 Competition on Robust Layout Segmentation in Corporate Documents

Christoph Auer, Ahmed Nassar, Maksym Lysak +3

Transforming documents into machine-processable representations is a challenging task due to their complex structures and variability in formats. Recovering the layout structure an…