1 citations · 3 across the 5 of their papers we have counts for
7 papers · 1 filter
Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images
A. Said Gurbuz, Ahmed Nassar, Christoph Auer +8
Document processing pipelines traditionally cascade optical character recognition (OCR) engines with downstream models for structured information extraction, leading to multi-stage…
Advanced Layout Analysis Models for Docling
Nikolaos Livathinos, Christoph Auer, Ahmed Nassar +16
This technical report documents the development of novel Layout Analysis models integrated into the Docling document-conversion pipeline. We trained several state-of-the-art object…
SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
Ahmed Nassar, Andres Marafioti, Matteo Omenetti +10
We introduce SmolDocling, an ultra-compact vision-language model targeting end-to-end document conversion. Our model comprehensively processes entire pages by generating DocTags, a…
Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence
Granite Vision Team, Leonid Karlinsky, Assaf Arbelle +60
We introduce Granite Vision, a lightweight large language model with vision capabilities, specifically designed to excel in enterprise use cases, particularly in visual document un…
ICDAR 2023 Competition on Robust Layout Segmentation in Corporate Documents
Christoph Auer, Ahmed Nassar, Maksym Lysak +3
Transforming documents into machine-processable representations is a challenging task due to their complex structures and variability in formats. Recovering the layout structure an…
Optimized Table Tokenization for Table Structure Recognition
Maksym Lysak, Ahmed Nassar, Nikolaos Livathinos +2
Extracting tables from documents is a crucial task in any document conversion pipeline. Recently, transformer-based models have demonstrated that table-structure can be recognized…