activity
20242026
collaborators

8 papers

cs.CV2026

Identify, Locate, Link: End-to-End Key-Value Extraction from Document Images

A. Said Gurbuz, Ahmed Nassar, Christoph Auer +8

Document processing pipelines traditionally cascade optical character recognition (OCR) engines with downstream models for structured information extraction, leading to multi-stage…

cs.CV2026

MarkushGrapher-2: End-to-end Multimodal Recognition of Chemical Structures

Tim Strohmeyer, Lucas Morin, Gerhard Ingmar Meijer +3

Automatically extracting chemical structures from documents is essential for the large-scale analysis of the literature in chemistry. Automatic pipelines have been developed to rec…

cs.CV2025

SubGrapher: Visual Fingerprinting of Chemical Structures

Lucas Morin, Gerhard Ingmar Meijer, Valéry Weber +2

Automatic extraction of chemical structures from scientific literature plays a crucial role in accelerating research across fields ranging from drug discovery to materials science.…

cs.CV2025

Advanced Layout Analysis Models for Docling

Nikolaos Livathinos, Christoph Auer, Ahmed Nassar +16

This technical report documents the development of novel Layout Analysis models integrated into the Docling document-conversion pipeline. We trained several state-of-the-art object…

cs.CV2025

MarkushGrapher: Joint Visual and Textual Recognition of Markush Structures

Lucas Morin, Valéry Weber, Ahmed Nassar +4

The automated analysis of chemical literature holds promise to accelerate discovery in fields such as material science and drug development. In particular, search capabilities for…

cs.CV2025

SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Ahmed Nassar, Andres Marafioti, Matteo Omenetti +10

We introduce SmolDocling, an ultra-compact vision-language model targeting end-to-end document conversion. Our model comprehensively processes entire pages by generating DocTags, a…