activity
20242026
most citedDocling: An Efficient Open-Source Toolkit for AI-driven Document Conversion

1 citations · 3 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV20261 cited

MarkushGrapher-2: End-to-end Multimodal Recognition of Chemical Structures

Tim Strohmeyer, Lucas Morin, Gerhard Ingmar Meijer +3

Automatically extracting chemical structures from documents is essential for the large-scale analysis of the literature in chemistry. Automatic pipelines have been developed to rec…

cs.CV2026

ScreenParse: Moving Beyond Sparse Grounding with Complete Screen Parsing Supervision

A. Said Gurbuz, Sunghwan Hong, Ahmed Nassar +2

Modern computer-use agents (CUA) must perceive a screen as a structured state, what elements are visible, where they are, and what text they contain, before they can reliably groun…

cs.CV2025

Advanced Layout Analysis Models for Docling

Nikolaos Livathinos, Christoph Auer, Ahmed Nassar +16

This technical report documents the development of novel Layout Analysis models integrated into the Docling document-conversion pipeline. We trained several state-of-the-art object…

cs.CV2025

MarkushGrapher: Joint Visual and Textual Recognition of Markush Structures

Lucas Morin, Valéry Weber, Ahmed Nassar +4

The automated analysis of chemical literature holds promise to accelerate discovery in fields such as material science and drug development. In particular, search capabilities for…

cs.CV2025

SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Ahmed Nassar, Andres Marafioti, Matteo Omenetti +10

We introduce SmolDocling, an ultra-compact vision-language model targeting end-to-end document conversion. Our model comprehensively processes entire pages by generating DocTags, a…

cs.CV20251 cited

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence

Granite Vision Team, Leonid Karlinsky, Assaf Arbelle +60

We introduce Granite Vision, a lightweight large language model with vision capabilities, specifically designed to excel in enterprise use cases, particularly in visual document un…