activity
20242026
collaborators

14 papers

cs.CL2026

DocAtlas: Multilingual Document Understanding Across 80+ Languages

Ahmed Heakl, Youssef Mohamed, Abdullah Sohail +6

Multilingual document understanding remains limited for low-resource languages due to scarce training data and model-based annotation pipelines that perpetuate existing biases. We…

cs.CV2026

Structured Layout Priors for Robust Out-of-Distribution Visual Document Understanding

Peter El Hachem, Ahmed Nassar, A. Said Gurbuz +2

Vision-Language Models (VLMs) parse documents end-to-end but frequently break down on layouts unlike those seen in training. We attribute this to a two-hop bottleneck: before the d…

cs.CV2026

ScreenParse: Moving Beyond Sparse Grounding with Complete Screen Parsing Supervision

A. Said Gurbuz, Sunghwan Hong, Ahmed Nassar +2

Modern computer-use agents (CUA) must perceive a screen as a structured state, what elements are visible, where they are, and what text they contain, before they can reliably groun…

cs.CV2026

ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding

Jovana Kondic, Pengyuan Li, Dhiraj Joshi +24

Understanding charts requires models to jointly reason over geometric visual patterns, structured numerical data, and natural language -- a capability where current vision-language…

cs.CV2026

MarkushGrapher-2: End-to-end Multimodal Recognition of Chemical Structures

Tim Strohmeyer, Lucas Morin, Gerhard Ingmar Meijer +3

Automatically extracting chemical structures from documents is essential for the large-scale analysis of the literature in chemistry. Automatic pipelines have been developed to rec…

cs.CV2025

SubGrapher: Visual Fingerprinting of Chemical Structures

Lucas Morin, Gerhard Ingmar Meijer, Valéry Weber +2

Automatic extraction of chemical structures from scientific literature plays a crucial role in accelerating research across fields ranging from drug discovery to materials science.…