Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Towards Hierarchical Structure Understanding of Newspaper Images
William Mocaër, Solène Tarride, Thomas Constum +7
Understanding newspaper images remains a challenging task due to their complex, nested hierarchical structures and dense, heterogeneous layouts. In this paper, we explore two compl…
cs.CV2026
Few-shot Writer Adaptation via Multimodal In-Context Learning
Tom Simon, Stephane Nicolas, Pierrick Tranouez +2
While state-of-the-art Handwritten Text Recognition (HTR) models perform well on standard benchmarks, they frequently struggle with writers exhibiting highly specific styles that a…
cs.CV2025
Classifying the Unknown: In-Context Learning for Open-Vocabulary Text and Symbol Recognition
Tom Simon, William Mocaer, Pierrick Tranouez +2
We introduce Rosetta, a multimodal model that leverages Multimodal In-Context Learning (MICL) to classify sequences of novel script patterns in documents by leveraging minimal exam…