From the 1 of 5 linked papers with an AI index.
5 papers
Towards Hierarchical Structure Understanding of Newspaper Images
William Mocaër, Solène Tarride, Thomas Constum +7
The paper proposes two methods for parsing the complex hierarchical layout of newspaper images: a modular bottom‑up pipeline using existing models (YOLO, LayoutReader) and a new en…
A Benchmark of State-Space Models vs. Transformers and BiLSTM-based Models for Historical Newspaper OCR
Merveilles Agbeti-Messan, Pierrick Tranouez, Stéphane Nicolas +2
End-to-end OCR for historical newspapers remains challenging, as models must handle long text sequences, degraded print quality, and complex layouts. While Transformer-based recogn…
Scaling State-Space Models from Lines to Paragraphs: An Ablation of Mamba-based OCR
Merveilles Agbeti-Messan, Pierrick Tranouez, Stéphane Nicolas +2
End-to-end OCR increasingly relies on autoregressive sequence models, where the quadratic cost of Transformer attention limits efficient transcription of long, paragraph-level text…
Few-shot Writer Adaptation via Multimodal In-Context Learning
Tom Simon, Stephane Nicolas, Pierrick Tranouez +2
While state-of-the-art Handwritten Text Recognition (HTR) models perform well on standard benchmarks, they frequently struggle with writers exhibiting highly specific styles that a…
Classifying the Unknown: In-Context Learning for Open-Vocabulary Text and Symbol Recognition
Tom Simon, William Mocaer, Pierrick Tranouez +2
We introduce Rosetta, a multimodal model that leverages Multimodal In-Context Learning (MICL) to classify sequences of novel script patterns in documents by leveraging minimal exam…