activity
20242026
collaborators

5 papers

cs.CV2026

A Benchmark of State-Space Models vs. Transformers and BiLSTM-based Models for Historical Newspaper OCR

Merveilles Agbeti-Messan, Pierrick Tranouez, Stéphane Nicolas +2

End-to-end OCR for historical newspapers remains challenging, as models must handle long text sequences, degraded print quality, and complex layouts. While Transformer-based recogn…

cs.CV2026

Scaling State-Space Models from Lines to Paragraphs: An Ablation of Mamba-based OCR

Merveilles Agbeti-Messan, Pierrick Tranouez, Stéphane Nicolas +2

End-to-end OCR increasingly relies on autoregressive sequence models, where the quadratic cost of Transformer attention limits efficient transcription of long, paragraph-level text…

cs.CV2026

Few-shot Writer Adaptation via Multimodal In-Context Learning

Tom Simon, Stephane Nicolas, Pierrick Tranouez +2

While state-of-the-art Handwritten Text Recognition (HTR) models perform well on standard benchmarks, they frequently struggle with writers exhibiting highly specific styles that a…

cs.CV2025

End-to-End Full-Page Optical Music Recognition for Pianoform Sheet Music

Antonio Ríos-Vila, Jorge Calvo-Zaragoza, David Rizo +1

Optical Music Recognition (OMR) has made significant progress since its inception, with various approaches now capable of accurately transcribing music scores into digital formats.…

cs.CV2024

Sheet Music Transformer: End-To-End Optical Music Recognition Beyond Monophonic Transcription

Antonio Ríos-Vila, Jorge Calvo-Zaragoza, Thierry Paquet

State-of-the-art end-to-end Optical Music Recognition (OMR) has, to date, primarily been carried out using monophonic transcription techniques to handle complex score layouts, such…