2 papers
cs.CV2026
When Low CER is Not Enough: An Analysis of Hallucinations in Vision-Language OCR Systems on Historical Uruguayan Documents
Marina Gardella, Camilo Mari{ñ}o, Diego Belzarena +3
Optical Character Recognition (OCR) is a key component in the digitization of historical archives. Recently, Vision-Language Models (VLMs) have emerged as strong alternatives to tr…
cs.CV2025
Improving OCR using internal document redundancy
Diego Belzarena, Seginus Mowlavi, Aitor Artola +9
Current OCR systems are based on deep learning models trained on large amounts of data. Although they have shown some ability to generalize to unseen data, especially in detection…