6 papers · 1 filter
Reading or Guessing? Visual Grounding Failures of Vision-Language Models for OCR in Ancient Greek Editions
Antonia Karamolegkou, Nicolas Angleraud, Benoît Sagot +1
Recent work has shown that Vision-Language Models (VLMs) used for optical character recognition (OCR) can generate plausible but visually unsupported text, suggesting reliance on l…
Pre-Editorial Normalization for Automatically Transcribed Medieval Manuscripts in Old French and Latin
Thibault Clérice, Rachel Bawden, Anthony Glaise +2
Recent advances in Automatic Text Recognition (ATR) have improved access to historical archives, yet a methodological divide persists between palaeographic transcriptions and norma…
How Should We Model the Probability of a Language?
Rasul Dent, Pedro Ortiz Suarez, Thibault Clérice +1
Of the over 7,000 languages spoken in the world, commercial language identification (LID) systems only reliably identify a few hundred in written form. Research-grade systems exten…
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett +94
Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heter…
KréyoLID From Language Identification Towards Language Mining
Rasul Dent, Pedro Ortiz Suarez, Thibault Clérice +1
Automatic language identification is frequently framed as a multi-class classification problem. However, when creating digital corpora for less commonly written languages, it may b…
Molyé: A Corpus-based Approach to Language Contact in Colonial France
Rasul Dent, Juliette Janès, Thibault Clérice +2
Whether or not several Creole languages which developed during the early modern period can be considered genetic descendants of European languages has been the subject of intense d…