4 papers
Pre-Editorial Normalization for Automatically Transcribed Medieval Manuscripts in Old French and Latin
Thibault Clérice, Rachel Bawden, Anthony Glaise +2
Recent advances in Automatic Text Recognition (ATR) have improved access to historical archives, yet a methodological divide persists between palaeographic transcriptions and norma…
How Should We Model the Probability of a Language?
Rasul Dent, Pedro Ortiz Suarez, Thibault Clérice +1
Of the over 7,000 languages spoken in the world, commercial language identification (LID) systems only reliably identify a few hundred in written form. Research-grade systems exten…
KréyoLID From Language Identification Towards Language Mining
Rasul Dent, Pedro Ortiz Suarez, Thibault Clérice +1
Automatic language identification is frequently framed as a multi-class classification problem. However, when creating digital corpora for less commonly written languages, it may b…
Diachronic Document Dataset for Semantic Layout Analysis
Thibault Clérice, Juliette Janes, Hugo Scheithauer +5
We present a novel, open-access dataset designed for semantic layout analysis, built to support document recreation workflows through mapping with the Text Encoding Initiative (TEI…