1 citations · 1 across the 6 of their papers we have counts for
10 papers · 1 filter
ICDAR 2026 HIPE-OCRepair Competition on LLM-Assisted OCR Post-Correction for Historical Documents
Maud Ehrmann, Emanuela Boros, Juri Opitz +3
We present the results of HIPE-OCRepair-2026, an ICDAR competition on LLM-assisted OCR post-correction of historical documents. OCR post-correction remains a long-standing challeng…
ALEE: Any-Language Evaluation of Embeddings via English-Centric Minimal Pairs
Andrianos Michail, Stylianos Psychias, Michelle Wastl +3
Text embeddings are standard for semantic similarity tasks, yet their evaluation remains an open challenge. Current benchmarks are static, cover only a limited set of languages, ar…
Overview of HIPE-2026: Person-Place Relation Extraction from Multilingual Historical Texts
Juri Opitz, Maud Ehrmann, Corina Raclé +3
Was this person ever at that place, and if so, when? Answering such questions from noisy, multilingual historical documents is the central challenge of HIPE-2026, the third edition…
Information Representation Fairness in Long-Document Embeddings: The Peculiar Interaction of Positional and Language Bias
Elias Schuhmacher, Andrianos Michail, Juri Opitz +2
To be discoverable in an embedding-based search process, each part of a document should be reflected in its embedding representation. To quantify any potential reflection biases, w…
Sentence Smith: Controllable Edits for Evaluating Text Embeddings
Hongji Li, Andrianos Michail, Reto Gubelmann +2
Controllable and transparent text generation has been a long-standing goal in NLP. Almost as long-standing is a general idea for addressing this challenge: Parsing text to a symbol…
MMTEB: Massive Multilingual Text Embedding Benchmark
Kenneth Enevoldsen, Isaac Chung, Imene Kerboua +83
Text embeddings are typically evaluated on a limited set of tasks, which are constrained by language, domain, and task diversity. To address these limitations and provide a more co…