activity
20242026
most citedCLEF HIPE-2026: Evaluating Accurate and Efficient Person-Place Relation Extraction from Multilingual Historical Texts

1 citations · 1 across the 6 of their papers we have counts for

collaborators
Showing cs.CLShow all

10 papers · 1 filter

cs.CL2026

ICDAR 2026 HIPE-OCRepair Competition on LLM-Assisted OCR Post-Correction for Historical Documents

Maud Ehrmann, Emanuela Boros, Juri Opitz +3

We present the results of HIPE-OCRepair-2026, an ICDAR competition on LLM-assisted OCR post-correction of historical documents. OCR post-correction remains a long-standing challeng…

cs.CL2026

ALEE: Any-Language Evaluation of Embeddings via English-Centric Minimal Pairs

Andrianos Michail, Stylianos Psychias, Michelle Wastl +3

Text embeddings are standard for semantic similarity tasks, yet their evaluation remains an open challenge. Current benchmarks are static, cover only a limited set of languages, ar…

cs.CL2026

Overview of HIPE-2026: Person-Place Relation Extraction from Multilingual Historical Texts

Juri Opitz, Maud Ehrmann, Corina Raclé +3

Was this person ever at that place, and if so, when? Answering such questions from noisy, multilingual historical documents is the central challenge of HIPE-2026, the third edition…

cs.CL2026

Information Representation Fairness in Long-Document Embeddings: The Peculiar Interaction of Positional and Language Bias

Elias Schuhmacher, Andrianos Michail, Juri Opitz +2

To be discoverable in an embedding-based search process, each part of a document should be reflected in its embedding representation. To quantify any potential reflection biases, w…

cs.CL2025

Sentence Smith: Controllable Edits for Evaluating Text Embeddings

Hongji Li, Andrianos Michail, Reto Gubelmann +2

Controllable and transparent text generation has been a long-standing goal in NLP. Almost as long-standing is a general idea for addressing this challenge: Parsing text to a symbol…

cs.CL2025

MMTEB: Massive Multilingual Text Embedding Benchmark

Kenneth Enevoldsen, Isaac Chung, Imene Kerboua +83

Text embeddings are typically evaluated on a limited set of tasks, which are constrained by language, domain, and task diversity. To address these limitations and provide a more co…