collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

Information Representation Fairness in Long-Document Embeddings: The Peculiar Interaction of Positional and Language Bias

Elias Schuhmacher, Andrianos Michail, Juri Opitz +2

To be discoverable in an embedding-based search process, each part of a document should be reflected in its embedding representation. To quantify any potential reflection biases, w…

cs.CL2025

Adapting Multilingual Embedding Models to Historical Luxembourgish

Andrianos Michail, Corina Julia Raclé, Juri Opitz +1

The growing volume of digitized historical texts requires effective semantic search using text embeddings. However, pre-trained multilingual models face challenges with historical…

cs.CL2025

Interpretable Text Embeddings and Text Similarity Explanation: A Survey

Juri Opitz, Lucas Möller, Andrianos Michail +2

Text embeddings are a fundamental component in many NLP tasks, including classification, regression, clustering, and semantic search. However, despite their ubiquitous application,…

cs.CL2025

Sentence Smith: Controllable Edits for Evaluating Text Embeddings

Hongji Li, Andrianos Michail, Reto Gubelmann +2

Controllable and transparent text generation has been a long-standing goal in NLP. Almost as long-standing is a general idea for addressing this challenge: Parsing text to a symbol…

cs.CL2025

MMTEB: Massive Multilingual Text Embedding Benchmark

Kenneth Enevoldsen, Isaac Chung, Imene Kerboua +83

Text embeddings are typically evaluated on a limited set of tasks, which are constrained by language, domain, and task diversity. To address these limitations and provide a more co…

cs.CL2025

Examining Multilingual Embedding Models Cross-Lingually Through LLM-Generated Adversarial Examples

Andrianos Michail, Simon Clematide, Rico Sennrich

The evaluation of cross-lingual semantic search models is often limited to existing datasets from tasks such as information retrieval and semantic textual similarity. We introduce…