collaborators

6 papers

cs.CL2025

Robustness of Misinformation Classification Systems to Adversarial Examples Through BeamAttack

Arnisa Fazla, Lucas Krauter, David Guzman Piedrahita +1

We extend BeamAttack, an adversarial attack algorithm designed to evaluate the robustness of text classification systems through word-level modifications guided by beam search. Our…

cs.CL2025

Adapting Multilingual Embedding Models to Historical Luxembourgish

Andrianos Michail, Corina Julia Raclé, Juri Opitz +1

The growing volume of digitized historical texts requires effective semantic search using text embeddings. However, pre-trained multilingual models face challenges with historical…

cs.CL2025

Interpretable Text Embeddings and Text Similarity Explanation: A Survey

Juri Opitz, Lucas Möller, Andrianos Michail +2

Text embeddings are a fundamental component in many NLP tasks, including classification, regression, clustering, and semantic search. However, despite their ubiquitous application,…

cs.CL2025

Sentence Smith: Controllable Edits for Evaluating Text Embeddings

Hongji Li, Andrianos Michail, Reto Gubelmann +2

Controllable and transparent text generation has been a long-standing goal in NLP. Almost as long-standing is a general idea for addressing this challenge: Parsing text to a symbol…

cs.CL2025

MMTEB: Massive Multilingual Text Embedding Benchmark

Kenneth Enevoldsen, Isaac Chung, Imene Kerboua +83

Text embeddings are typically evaluated on a limited set of tasks, which are constrained by language, domain, and task diversity. To address these limitations and provide a more co…

cs.CL2025

Examining Multilingual Embedding Models Cross-Lingually Through LLM-Generated Adversarial Examples

Andrianos Michail, Simon Clematide, Rico Sennrich

The evaluation of cross-lingual semantic search models is often limited to existing datasets from tasks such as information retrieval and semantic textual similarity. We introduce…