5 papers
ALEE: Any-Language Evaluation of Embeddings via English-Centric Minimal Pairs
Andrianos Michail, Stylianos Psychias, Michelle Wastl +3
Text embeddings are standard for semantic similarity tasks, yet their evaluation remains an open challenge. Current benchmarks are static, cover only a limited set of languages, ar…
Scaling Unsupervised Word Alignment to Documents via Structural Constraints
Michelle Wastl, Jannis Vamvas, Rico Sennrich
Word alignment has traditionally been studied between sentences, but many cross-lingual tasks increasingly require correspondences across full documents. While recent multilingual…
It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief
Kevin Du, Clara Kümpel, Michelle Wastl +1
Users frequently express their beliefs to large language models (LLMs). In some situations, the LLM should accept these contextual beliefs as true. In others, they should stick to…
SwissGov-RSD: A Human-annotated, Cross-lingual Benchmark for Token-level Recognition of Semantic Differences Between Related Documents
Michelle Wastl, Jannis Vamvas, Rico Sennrich
Recognizing semantic differences across documents is crucial for text generation evaluation and content alignment, especially in cross-lingual settings. However, as a standalone ta…
20min-XD: A Comparable Corpus of Swiss News Articles
Michelle Wastl, Jannis Vamvas, Selena Calleri +1
We present 20min-XD (20 Minuten cross-lingual document-level), a French-German, document-level comparable corpus of news articles, sourced from the Swiss online news outlet 20 Minu…