17 papers
Scaling Unsupervised Word Alignment to Documents via Structural Constraints
Michelle Wastl, Jannis Vamvas, Rico Sennrich
Word alignment has traditionally been studied between sentences, but many cross-lingual tasks increasingly require correspondences across full documents. While recent multilingual…
Reinforcement Learning Elicits Contextual Learning of Unseen Language Translation
Hanxu Hu, ZdenÄk Å najdr, Pinzhen Chen +2
Prior work has shown that large language models (LLMs) can translate unseen or low-resource languages by undergoing continued training or even by encoding a grammar book in their c…
Attention Calibration for Position-Fair Dense Information Retrieval
Andrianos Michail, Elias Schuhmacher, Juri Opitz +2
Dense retrieval models exhibit positional bias: retrieval effectiveness degrades when relevant information appears later in a passage (Zeng et al., 2025). We ask whether this bias…
SwissGov-RSD: A Human-annotated, Cross-lingual Benchmark for Token-level Recognition of Semantic Differences Between Related Documents
Michelle Wastl, Jannis Vamvas, Rico Sennrich
Recognizing semantic differences across documents is crucial for text generation evaluation and content alignment, especially in cross-lingual settings. However, as a standalone ta…
Translation Asymmetry in LLMs as a Data Augmentation Factor: A Case Study for 6 Romansh Language Varieties
Jannis Vamvas, Ignacio Pérez Prat, Angela Heldstab +3
Recent strategies for low-resource machine translation rely on LLMs to generate synthetic data from higher-resource languages. We find that this method fails for Romansh, because L…
CHARM: Calibrating Reward Models With Chatbot Arena Scores
Xiao Zhu, Chenmien Tan, Pinzhen Chen +4
Reward models (RMs) play a crucial role in Reinforcement Learning from Human Feedback by serving as proxies for human preferences in aligning large language models. However, they s…