7 papers
Robust Language Identification for Romansh Varieties
Charlotte Model, Sina Ahmadi, Jannis Vamvas
The Romansh language has several regional varieties, called idioms, which sometimes have limited mutual intelligibility. Despite this linguistic diversity, there has been a lack of…
SwissGov-RSD: A Human-annotated, Cross-lingual Benchmark for Token-level Recognition of Semantic Differences Between Related Documents
Michelle Wastl, Jannis Vamvas, Rico Sennrich
Recognizing semantic differences across documents is crucial for text generation evaluation and content alignment, especially in cross-lingual settings. However, as a standalone ta…
RUMLEM: A Dictionary-Based Lemmatizer for Romansh
Dominic P. Fischer, Zachary Hopton, Jannis Vamvas
Lemmatization -- the task of mapping an inflected word form to its dictionary form -- is a crucial component of many NLP applications. In this paper, we present RUMLEM, a lemmatize…
The Mediomatix Corpus: Parallel Data for Romansh Language Varieties via Comparable Schoolbooks
Zachary Hopton, Jannis Vamvas, Andrin Büchler +3
The five idioms (i.e., varieties) of the Romansh language are largely standardized and are taught in the schools of the respective communities in Switzerland. In this paper, we pre…
QueST: Incentivizing LLMs to Generate Difficult Problems
Hanxu Hu, Xingxing Zhang, Jannis Vamvas +2
Large Language Models have achieved strong performance on reasoning tasks, solving competition-level coding and math problems. However, their scalability is limited by human-labele…
20min-XD: A Comparable Corpus of Swiss News Articles
Michelle Wastl, Jannis Vamvas, Selena Calleri +1
We present 20min-XD (20 Minuten cross-lingual document-level), a French-German, document-level comparable corpus of news articles, sourced from the Swiss online news outlet 20 Minu…