8 papers
Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness
Pinzhen Chen, Koel Dutta Chowdhury, Xiaoya Xu +20
Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone…
Want Better Synthetic Data? Steer It: Activation Steering for Low-Resource Language Generation
Jan Cegin, Daniil Gurgurov, Yusser Al Ghussin +1
Large language models (LLMs) have become an effective tool for synthetic data generation, including for low-resource languages, where generated data can improve downstream task per…
DFKI-MLT at SemEval-2026 TASK 7: Steering Multilingual Models Towards Cultural Knowledge
Yusser Al Ghussin, Daniil Gurgurov, Yasser Hamidullah +3
Large language models (LLMs) are increasingly used across diverse linguistic and cultural contexts, yet their cultural knowledge remains uneven across regions and languages. We pre…
Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection
Yusser Al Ghussin, Daniil Gurgurov, Tanja Baeumel +3
Sparse autoencoders (SAEs) enable feature-level mechanistic interpretability and activation steering in large language models (LLMs), but SAE-based language control remains unrelia…
CLaS-Bench: A Cross-Lingual Alignment and Steering Benchmark
Daniil Gurgurov, Yusser Al Ghussin, Tanja Baeumel +5
Understanding and controlling the behavior of large language models (LLMs) is an increasingly important topic in multilingual NLP. Beyond prompting or fine-tuning, , i.e.,~manipula…
Modular Arithmetic: Language Models Solve Math Digit by Digit
Tanja Baeumel, Daniil Gurgurov, Yusser al Ghussin +2
While recent work has begun to uncover the internal strategies that Large Language Models (LLMs) employ for simple arithmetic tasks, a unified understanding of their underlying mec…