3 papers
cs.CL2026
How Far Can Machine Translation Quality Take You? Extrinsic Discourse Evaluation in Goal-Oriented Setups
Wafaa Mohammed, Kata Naszadi, Vlad Niculae
Existing machine translation (MT) metrics and discourse-focused evaluations primarily assess translation quality intrinsically, without measuring the downstream consequences of tra…
cs.CL2026
Do Language Models Reason Across Languages?
Yan Meng, Wafaa Mohammed, Christof Monz
The real-world information sources are inherently multilingual, which naturally raises a question about whether language models can synthesize information across languages. In this…
cs.LG2025
Control the Temperature: Selective Sampling for Diverse and High-Quality LLM Outputs
Sergey Troshin, Wafaa Mohammed, Yan Meng +3
Diversity is an essential metric for evaluating the creativity of outputs generated by language models. Temperature-based sampling is a common strategy to increase diversity. Howev…