15 papers
Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+
Sherzod Hakimov, Karl Osswald, Jelle Psurek +3
We evaluate large language models (LLMs) as language agents playing goal-directed dialogue games in self-play across 30 languages: the 24 official EU languages plus six others. Unl…
The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue
Sherzod Hakimov, Mattia D'Agostini, Ivan Samodelkin +1
We introduce the Image Reconstruction Game, a fully automated benchmark in which a vision-language model issues corrective instructions to an image generator across multiple turns,…
Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely
Chalamalasetti Kranti, Sherzod Hakimov, David Schlangen
Robots operating in diverse environments rely on visual input to interpret objects and spatial layouts. In human-collaborative tasks, they are expected to communicate this understa…
What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025
Jing Yang, Nils Feldhus, Salar Mohtaj +10
As Natural Language Generation (NLG) dominates modern NLP, scalable evaluation remains a critical bottleneck. Consequently, LLM-as-a-judge (LaaJ) adoption has accelerated rapidly,…
TurkicNLP: An NLP Toolkit for Turkic Languages
Sherzod Hakimov
Natural language processing for the Turkic language family, spoken by over 200 million people across Eurasia, remains fragmented, with most languages lacking unified tooling and re…
A Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembench
David Schlangen, Sherzod Hakimov, Chalamalasetti Kranti +2
There are currently two main paradigms for evaluating large language models (LLMs), reference-based evaluation and preference-based evaluation. The first, carried over from the eva…