activity
20242026
collaborators

15 papers

cs.CL2026

Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+

Sherzod Hakimov, Karl Osswald, Jelle Psurek +3

We evaluate large language models (LLMs) as language agents playing goal-directed dialogue games in self-play across 30 languages: the 24 official EU languages plus six others. Unl…

cs.CV2026

The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue

Sherzod Hakimov, Mattia D'Agostini, Ivan Samodelkin +1

We introduce the Image Reconstruction Game, a fully automated benchmark in which a vision-language model issues corrective instructions to an image generator across multiple turns,…

cs.CL2026

Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely

Chalamalasetti Kranti, Sherzod Hakimov, David Schlangen

Robots operating in diverse environments rely on visual input to interpret objects and spatial layouts. In human-collaborative tasks, they are expected to communicate this understa…

cs.CL2026

What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025

Jing Yang, Nils Feldhus, Salar Mohtaj +10

As Natural Language Generation (NLG) dominates modern NLP, scalable evaluation remains a critical bottleneck. Consequently, LLM-as-a-judge (LaaJ) adoption has accelerated rapidly,…

cs.CL2026

TurkicNLP: An NLP Toolkit for Turkic Languages

Sherzod Hakimov

Natural language processing for the Turkic language family, spoken by over 200 million people across Eurasia, remains fragmented, with most languages lacking unified tooling and re…

cs.CL2026

A Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembench

David Schlangen, Sherzod Hakimov, Chalamalasetti Kranti +2

There are currently two main paradigms for evaluating large language models (LLMs), reference-based evaluation and preference-based evaluation. The first, carried over from the eva…