2 papers
cs.CL2026
Multimodal Evaluation of Russian-language Architectures
Artem Chervyakov, Ulyana Isaeva, Anton Emelyanov +15
Multimodal large language models (MLLMs) are currently at the center of research attention, showing rapid progress in scale and capabilities, yet their intelligence, limitations, a…
cs.CL2025
Eye of Judgement: Dissecting the Evaluation of Russian-speaking LLMs with POLLUX
Nikita Martynov, Anastasia Mordasheva, Dmitriy Gorbetskiy +8
We introduce POLLUX, a comprehensive open-source benchmark designed to evaluate the generative capabilities of large language models (LLMs) in Russian. Our main contribution is a n…