2 papers
cs.CL2025
Eye of Judgement: Dissecting the Evaluation of Russian-speaking LLMs with POLLUX
Nikita Martynov, Anastasia Mordasheva, Dmitriy Gorbetskiy +8
We introduce POLLUX, a comprehensive open-source benchmark designed to evaluate the generative capabilities of large language models (LLMs) in Russian. Our main contribution is a n…
cs.CL2024
MERA: A Comprehensive LLM Evaluation in Russian
Alena Fenogenova, Artem Chervyakov, Nikita Martynov +16
Over the past few years, one of the most notable advancements in AI research has been in foundation models (FMs), headlined by the rise of language models (LMs). As the models' siz…