collaborators

6 papers

cs.CL2025

GigaChat Family: Efficient Russian Language Modeling Through Mixture of Experts Architecture

GigaChat team, Mamedov Valentin, Evgenii Kosarev +31

Generative large language models (LLMs) have become crucial for modern NLP research and applications across various languages. However, the development of foundational models speci…

cs.CL2025

The Russian-focused embedders' exploration: ruMTEB benchmark and Russian embedding model design

Artem Snegirev, Maria Tikhonova, Anna Maksimova +2

Embedding models play a crucial role in Natural Language Processing (NLP) by creating text embeddings used in various tasks such as information retrieval and assessing semantic tex…

cs.CL2024

RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs

Ekaterina Taktasheva, Maxim Bazhukov, Kirill Koncha +3

Minimal pairs are a well-established approach to evaluating the grammatical knowledge of language models. However, existing resources for minimal pairs address a limited number of…

cs.CL2024

Long Input Benchmark for Russian Analysis

Igor Churin, Murat Apishev, Maria Tikhonova +5

Recent advancements in Natural Language Processing (NLP) have fostered the development of Large Language Models (LLMs) that can solve an immense variety of tasks. One of the key as…

cs.CL2024

MERA: A Comprehensive LLM Evaluation in Russian

Alena Fenogenova, Artem Chervyakov, Nikita Martynov +16

Over the past few years, one of the most notable advancements in AI research has been in foundation models (FMs), headlined by the rise of language models (LMs). As the models' siz…

cs.CL2024

A Family of Pretrained Transformer Language Models for Russian

Dmitry Zmitrovich, Alexander Abramov, Andrey Kalmykov +10

Transformer language models (LMs) are fundamental to NLP research methodologies and applications in various languages. However, developing such models specifically for the Russian…