3 papers
cs.CL2025
Iterative Layer-wise Distillation for Efficient Compression of Large Language Models
Grigory Kovalev, Mikhail Tikhomirov
This work investigates distillation methods for large language models (LLMs) with the goal of developing compact models that preserve high performance. Several existing approaches…
cs.IR2025
Wikipedia-based Datasets in Russian Information Retrieval Benchmark RusBEIR
Grigory Kovalev, Natalia Loukachevitch, Mikhail Tikhomirov +2
In this paper, we present a novel series of Russian information retrieval datasets constructed from the "Did you know..." section of Russian Wikipedia. Our datasets support a range…
cs.IR2025
Building Russian Benchmark for Evaluation of Information Retrieval Models
Grigory Kovalev, Mikhail Tikhomirov, Evgeny Kozhevnikov +2
We introduce RusBEIR, a comprehensive benchmark designed for zero-shot evaluation of information retrieval (IR) models in the Russian language. Comprising 17 datasets from various…