activity
20232026
collaborators

6 papers

cs.CL2026

BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models

Liubov Chubarova, Alexandra Kuleshova, Daniil Volkov +2

While Multimodal Large Language Models (MLLMs) have made significant strides in visual comprehension, their ability to reason about text-dense, professional documents remains incom…

cs.LG2025

Selective Adversarial Attacks on LLM Benchmarks

Ivan Dubrovsky, Anastasia Orlova, Illarion Iov +3

Benchmarking outcomes increasingly govern trust, selection, and deployment of LLMs, yet these evaluations remain vulnerable to semantically equivalent adversarial perturbations. Pr…

cs.CL2025

Language steering in latent space to mitigate unintended code-switching

Andrey Goncharov, Nikolai Kondusov, Alexey Zaytsev

Multilingual Large Language Models (LLMs) often exhibit hallucinations such as unintended code-switching, reducing reliability in downstream tasks. We propose latent-space language…

cs.LG2025

Complexity-aware fine-tuning

Andrey Goncharov, Daniil Vyazhev, Petr Sychev +2

General-purpose Large Language Models (LLMs) are frequently fine-tuned through supervised fine-tuning (SFT) to enhance performance in specific domains. Better results can be achiev…

cs.CL2025

When an LLM is apprehensive about its answers -- and when its uncertainty is justified

Petr Sychev, Andrey Goncharov, Daniil Vyazhev +2

Uncertainty estimation is crucial for evaluating Large Language Models (LLMs), particularly in high-stakes domains where incorrect answers result in significant consequences. Numer…

cs.LG2023

RepQ: Generalizing Quantization-Aware Training for Re-Parametrized Architectures

Anastasiia Prutianova, Alexey Zaytsev, Chung-Kuei Lee +2

Existing neural networks are memory-consuming and computationally intensive, making deploying them challenging in resource-constrained environments. However, there are various meth…