2 papers
cs.CL2025
T-pro 2.0: An Efficient Russian Hybrid-Reasoning Model and Playground
Dmitrii Stoianov, Danil Taranets, Olga Tsymboi +12
We introduce T-pro 2.0, an open-weight Russian LLM for hybrid reasoning and efficient inference. The model supports direct answering and reasoning-trace generation, using a Cyrilli…
cs.LG2025
Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models
Alina Shutova, Vladimir Malinovskii, Vage Egiazarian +5
Efficient real-world deployments of large language models (LLMs) rely on Key-Value (KV) caching for processing and generating long outputs, reducing the need for repetitive computa…