2 papers
cs.AI2026
Minima-KV: Retention-Preserving KV Cache Compression with Mixed-Format Paged Attention
Sergii Kozyrev, Davyd Maiboroda
The key-value (KV) cache is a primary capacity and bandwidth bottleneck in long-context LLM serving. We present Minima-KV, a retention-preserving hierarchy for mixed-format paged a…
cs.LG2026
A Practical Tensor-Network Compression Pipeline for Production-Scale Large Language Models
Sergii Kozyrev, Davyd Maiboroda
Large language models are limited in deployment by GPU memory and inference latency. We present Minima, a production compression pipeline that learns where and how to structurally…