1 citations · 1 across the 1 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Compactor: Calibrated Query-Agnostic KV Cache Compression with Approximate Leverage Scores
Vivek Chari, Benjamin Van Durme
Modern Large Language Models (LLMs) are increasingly trained to support very large context windows. We present Compactor, a training-free, query-agnostic KV compression strategy th…
cs.CL2025
KV-Distill: Nearly Lossless Learnable Context Compression for LLMs
Vivek Chari, Guanghui Qin, Benjamin Van Durme
Sequence-to-sequence tasks often benefit from long contexts, but the quadratic complexity of self-attention in standard Transformers renders this non-trivial. During generation, te…