Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction
Enshuai Zhou, Yifan Hao, Chao Wang +7
Long-context inference in Large Language Models (LLMs) is bottlenecked by the linear growth of Key-Value (KV) cache memory. Existing KV cache compression paradigms are fundamentall…
cs.LG2025
Outlier Smoothing with Closed-Form Rotations for W4A4 Large Language Model Quantization
Jinying Xiao, Bin Ji, Shasha Li +8
Large Language Models (LLMs) quantization facilitates deploying LLMs in resource-limited settings, but existing methods that combine incompatible gradient optimization and quantiza…