3 papers
cs.CL2025
Value-Guided KV Compression for LLMs via Approximated CUR Decomposition
Ayan Sengupta, Siddhant Chaudhary, Tanmoy Chakraborty
Key-value (KV) cache compression has emerged as a critical technique for reducing the memory and latency overhead of autoregressive language models during inference. Prior approach…
math.FA2025
Order-preserving unique Hahn-Banach extensions
Tanmoy Paul, T. S. S. R. K. Rao
Let be a real Banach lattice with a unit, and let be a closed subspace containing the unit. In this paper, we study the order-theoretic (also isometric) structu…
cs.CL2025
You Only Prune Once: Designing Calibration-Free Model Compression With Policy Learning
Ayan Sengupta, Siddhant Chaudhary, Tanmoy Chakraborty
The ever-increasing size of large language models (LLMs) presents significant challenges for deployment due to their heavy computational and memory requirements. Current model prun…