2 papers
cs.CL2026
Reformulating KV Cache Eviction Problem for Long-Context LLM Inference
Tho Mai, Joo-Young Kim
Large language models (LLMs) support long-context inference but suffer from substantial memory and runtime overhead due to Key-Value (KV) Cache growth. Existing KV Cache eviction m…
cs.AI2026
POP: Online Structural Pruning Enables Efficient Inference of Large Foundation Models
Yi Chen, Wonjin Shin, Shuhong Liu +6
Large foundation models (LFMs) achieve strong performance through scaling, yet current structural pruning methods derive fixed pruning decisions during inference, overlooking spars…