2 papers
cs.LG2026
Every Cache Entry Earns Its Place: Global Allocation of Resolution and Coverage for KV Cache Compression
Haolin Tian, Yuzhe Liu, Tonghan Wang
As large language models (LLMs) process increasingly long contexts, KV cache storage and repeated access have become a major bottleneck. Existing KV cache compression methods rely…
cs.MA2026
A Two-Tier Perspective on Inference-Time Parallelism in Multi-Agent LLM Systems
Zihan Xu, Haolin Tian, Hai Jiang
Large language model (LLM)-driven multi-agent systems typically require multiple model invocations and complex coordination during inference, and their execution strategies directl…