2 papers
cs.LG2026
When Does Value-Aware KV Eviction Help? A Fixed-Contract Diagnostic for Non-Monotone Cache Compression
Ruijie Zhang, Haozhe Liang, Da Chang +4
Long-context LLM inference is bottlenecked by the memory and bandwidth cost of reading large KV caches during decoding. KV compression reduces this cost by keeping only part of the…
cs.RO2026
Aegis: Automated Error Generation and Attribution for Multi-Agent Systems
Fanqi Kong, Ruijie Zhang, Huaxiao Yin +7
Large language model based multi-agent systems (MAS) have unlocked significant advancements in tackling complex problems, but their increasing capability introduces a structural fr…