1 paper
Yen-Chieh Huang, Pi-Cheng Hsiu, Rui Fang +1
Long-context LLM inference is bottlenecked by the quadratic attention complexity and linear KV cache growth. Prior approaches mitigate this via post-hoc selection or eviction but o…