1 paper
Yuxuan Yang, Feiyang Ren, Bowen Zeng +4
Long-context LLM inference faces a fundamental conflict: head-adaptive compression algorithms (e.g., Top-p nucleus sampling) offer superior accuracy by dynamically fluctuating me…