1 paper · 1 filter
Aojie Yuan, Tianqi Shen, Dajun Zhang
Reasoning LLMs produce thousands of chain-of-thought tokens whose KV cache must reside in scarce GPU HBM. The dominant response -- permanently evicting low-importance tokens -- is…