1 paper
Shen Han, Yuyang Wu, Junpu Yu +1
Reasoning language models often generate long chain-of-thought (CoT), which accumulates a massive KV cache during the decoding phase and incurs high decoding latency and limited th…