1 paper · 1 filter
Gengyang Li, Yifeng Gao, Yuming Li +1
While Chain-of-Thought (CoT) prompting improves reasoning in large language models (LLMs), the excessive length of reasoning tokens increases latency and KV cache memory usage, and…