1 paper
Yirou Ge, Yixi Li, Alec Chiu +8
Large language models (LLMs) are increasingly being deployed in cost and latency-sensitive settings. While chain-of-thought improves reasoning, it can waste tokens on simple reques…