1 paper
Fabio Valerio Massoli, Andrey Kuzmin, Arash Behboodi
\ac{CoT} prompting improves LLM accuracy on complex tasks but often increases token usage and inference cost. Existing ``Budget Forcing'' methods reduce cost via fine-tuning with h…