1 paper
Hanbing Liu, Lang Cao, Yuanyi Ren +5
Large language models (LLMs) show strong reasoning abilities but often produce unnecessarily long explanations that reduce efficiency. Although reinforcement learning (RL) has been…