1 paper · 1 filter
Siwei Chen, Siqi Chen, Xupeng Miao +1
Recent large reasoning models often develop long chain-of-thought responses during reinforcement learning (RL), resulting in high inference latency and deployment cost. Existing me…