1 paper · 1 filter
Xinge Ye, Rui Wang, Yuchuan Wu +4
Reinforcement Learning Fine-Tuning (RLFT) has achieved notable success in tasks with objectively verifiable answers (e.g., code generation, mathematical reasoning), yet struggles w…