1 paper
Xinge Ye, Rui Wang, Yuchuan Wu +4
Reinforcement Learning Fine-Tuning (RLFT) has achieved notable success in tasks with objectively verifiable answers (e.g., code generation, mathematical reasoning), yet struggles w…