1 paper · 1 filter
Yifu Han, Geo Zhang
This study investigates the effectiveness of reinforcement learning (RL) fine-tuning techniques on a compact language model (Qwen2.5-0.5B Base) for two challenging tasks: instructi…