1 paper
Ryozo Masukawa, Sanggeon Yun, Hyunwoo Oh +8
Recent progress in reinforcement learning with verifiable rewards (RLVR) shows that small, specialized language models (SLMs) can exhibit structured reasoning without relying on la…