1 paper
Yujia Chen, Yang Ye, Xiao Chu +2
Reinforcement learning (RL) with verifiable rewards has proven effective at post-training LLMs for coding, yet deploying separate task-specific specialists incurs costs that scale…