1 paper · 1 filter
Ziqin Gong, Ning Li, Huaikang Zhou
Reinforcement-learned reasoning has powered recent AI leaps on verifiable tasks, including mathematics, code, and structure prediction. The harder bottleneck is evaluative judgment…