1 paper
Zijun Weng, Xiaohui Hu, Shuangyong Song +3
Open-ended post-training benefits from rewards that make prompt-specific success conditions explicit, rather than relying only on post-hoc scalar scores. In instruction following,…