1 paper
Mingxuan Xia, Yuhang Yang, Chao Ye +7
Rubric-based RL has recently shown promise in improving LLMs on open-ended tasks. A widely recognized limitation of rubric-based RL is limited exploration: criteria that no rollout…