2 papers
cs.CL2026
Open Rubric System: Scaling Reinforcement Learning with Pairwise Adaptive Rubric
Ruipeng Jia, Yunyi Yang, Wen Wang +7
Scalar reward models compress multi-dimensional human preferences into a single opaque score, creating an information bottleneck that often leads to brittleness and reward hacking…
cs.CL2025
Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards
Ruipeng Jia, Yunyi Yang, Yongbo Gai +5
Reinforcement learning with verifiable rewards (RLVR) has enabled large language models (LLMs) to achieve remarkable breakthroughs in reasoning tasks with objective ground-truth an…