1 paper
Xinyu Qiu, Heng Jia, Zhengwen Zeng +4
Parallel test-time scaling typically trains separate generation and verification models, incurring high training and inference costs. We propose Advantage Decoupled Preference Opti…