collaborators

5 papers

cs.CV2026

Grounding the Score: Explicit Visual Premise Verification for Reliable Vision-Language Process Reward Models

Junxin Wang, Dai Guan, Weijie Qiu +7

Vision-language process reward models (VL-PRMs) are increasingly used to score intermediate reasoning steps and rerank candidates under test-time scaling. However, they often funct…

cs.CV2026

Rationale Matters: Learning Transferable Rubrics via Proxy-Guided Critique for VLM Reward Models

Weijie Qiu, Dai Guan, Junxin Wang +6

Generative reward models (GRMs) for vision-language models (VLMs) often evaluate outputs via a three-stage pipeline: rubric generation, criterion-based scoring, and a final verdict…

cs.LG2026

CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR

Sijia Cui, Pengyu Cheng, Jiajun Song +6

Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced the reasoning capacity of Large Language Models (LLMs). However, RLVR solely relies on final answer…

cs.CL2026

Open Rubric System: Scaling Reinforcement Learning with Pairwise Adaptive Rubric

Ruipeng Jia, Yunyi Yang, Wen Wang +7

Scalar reward models compress multi-dimensional human preferences into a single opaque score, creating an information bottleneck that often leads to brittleness and reward hacking…

cs.CL2025

Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards

Ruipeng Jia, Yunyi Yang, Yongbo Gai +5

Reinforcement learning with verifiable rewards (RLVR) has enabled large language models (LLMs) to achieve remarkable breakthroughs in reasoning tasks with objective ground-truth an…