22 papers
GR2 Technical Report
Yufei Li, Zaiwei Zhang, Mingfu Liang +67
Industrial recommendation systems serve billions of users through a multi-stage funnel -- retrieval, early-stage ranking, and re-ranking -- where the final re-ranking step dispropo…
Dual-Uncertainty Guided Policy Learning for Multimodal Reasoning
Rui Liu, Dian Yu, Tong Zheng +8
Reinforcement learning with verifiable rewards (RLVR) has advanced reasoning capabilities in multimodal large language models. However, existing methods typically treat visual inpu…
One Token to Fool LLM-as-a-Judge
Yulai Zhao, Haolin Liu, Dian Yu +4
Large language models (LLMs) are increasingly trusted as automated judges, assisting evaluation and providing reward signals for training other models, particularly in reference-ba…
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis
Yucheng Shi, Zhenwen Liang, Kishan Panaganti +3
We pursue a vision for self-improving language models in which the model does not merely generate problems or traces to imitate, but constructs the environments that train it. In z…
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification
Rui Liu, Dian Yu, Zhenwen Liang +6
Aligning Multimodal Large Language Models (MLLMs) requires reliable reward models, yet existing single-step evaluators can suffer from lazy judging, exploiting language priors over…
Reinforcing Multimodal Reasoning Against Visual Degradation
Rui Liu, Dian Yu, Haolin Liu +6
Reinforcement Learning has significantly advanced the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet the resulting policies remain brittle against real-wor…