6 papers
V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning
Haoxiang Sun, Zhihang Yi, Langxuan Deng +6
Fine-grained visual reasoning requires multimodal large language models (MLLMs) to identify task-relevant visual evidence and ground their reasoning in local image regions. Existin…
GraphPO: Graph-based Policy Optimization for Reasoning Models
Yuliang Zhan, Xinyu Tang, Jian Li +7
Reinforcement Learning with Verifiable Rewards (RLVR) has become a standard paradigm for enhancing the capability of large reasoning models. RLVR typically samples responses indepe…
OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models
Hao Sun, Yunyi Shen, Mihaela van der Schaar
In the era of large language models (LLMs), high-quality, domain-rich, and continuously evolving datasets capturing expert-level knowledge, core human values, and reasoning are inc…
Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs
Hao Sun, Yunyi Shen, Jean-Francois Ton +1
Large Language Models (LLMs) have made substantial strides in structured tasks through Reinforcement Learning (RL), demonstrating proficiency in mathematical reasoning and code gen…
Reviving The Classics: Active Reward Modeling in Large Language Model Alignment
Yunyi Shen, Hao Sun, Jean-François Ton
Building neural reward models from human preferences is a pivotal component in reinforcement learning from human feedback (RLHF) and large language model alignment research. Given…
Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives
Hao Sun, Yunyi Shen, Jean-Francois Ton
The Bradley-Terry (BT) model is a common and successful practice in reward modeling for Large Language Model (LLM) alignment. However, it remains unclear why this model -- original…