collaborators

6 papers

cs.CV2026

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning

Haoxiang Sun, Zhihang Yi, Langxuan Deng +6

Fine-grained visual reasoning requires multimodal large language models (MLLMs) to identify task-relevant visual evidence and ground their reasoning in local image regions. Existin…

cs.CL2026

GraphPO: Graph-based Policy Optimization for Reasoning Models

Yuliang Zhan, Xinyu Tang, Jian Li +7

Reinforcement Learning with Verifiable Rewards (RLVR) has become a standard paradigm for enhancing the capability of large reasoning models. RLVR typically samples responses indepe…

cs.CY2025

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models

Hao Sun, Yunyi Shen, Mihaela van der Schaar

In the era of large language models (LLMs), high-quality, domain-rich, and continuously evolving datasets capturing expert-level knowledge, core human values, and reasoning are inc…

cs.CL2025

Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs

Hao Sun, Yunyi Shen, Jean-Francois Ton +1

Large Language Models (LLMs) have made substantial strides in structured tasks through Reinforcement Learning (RL), demonstrating proficiency in mathematical reasoning and code gen…

cs.CL2025

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment

Yunyi Shen, Hao Sun, Jean-François Ton

Building neural reward models from human preferences is a pivotal component in reinforcement learning from human feedback (RLHF) and large language model alignment research. Given…

cs.AI2025

Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Hao Sun, Yunyi Shen, Jean-Francois Ton

The Bradley-Terry (BT) model is a common and successful practice in reward modeling for Large Language Model (LLM) alignment. However, it remains unclear why this model -- original…