5 papers
MASS-DPO: Multi-negative Active Sample Selection for Direct Policy Optimization
Rohan Surana, Xintong Li, Sheldon Yu +7
Multi-negative preference optimization under the Plackett--Luce (PL) model extends Direct Preference Optimization (DPO) by leveraging comparative signals across one preferred and m…
RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees
Yichen Xu, Yuanhang Liu, Chuhan Wang +5
While Multimodal Large Language Models (MLLMs) excel at generic video understanding, their ability to support specialized, rule-grounded decision-making remains insufficiently expl…
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning
Rohan Surana, Gagan Mundada, Xunyi Jiang +19
Reinforcement learning (RL) has become a central post-training tool for improving the reasoning abilities of large language models (LLMs). In these systems, the rollout, the trajec…
SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes
Chuhan Wang, Xintong Li, Jennifer Yuntong Zhang +5
Multimodal large language models often struggle with faithful reasoning in complex visual scenes, where intricate entities and relations require precise visual grounding at each st…
Importance Sampling for Multi-Negative Multimodal Direct Preference Optimization
Xintong Li, Chuhan Wang, Junda Wu +4
Direct Preference Optimization (DPO) has recently been extended from text-only models to vision-language models. However, existing methods rely on oversimplified pairwise compariso…