8 papers
BAS-OPD: Budget-Aware Selective On-Policy Self-Distillation for Fine-Grained Multimodal Perception
Zihan Chen, Hengguang Zhou, Yuan Kang +5
Multimodal large language models (MLLMs) often struggle with fine-grained visual perception when processing complete images, as critical evidence may only appear in local regions.…
Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models
Yuanhao Ban, Jiaqi Feng, Hengguang Zhou +3
Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objectives optimize local frame prediction rather than the geometry an…
Self-Evolving Visual Questioner
Yijun Liang, Hengguang Zhou, Ming Li +3
Vision-language models (VLMs) are typically trained as passive answerers, while their ability to actively ask diverse, non-trivial, visual-centric and grounded questions remains un…
Do Prompt-Elicited Trajectories Reflect Training-Time Reward Hacking? A Systematic Study on Monitoring Training-Time Reward Hacking in Code Generation
Lichen Li, Hengguang Zhou, Yijun Liang +2
Reward hacking in code generation, where models exploit evaluation loopholes to obtain high reward without correctly solving the intended task, poses a critical challenge for Reinf…
Understanding Reward Hacking in Text-to-Image Reinforcement Learning
Yunqi Hong, Kuei-Chun Kao, Hengguang Zhou +1
Reinforcement learning (RL) has become a standard approach for post-training large language models and, more recently, for improving image generation models, which uses reward func…
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
Zihan Chen, Yiming Zhang, Hengguang Zhou +3
Current benchmarks are inadequate for evaluating progress in reinforcement learning (RL) for large language models (LLMs).Despite recent benchmark gains reported for RL, we find th…