13 papers
Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models
Yuanhao Ban, Jiaqi Feng, Hengguang Zhou +3
Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objectives optimize local frame prediction rather than the geometry an…
Self-Evolving Visual Questioner
Yijun Liang, Hengguang Zhou, Ming Li +3
Vision-language models (VLMs) are typically trained as passive answerers, while their ability to actively ask diverse, non-trivial, visual-centric and grounded questions remains un…
When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
Tong Xie, Andrew Bai, Yuanhao Ban +3
Reward models are central to Large Language Model (LLM) alignment within the framework of RLHF. The standard objective used in reward modeling is the Bradley-Terry (BT) loss, which…
One-Forcing: Towards Stable One-Step Autoregressive Video Generation
Jiaqi Feng, Justin Cui, Yuanhao Ban +1
Recent advances have substantially improved real-time interactive video generation in the autoregressive regime. However, most existing few-step autoregressive video generation met…
AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment
Kuei-Chun Kao, Daixuan Huo, Yuanhao Ban +1
Aligning Text-to-Image (T2I) generation models with human preferences increasingly relies on image reward models that score or rank generated images according to prompt alignment a…
IRIS: Intrinsic Reward Image Synthesis
Yihang Chen, Yuanhao Ban, Yunqi Hong +1
Despite the success of Reinforcement Learning from Human Feedback (RLHF) in language reasoning, its application to autoregressive Text-to-Image (T2I) generation is often constraine…