9 papers
Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist
Yuanhao Ban, Tong Xie, Sohyun An +6
Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) models. Existing faithfulness ben…
A Unifying Lens on Supervised Fine-Tuning Through Target Distribution Design
Tong Xie, Yuanhao Ban, Yunqi Hong +3
Supervised fine-tuning (SFT) typically maximizes the likelihood of every token in a demonstrated trajectory. However, an observed token can be non-unique, noisy, or misaligned with…
When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
Tong Xie, Andrew Bai, Yuanhao Ban +3
Reward models are central to Large Language Model (LLM) alignment within the framework of RLHF. The standard objective used in reward modeling is the Bradley-Terry (BT) loss, which…
One-Forcing: Towards Stable One-Step Autoregressive Video Generation
Jiaqi Feng, Justin Cui, Yuanhao Ban +1
Recent advances have substantially improved real-time interactive video generation in the autoregressive regime. However, most existing few-step autoregressive video generation met…
AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment
Kuei-Chun Kao, Daixuan Huo, Yuanhao Ban +1
Aligning Text-to-Image (T2I) generation models with human preferences increasingly relies on image reward models that score or rank generated images according to prompt alignment a…
Reward-Forcing: Autoregressive Video Generation with Reward Feedback
Jingran Zhang, Ning Li, Yuanhao Ban +2
While most prior work in video generation relies on bidirectional architectures, recent efforts have sought to adapt these models into autoregressive variants to support near real-…