activity
20242026
collaborators

6 papers

cs.CV2026

Unified Reward Model for Multimodal Understanding and Generation

Yibin Wang, Yuhang Zang, Hao Li +2

Recent advances in human preference alignment have significantly improved multimodal generation and understanding. A key approach is to train reward models that provide supervision…

cs.CV2025

Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning

Yibin Wang, Zhimin Li, Yuhang Zang +4

Recent advances in multimodal Reward Models (RMs) have shown significant promise in delivering reward signals to align vision models with human preferences. However, current RMs ar…

cs.CV2025

Enhancing Object Coherence in Layout-to-Image Synthesis

Yibin Wang, Changhai Zhou, Honghui Xu

Layout-to-image synthesis is an emerging technique in conditional image generation. It aims to generate complex scenes, where users require fine control over the layout of the obje…

cs.CV2025

DreamText: High Fidelity Scene Text Synthesis

Yibin Wang, Weizhong Zhang, Honghui Xu +1

Scene text synthesis involves rendering specified texts onto arbitrary images. Current methods typically formulate this task in an end-to-end manner but lack effective character-le…

cs.CV2025

LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment

Yibin Wang, Zhiyu Tan, Junyan Wang +3

Recent advances in text-to-video (T2V) generative models have shown impressive capabilities. However, these models are still inadequate in aligning synthesized videos with human pr…

cs.CV2024

MagicFace: Training-free Universal-Style Human Image Customized Synthesis

Yibin Wang, Weizhong Zhang, Cheng Jin

Current human image customization methods leverage Stable Diffusion (SD) for its rich semantic prior. However, since SD is not specifically designed for human-oriented generation,…