8 papers · 1 filter
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
Yibin Wang, Zhimin Li, Yuhang Zang +4
Recent advances in multimodal Reward Models (RMs) have shown significant promise in delivering reward signals to align vision models with human preferences. However, current RMs ar…
Unified Reward Model for Multimodal Understanding and Generation
Yibin Wang, Yuhang Zang, Hao Li +2
Recent advances in human preference alignment have significantly improved multimodal generation and understanding. A key approach is to train reward models that provide supervision…
LiFT: Leveraging Human Feedback for Text-to-Video Model Alignment
Yibin Wang, Zhiyu Tan, Junyan Wang +3
Recent advances in text-to-video (T2V) generative models have shown impressive capabilities. However, these models are still inadequate in aligning synthesized videos with human pr…
MagicFace: Training-free Universal-Style Human Image Customized Synthesis
Yibin Wang, Weizhong Zhang, Cheng Jin
Current human image customization methods leverage Stable Diffusion (SD) for its rich semantic prior. However, since SD is not specifically designed for human-oriented generation,…
DreamText: High Fidelity Scene Text Synthesis
Yibin Wang, Weizhong Zhang, Honghui Xu +1
Scene text synthesis involves rendering specified texts onto arbitrary images. Current methods typically formulate this task in an end-to-end manner but lack effective character-le…
PrimeComposer: Faster Progressively Combined Diffusion for Image Composition with Attention Steering
Yibin Wang, Weizhong Zhang, Jianwei Zheng +1
Image composition involves seamlessly integrating given objects into a specific visual context. Current training-free methods rely on composing attention weights from several sampl…