5 papers
PrefGen: Multimodal Preference Learning for Preference-Conditioned Image Generation
Wenyi Mo, Tianyu Zhang, Yalong Bai +3
Preference-conditioned image generation seeks to adapt generative models to individual users, producing outputs that reflect personal aesthetic choices beyond the given textual pro…
Learning User Preferences for Image Generation Model
Wenyi Mo, Ying Ba, Tianyu Zhang +2
User preference prediction requires a comprehensive and accurate understanding of individual tastes. This includes both surface-level attributes, such as color and style, and deepe…
Enhancing Reward Models for High-quality Image Generation: Beyond Text-Image Alignment
Ying Ba, Tianyu Zhang, Yalong Bai +4
Contemporary image generation systems have achieved high fidelity and superior aesthetic quality beyond basic text-image alignment. However, existing evaluation frameworks have fai…
V2Flow: Unifying Visual Tokenization and Large Language Model Vocabularies for Autoregressive Image Generation
Guiwei Zhang, Tianyu Zhang, Mohan Zhou +2
We propose V2Flow, a novel tokenizer that produces discrete visual tokens capable of high-fidelity reconstruction, while ensuring structural and latent distribution alignment with…
Uniform Attention Maps: Boosting Image Fidelity in Reconstruction and Editing
Wenyi Mo, Tianyu Zhang, Yalong Bai +2
Text-guided image generation and editing using diffusion models have achieved remarkable advancements. Among these, tuning-free methods have gained attention for their ability to p…