4 papers
MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale
Zhicong Tang, Zhao Zhang, Jingye Chen +6
Layered image generation and editing is a fundamental capability that enables layer-wise reuse, editing, and composition of generated visual content, analogous to word-level editin…
Pareto-Guided Optimal Transport for Multi-Reward Alignment
Ying Ba, Tianyu Zhang, Mohan Zhou +5
Text-to-image generation models have achieved remarkable progress in preference optimization, yet achieving robust alignment across diverse reward models remains a significant chal…
V2Flow: Unifying Visual Tokenization and Large Language Model Vocabularies for Autoregressive Image Generation
Guiwei Zhang, Tianyu Zhang, Mohan Zhou +2
We propose V2Flow, a novel tokenizer that produces discrete visual tokens capable of high-fidelity reconstruction, while ensuring structural and latent distribution alignment with…
STAR: Scale-wise Text-conditioned AutoRegressive image generation
Xiaoxiao Ma, Mohan Zhou, Tao Liang +5
We introduce STAR, a text-to-image model that employs a scale-wise auto-regressive paradigm. Unlike VAR, which is constrained to class-conditioned synthesis for images up to 256$\t…