activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale

Zhicong Tang, Zhao Zhang, Jingye Chen +6

Layered image generation and editing is a fundamental capability that enables layer-wise reuse, editing, and composition of generated visual content, analogous to word-level editin…

cs.CV2026

Pareto-Guided Optimal Transport for Multi-Reward Alignment

Ying Ba, Tianyu Zhang, Mohan Zhou +5

Text-to-image generation models have achieved remarkable progress in preference optimization, yet achieving robust alignment across diverse reward models remains a significant chal…

cs.CV2025

V2Flow: Unifying Visual Tokenization and Large Language Model Vocabularies for Autoregressive Image Generation

Guiwei Zhang, Tianyu Zhang, Mohan Zhou +2

We propose V2Flow, a novel tokenizer that produces discrete visual tokens capable of high-fidelity reconstruction, while ensuring structural and latent distribution alignment with…

cs.CV2025

STAR: Scale-wise Text-conditioned AutoRegressive image generation

Xiaoxiao Ma, Mohan Zhou, Tao Liang +5

We introduce STAR, a text-to-image model that employs a scale-wise auto-regressive paradigm. Unlike VAR, which is constrained to class-conditioned synthesis for images up to 256$\t…

cs.CV2024

StyleInject: Parameter Efficient Tuning of Text-to-Image Diffusion Models

Mohan Zhou, Yalong Bai, Qing Yang +1

The ability to fine-tune generative models for text-to-image generation tasks is crucial, particularly facing the complexity involved in accurately interpreting and visualizing tex…