6 citations · 15 across the 18 of their papers we have counts for
14 papers · 1 filter
StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training
Bao Tang, Jiahao Guo, Haoxiang Cao +4
Vector Quantization (VQ) is fundamental to discrete visual tokenizers that power modern autoregressive and masked image generation models. While recent shared-projection codebook m…
Dior: Drawing the Light of Image via Material-Decoupled Illumination Representation
Xuanpu Zhang, Xuesong Niu, Haoxiang Cao +3
Controllable image relighting is an important problem in image editing, and hand-drawn scribbles provide an intuitive interface for specifying the desired illumination. However, ex…
AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss
Mingju Gao, Jingkai Zhou, Kun Gai +2
Fréchet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-level diffusion and flow-match…
Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback
Huaisong Zhang, Hao Yu, Yuxuan Zhang +7
Despite generating increasingly photorealistic images, text-to-image (T2I) models still exhibit localized, subtle, and structurally complex failures. Diagnosing these failures requ…
MaskAlign: Token-Subset Representation Alignment for Efficient Diffusion Training
Lianyu Pang, Tianlin Pan, Cheng Da +5
Representation alignment with pretrained vision models has recently shown strong potential for accelerating diffusion transformer training. By aligning intermediate diffusion featu…
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization
Yifu Luo, Haoyuan Sun, Xinhao Hu +12
Recent Progress in post-training flow matching for text-to-image (T2I) generation with Group Relative Policy Optimization (GRPO) has demonstrated strong potential. However, it is h…