papers

Publications (34)

cs.CV2026

Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback

Huaisong Zhang, Hao Yu, Yuxuan Zhang +7

Despite generating increasingly photorealistic images, text-to-image (T2I) models still exhibit localized, subtle, and structurally complex failures. Diagnosing these failures requ…

cs.CV2022

CORE: Consistent Representation Learning for Face Forgery Detection

Yunsheng Ni, Depu Meng, Changqian Yu +3

Face manipulation techniques develop rapidly and arouse widespread public concerns. Despite that vanilla convolutional neural networks achieve acceptable performance, they suffer f…

cs.CV2021

CondNet: Conditional Classifier for Scene Segmentation

Changqian Yu, Yuanjie Shao, Changxin Gao +1

The fully convolutional network (FCN) has achieved tremendous success in dense visual recognition tasks, such as scene segmentation. The last layer of FCN is typically a global cla…

cs.CV2026

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization

Yifu Luo, Haoyuan Sun, Xinhao Hu +12

Recent Progress in post-training flow matching for text-to-image (T2I) generation with Group Relative Policy Optimization (GRPO) has demonstrated strong potential. However, it is h…

cs.CV2024

Scaling Diffusion Transformers to 16 Billion Parameters

Zhengcong Fei, Mingyuan Fan, Changqian Yu +2

In this paper, we present DiT-MoE, a sparse version of the diffusion Transformer, that is scalable and competitive with dense networks while exhibiting highly optimized inference.…

cs.SD2024

FLUX that Plays Music

Zhengcong Fei, Mingyuan Fan, Changqian Yu +1

This paper explores a simple extension of diffusion-based rectified flow Transformers for text-to-music generation, termed as FluxMusic. Generally, along with design in advanced Fl…