activity
20242026
collaborators

9 papers

cs.CV2026

ViGoR-Bench: How Far Are Visual Generative Models From Zero-Shot Visual Reasoners?

Haonan Han, Jiancheng Huang, Xiaopeng Sun +7

Beneath the stunning visual fidelity of modern AIGC models lies a "logical desert", where systems fail tasks that require physical, causal, or complex spatial reasoning. Current ev…

cs.CV2026

Forge-and-Quench: Enhancing Image Generation for Higher Fidelity in Unified Multimodal Models

Yanbing Zeng, Jia Wang, Hanghang Ma +4

Integrating image generation and understanding into a single framework has become a pivotal goal in the multimodal domain. However, how understanding can effectively assist generat…

cs.CV2025

MagicMirror: A Large-Scale Dataset and Benchmark for Fine-Grained Artifacts Assessment in Text-to-Image Generation

Jia Wang, Jie Hu, Xiaoqi Ma +3

Text-to-image (T2I) generation has achieved remarkable progress in instruction following and aesthetics. However, a persistent challenge is the prevalence of physical artifacts, su…

cs.CV2025

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models

Huijie Liu, Jingyun Wang, Shuai Ma +3

Motion customization aims to adapt the diffusion model (DM) to generate videos with the motion specified by a set of video clips with the same motion concept. To realize this goal,…

cs.CV2025

Omni-Dish: Photorealistic and Faithful Image Generation and Editing for Arbitrary Chinese Dishes

Huijie Liu, Bingcan Wang, Jie Hu +2

Dish images play a crucial role in the digital era, with the demand for culturally distinctive dish images continuously increasing due to the digitization of the food industry and…

cs.GR2025

Image Editing with Diffusion Models: A Survey

Jia Wang, Jie Hu, Xiaoqi Ma +3

With deeper exploration of diffusion model, developments in the field of image generation have triggered a boom in image creation. As the quality of base-model generated images con…