5 papers · 1 filter
How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing
Huanyu Zhang, Xuehai Bai, Chengzu Li +9
Recent generative models have achieved remarkable progress in image editing. However, existing systems and benchmarks remain largely text-guided. In contrast, human communication i…
Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?
Xinchen Yan, Chen Liang, Lijun Yu +3
This paper investigates the scaling properties of autoregressive next-pixel prediction, a simple, end-to-end yet under-explored framework for unified vision models. Starting with i…
CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation
Chengzhuo Tong, Mingkun Chang, Shenglong Zhang +12
Recent video generation models have revealed the emergence of Chain-of-Frame (CoF) reasoning, enabling frame-by-frame visual inference. With this capability, video models have been…
In-Context LoRA for Diffusion Transformers
Lianghua Huang, Wei Wang, Zhi-Fan Wu +6
Recent research arXiv:2410.15027 has explored the use of diffusion transformers (DiTs) for task-agnostic image generation by simply concatenating attention tokens across images. Ho…
Group Diffusion Transformers are Unsupervised Multitask Learners
Lianghua Huang, Wei Wang, Zhi-Fan Wu +6
While large language models (LLMs) have revolutionized natural language processing with their task-agnostic capabilities, visual generation tasks such as image translation, style t…