4 papers
NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
Huichao Zhang, Liao Qu, Yiheng Liu +33
We present NextFlow, a unified decoder-only autoregressive transformer trained on 6 trillion interleaved text-image discrete tokens. By leveraging a unified vision representation w…
DreamO: A Unified Framework for Image Customization
Chong Mou, Yanze Wu, Wenxu Wu +15
Recently, extensive research on image customization (e.g., identity, subject, style, background, etc.) demonstrates strong customization capabilities in large-scale generative mode…
RealCustom++: Representing Images as Real Textual Word for Real-Time Customization
Zhendong Mao, Mengqi Huang, Fei Ding +3
Given a text and an image of a specific subject, text-to-image customization aims to generate new images that align with both the text and the subject's appearance. Existing works…
Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset
Zhuowei Chen, Bingchuan Li, Tianxiang Ma +8
Subject-to-video generation has witnessed substantial progress in recent years. However, existing models still face significant challenges in faithfully following textual instructi…