10 papers · 1 filter
Twins: Learn to Predict Unified Representations with Focal Loss
Kaixiong Gong, Xin Cai, Bin Lin +9
Unified multimodal models seek a shared visual token space that supports both multimodal understanding and image generation. Discrete methods unify the interface via a shared codeb…
Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models
Junhao Liu, Jian-Wei Zhang, Tao Huang +3
Unified Multimodal Large Language Models (MLLMs) offer a promising paradigm for unifying visual understanding and generation, yet they still struggle to follow complex spatial inst…
Rosetta: Composable Native Multimodal Pretraining
Xiangyue Liu, Zijian Zhang, Miles Yang +3
Achieving true artificial general intelligence requires foundation models capable of integrating new modalities without forgetting prior knowledge. However, accommodating continuou…
GEAR: Guided End-to-End AutoRegression for Image Synthesis
Bin Lin, Zheyuan Liu, Chenguo Lin +8
Visual generative models are typically trained in two stages. A tokenizer is first trained for reconstruction and then frozen, after which a generator is trained on its discrete in…
Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding
Xiangyue Liu, Zijian Zhang, Miles Yang +3
Empowering Large Multimodal Models (LMMs) with image generation often leads to catastrophic forgetting in understanding tasks due to severe gradient conflicts. While existing parad…
HunyuanImage 3.0 Technical Report
Tencent Hunyuan Foundation Model Team
We present HunyuanImage 3.0, a native multimodal model that unifies multimodal understanding and generation within an autoregressive framework, with its image generation module pub…