1 paper
Fuxiang Zhai, Sixiang Chen, Yingjin Li +4
Unified multimodal models can encode visual understanding and image generation within a shared backbone, yet understanding does not automatically translate into control: models may…