1 paper
Yiren Zheng, Shibo Li, Jiaming Liu +2
Current multimodal approaches predominantly treat visual generation as an external process, relying on pixel rendering or code execution, thereby overlooking the native visual repr…