2 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.CV2026
VoT: Vision-of-Thought for Unified Multimodal Representation Alignment
Jingxiang Sun, Chao Liao, Zhengxiong Luo +6
Current text-to-image systems typically employ a "text encoder plus diffusion decoder" paradigm, in which text semantics directly modulate continuous latent noise. Despite their su…
cs.CV2025★ 2 cited
Seedream 4.0: Toward Next-generation Multimodal Image Generation
Team Seedream, :, Yunpeng Chen +48
We introduce Seedream 4.0, an efficient and high-performance multimodal image generation system that unifies text-to-image (T2I) synthesis, image editing, and multi-image compositi…
cs.CV2025
Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
Chao Liao, Liyang Liu, Xun Wang +7
Recent progress in unified models for image understanding and generation has been impressive, yet most approaches remain limited to single-modal generation conditioned on multiple…