42 citations · 184 across the 29 of their papers we have counts for
46 papers · 1 filter
VoT: Vision-of-Thought for Unified Multimodal Representation Alignment
Jingxiang Sun, Chao Liao, Zhengxiong Luo +6
Current text-to-image systems typically employ a "text encoder plus diffusion decoder" paradigm, in which text semantics directly modulate continuous latent noise. Despite their su…
Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning
Jiahao Shao, Yuanbo Yang, Yiyi Liao +3
Tool-augmented vision-language models increasingly "think with images": they call crop, zoom, or code tools and reason over the returned pixels. However, recent work using blind te…
Representation Forcing for Bottleneck-Free Unified Multimodal Models
Yuqing Wang, Zhijie Lin, Ceyuan Yang +10
Unified multimodal models (UMMs) aim to handle perception and generation in a single model. Yet existing UMMs still rely on a frozen, separately pretrained VAE for image generation…
Context Unrolling in Omni Models
Ceyuan Yang, Zhijie Lin, Yang Zhao +16
We present Omni, a unified multimodal model natively trained on diverse modalities, including text, images, videos, 3D geometry, and hidden representations. We find that such train…
VTok: A Unified Video Tokenizer with Decoupled Spatial-Temporal Latents
Feng Wang, Yichun Shi, Ceyuan Yang +4
This work presents VTok, a unified video tokenization framework that can be used for both generation and understanding tasks. Unlike the leading vision-language systems that tokeni…
Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model
Team Seedance, Heyi Chen, Siyan Chen +194
Recent strides in video generation have paved the way for unified audio-visual generation. In this work, we present Seedance 1.5 pro, a foundational model engineered specifically f…