27 papers
Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning
Jiahao Shao, Yuanbo Yang, Yiyi Liao +3
Tool-augmented vision-language models increasingly "think with images": they call crop, zoom, or code tools and reason over the returned pixels. However, recent work using blind te…
Representation Forcing for Bottleneck-Free Unified Multimodal Models
Yuqing Wang, Zhijie Lin, Ceyuan Yang +10
Unified multimodal models (UMMs) aim to handle perception and generation in a single model. Yet existing UMMs still rely on a frozen, separately pretrained VAE for image generation…
End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
Yuwei Guo, Ceyuan Yang, Hao He +5
Autoregressive video diffusion models hold promise for world simulation but are vulnerable to exposure bias arising from the train-test mismatch. While recent works address this vi…
Explicit Critic Guidance for Aligning Diffusion Models
Zhengyang Liang, Qihang Zhang, Ceyuan Yang
Online reinforcement learning is becoming increasingly important for aligning diffusion models with non-differentiable objectives. However, existing methods still face limitations…
Adversarial Flow Models
Shanchuan Lin, Ceyuan Yang, Zhijie Lin +2
We present adversarial flow models, a class of generative models that belongs to both the adversarial and flow families. Our method supports native one-step and multi-step generati…
Context Unrolling in Omni Models
Ceyuan Yang, Zhijie Lin, Yang Zhao +16
We present Omni, a unified multimodal model natively trained on diverse modalities, including text, images, videos, 3D geometry, and hidden representations. We find that such train…