collaborators

27 papers

cs.CV2026

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning

Jiahao Shao, Yuanbo Yang, Yiyi Liao +3

Tool-augmented vision-language models increasingly "think with images": they call crop, zoom, or code tools and reason over the returned pixels. However, recent work using blind te…

cs.CV2026

Representation Forcing for Bottleneck-Free Unified Multimodal Models

Yuqing Wang, Zhijie Lin, Ceyuan Yang +10

Unified multimodal models (UMMs) aim to handle perception and generation in a single model. Yet existing UMMs still rely on a frozen, separately pretrained VAE for image generation…

cs.CV2026

End-to-End Training for Autoregressive Video Diffusion via Self-Resampling

Yuwei Guo, Ceyuan Yang, Hao He +5

Autoregressive video diffusion models hold promise for world simulation but are vulnerable to exposure bias arising from the train-test mismatch. While recent works address this vi…

cs.LG2026

Explicit Critic Guidance for Aligning Diffusion Models

Zhengyang Liang, Qihang Zhang, Ceyuan Yang

Online reinforcement learning is becoming increasingly important for aligning diffusion models with non-differentiable objectives. However, existing methods still face limitations…

cs.LG2026

Adversarial Flow Models

Shanchuan Lin, Ceyuan Yang, Zhijie Lin +2

We present adversarial flow models, a class of generative models that belongs to both the adversarial and flow families. Our method supports native one-step and multi-step generati…

cs.CV2026

Context Unrolling in Omni Models

Ceyuan Yang, Zhijie Lin, Yang Zhao +16

We present Omni, a unified multimodal model natively trained on diverse modalities, including text, images, videos, 3D geometry, and hidden representations. We find that such train…