activity
20242026
collaborators

7 papers

cs.CV2026

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

Sicheng Mo, Yuheng Li, Ziyang Leng +2

Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing a…

cs.CV2026

UniTemp: Unlocking Video Generation in Any Temporal Order via Bidirectional Distillation

Lin Zhang, Sicheng Mo, Zefan Cai +6

Autoregressive video diffusion models have emerged as a promising approach for long video generation, achieving strong performance in streaming settings. However, existing methods…

cs.CV2026

Relational Visual Similarity

Thao Nguyen, Sicheng Mo, Krishna Kumar Singh +6

Humans do not just see attribute similarity -- we also see relational similarity. An apple is like a peach because both are reddish fruit, but the Earth is also like a peach: its c…

cs.CV2025

Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration

Sicheng Mo, Thao Nguyen, Richard Zhang +7

In this work, we explore an untapped signal in diffusion model inference. While all previous methods generate images independently at inference, we instead ask if samples can be ge…

cs.CV2025

Dreamland: Controllable World Creation with Simulator and Generative Models

Sicheng Mo, Ziyang Leng, Leon Liu +3

Large-scale video generative models can synthesize diverse and realistic visual content for dynamic world creation, but they often lack element-wise controllability, hindering thei…

cs.CV2025

X-Fusion: Introducing New Modality to Frozen Large Language Models

Sicheng Mo, Thao Nguyen, Xun Huang +9

We propose X-Fusion, a framework that extends pretrained Large Language Models (LLMs) for multimodal tasks while preserving their language capabilities. X-Fusion employs a dual-tow…