activity
20242026
collaborators

12 papers

cs.CV2026

Olaf-World: Orienting Latent Actions for Video World Modeling

Yuxin Jiang, Yuchao Gu, Ivor W. Tsang +1

Scaling action-controllable world models is limited by the scarcity of action labels. While latent action learning promises to extract control interfaces from unlabeled video, lear…

cs.CV2026

MIND: Benchmarking Memory Consistency and Action Control in World Models

Yixuan Ye, Xuanyu Lu, Yuxin Jiang +7

World models aim to understand, remember, and predict dynamic visual environments, yet a unified benchmark for evaluating their fundamental abilities remains lacking. To address th…

cs.CV2025

Personalized Vision via Visual In-Context Learning

Yuxin Jiang, Yuchao Gu, Yiren Song +2

Modern vision models, trained on large-scale annotated datasets, excel at predefined tasks but struggle with personalized vision -- tasks defined at test time by users with customi…

cs.CV2025

UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models

Lan Chen, Yuchao Gu, Qi Mao

Large language models, trained on extensive corpora, successfully unify diverse linguistic tasks within a single generative framework. Inspired by this, recent works like Large Vis…

cs.CV2025

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Jinheng Xie, Weijia Mao, Zechen Bai +7

We present a unified transformer, i.e., Show-o, that unifies multimodal understanding and generation. Unlike fully autoregressive models, Show-o unifies autoregressive and (discret…

cs.GR2025

FramePrompt: In-context Controllable Animation with Zero Structural Changes

Guian Fang, Yuchao Gu, Mike Zheng Shou

Generating controllable character animation from a reference image and motion guidance remains a challenging task due to the inherent difficulty of injecting appearance and motion…