collaborators

7 papers

cs.CV2025

Personalized Vision via Visual In-Context Learning

Yuxin Jiang, Yuchao Gu, Yiren Song +2

Modern vision models, trained on large-scale annotated datasets, excel at predefined tasks but struggle with personalized vision -- tasks defined at test time by users with customi…

cs.CV2025

UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models

Lan Chen, Yuchao Gu, Qi Mao

Large language models, trained on extensive corpora, successfully unify diverse linguistic tasks within a single generative framework. Inspired by this, recent works like Large Vis…

cs.GR2025

FramePrompt: In-context Controllable Animation with Zero Structural Changes

Guian Fang, Yuchao Gu, Mike Zheng Shou

Generating controllable character animation from a reference image and motion guidance remains a challenging task due to the inherent difficulty of injecting appearance and motion…

cs.CV2025

Tuning-Free Image Editing with Fidelity and Editability via Unified Latent Diffusion Model

Qi Mao, Lan Chen, Yuchao Gu +2

Balancing fidelity and editability is essential in text-based image editing (TIE), where failures commonly lead to over- or under-editing issues. Existing methods typically rely on…

cs.CV2025

Long-Context Autoregressive Video Modeling with Next-Frame Prediction

Yuchao Gu, Weijia Mao, Mike Zheng Shou

Long-context video modeling is essential for enabling generative models to function as world simulators, as they must maintain temporal coherence over extended time spans. However,…

cs.CV2025

Edit Transfer: Learning Image Editing via Vision In-Context Relations

Lan Chen, Qi Mao, Yuchao Gu +1

We introduce a new setting, Edit Transfer, where a model learns a transformation from just a single source-target example and applies it to a new query image. While text-based meth…