collaborators

15 papers

cs.CV2026

Addressable Memory for Video World Models

Xindi Wu, Sven Elflein, James Lucas +5

We study visual persistence in interactive video world models. These models rely on a Key-Value (KV) cache as a growing visual memory to carry forward previously generated frames.…

cs.CV2026

Motion Attribution for Video Generation

Xindi Wu, Despoina Paschalidou, Jun Gao +5

Despite the rapid progress of video generation models, the role of data in influencing motion is poorly understood. We present Motive (MOTIon attribution for Video gEneration), a m…

cs.CV2026

Beyond Objects: Contextual Synthetic Data Generation for Fine-Grained Classification

William Yang, Xindi Wu, Zhiwei Deng +2

Text-to-image (T2I) models are increasingly used for synthetic dataset generation, but generating effective synthetic training data for classification remains challenging. Fine-tun…

cs.CV2026

GASS: Geometry-Aware Spherical Sampling for Disentangled Diversity Enhancement in Text-to-Image Generation

Ye Zhu, Kaleb S. Newman, Johannes F. Lutzeyer +3

Despite high semantic alignment, modern text-to-image (T2I) generative models still struggle to synthesize diverse images from a given prompt. In this work, we enhance the T2I dive…

cs.CV2026

Personalized Generative Models for Contextual Debiasing

Xinran Liang, Esin Tureci, Prachi Sinha +3

Different visual patterns appear with different frequencies in the world: e.g., beach balls appear on sand more often than they do on a road. These statistics are reflected in visi…

cs.CV2026

Visual Compositional Tuning

Xindi Wu, Hee Seung Hwang, Polina Kirichenko +2

Visual instruction tuning (VIT) datasets have grown rapidly in scale, yet the informativeness of individual training samples has largely been overlooked. Recent dataset selection m…