10 papers
Addressable Memory for Video World Models
Xindi Wu, Sven Elflein, James Lucas +5
We study visual persistence in interactive video world models. These models rely on a Key-Value (KV) cache as a growing visual memory to carry forward previously generated frames.…
Explain Before You Answer: A Survey on Compositional Visual Reasoning
Fucai Ke, Joy Hsu, Zhixi Cai +10
Compositional visual reasoning has emerged as a key research frontier in multimodal AI, aiming to endow machines with the human-like ability to decompose visual scenes, ground inte…
Motion Attribution for Video Generation
Xindi Wu, Despoina Paschalidou, Jun Gao +5
Despite the rapid progress of video generation models, the role of data in influencing motion is poorly understood. We present Motive (MOTIon attribution for Video gEneration), a m…
Beyond Objects: Contextual Synthetic Data Generation for Fine-Grained Classification
William Yang, Xindi Wu, Zhiwei Deng +2
Text-to-image (T2I) models are increasingly used for synthetic dataset generation, but generating effective synthetic training data for classification remains challenging. Fine-tun…
Variance Reduction for Expectations with Diffusion Teachers
Jesse Bettencourt, Xindi Wu, Matan Atzmon +2
Pretrained diffusion models serve as frozen teachers feeding downstream pipelines such as text-to-3D, single-step distillation, and data attribution. The teacher gradients these pi…
Visual Compositional Tuning
Xindi Wu, Hee Seung Hwang, Polina Kirichenko +2
Visual instruction tuning (VIT) datasets have grown rapidly in scale, yet the informativeness of individual training samples has largely been overlooked. Recent dataset selection m…