1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.LG2025★ 1 cited
I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models
Zhenxing Mi, Kuan-Chieh Wang, Guocheng Qian +5
This paper presents ThinkDiff, a novel alignment paradigm that empowers text-to-image diffusion models with multimodal in-context understanding and reasoning capabilities by integr…
cs.CV2024
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement
Zhefan Rao, Liya Ji, Yazhou Xing +6
Text-to-video (T2V) generation has gained significant attention recently. However, the costs of training a T2V model from scratch remain persistently high, and there is considerabl…