4 papers · 1 filter
SeMv-3D: Towards Concurrency of Semantic and Multi-view Consistency in General Text-to-3D Generation
Xiao Cai, Pengpeng Zeng, Lianli Gao +3
General Text-to-3D (GT23D) generation is crucial for creating diverse 3D content across objects and scenes, yet it faces two key challenges: 1) ensuring semantic consistency betwee…
CoIN: A Benchmark of Continual Instruction tuNing for Multimodel Large Language Model
Cheng Chen, Junchen Zhu, Xu Luo +3
Instruction tuning represents a prevalent strategy employed by Multimodal Large Language Models (MLLMs) to align with human instructions and adapt to new tasks. Nevertheless, MLLMs…
AICL: Action In-Context Learning for Video Diffusion Model
Jianzhi Liu, Junchen Zhu, Lianli Gao +2
The open-domain video generation models are constrained by the scale of the training video datasets, and some less common actions still cannot be generated. Some researchers explor…
Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation
Wenjing Wang, Huan Yang, Zixi Tuo +4
With the explosive popularity of AI-generated content (AIGC), video generation has recently received a lot of attention. Generating videos guided by text instructions poses signifi…