5 citations · 7 across the 4 of their papers we have counts for
8 papers · 1 filter
SeMv-3D: Towards Concurrency of Semantic and Multi-view Consistency in General Text-to-3D Generation
Xiao Cai, Pengpeng Zeng, Lianli Gao +3
General Text-to-3D (GT23D) generation is crucial for creating diverse 3D content across objects and scenes, yet it faces two key challenges: 1) ensuring semantic consistency betwee…
AICL: Action In-Context Learning for Video Diffusion Model
Jianzhi Liu, Junchen Zhu, Lianli Gao +2
The open-domain video generation models are constrained by the scale of the training video datasets, and some less common actions still cannot be generated. Some researchers explor…
CoIN: A Benchmark of Continual Instruction tuNing for Multimodel Large Language Model
Cheng Chen, Junchen Zhu, Xu Luo +3
Instruction tuning represents a prevalent strategy employed by Multimodal Large Language Models (MLLMs) to align with human instructions and adapt to new tasks. Nevertheless, MLLMs…
Training-Free Semantic Video Composition via Pre-trained Diffusion Model
Jiaqi Guo, Sitong Su, Junchen Zhu +2
The video composition task aims to integrate specified foregrounds and backgrounds from different videos into a harmonious composite. Current approaches, predominantly trained on v…
CUCL: Codebook for Unsupervised Continual Learning
Chen Cheng, Jingkuan Song, Xiaosu Zhu +3
The focus of this study is on Unsupervised Continual Learning (UCL), as it presents an alternative to Supervised Continual Learning which needs high-quality manual labeled data. Th…
MobileVidFactory: Automatic Diffusion-Based Social Media Video Generation for Mobile Devices from Text
Junchen Zhu, Huan Yang, Wenjing Wang +8
Videos for mobile devices become the most popular access to share and acquire information recently. For the convenience of users' creation, in this paper, we present a system, name…