12 citations · 23 across the 2 of their papers we have counts for
2 papers
cs.CV2023★ 11 cited
VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending
Xingjian He, Sihan Chen, Fan Ma +7
Large-scale image-text contrastive pre-training models, such as CLIP, have been demonstrated to effectively learn high-quality multimodal representations. However, there is limited…
cs.CV2023★ 12 cited
Sounding Video Generator: A Unified Framework for Text-guided Sounding Video Generation
Jiawei Liu, Weining Wang, Sihan Chen +2
As a combination of visual and audio signals, video is inherently multi-modal. However, existing video generation methods are primarily intended for the synthesis of visual frames,…