2 citations · 2 across the 1 of their papers we have counts for
1 paper
Sihan Chen, Xingjian He, Handong Li +3
Due to the limited scale and quality of video-text training corpus, most vision-language foundation models employ image-text datasets for pretraining and primarily focus on modelin…