28 citations · 28 across the 1 of their papers we have counts for
1 paper · 1 filter
Sihan Chen, Handong Li, Qunbo Wang +4
Vision and text have been fully explored in contemporary video-text foundational models, while other modalities such as audio and subtitles in videos have not received sufficient a…