1 paper
Sihan Chen, Handong Li, Qunbo Wang +4
Vision and text have been fully explored in contemporary video-text foundational models, while other modalities such as audio and subtitles in videos have not received sufficient a…