1 citations · 1 across the 21 of their papers we have counts for
1 paper · 2 filters
Gengyuan Zhang, Jinhe Bi, Jindong Gu +2
Understanding videos is an important research topic for multimodal learning. Leveraging large-scale datasets of web-crawled video-text pairs as weak supervision has become a pre-tr…