21 citations · 21 across the 2 of their papers we have counts for
1 paper · 1 filter
Yifei Chen, Dapeng Chen, Ruijin Liu +3
Large-scale visual-language pre-trained models have achieved significant success in various video tasks. However, most existing methods follow an "adapt then align" paradigm, which…