1 paper
Wufei Ma, Kai Li, Zhongshi Jiang +5
Recent video-text foundation models have demonstrated strong performance on a wide variety of downstream video understanding tasks. Can these video-text models genuinely understand…