2 citations · 4 across the 9 of their papers we have counts for
1 paper · 1 filter
Rwiddhi Chakraborty, Yinong, Wang +7
Large video models have exhibited impressive performance on a wide range of visual question answering tasks, owing to the rise of powerful, pretrained text and vision encoders. The…