1 paper · 1 filter
Haifeng Huang, Yang Li
Video Large Language Models (Video LLMs) achieve strong performance on video understanding tasks but suffer from high inference costs due to the large number of visual tokens. We p…