1 paper · 1 filter
Yicheng Ji, Jun Zhang, Heming Xia +4
Video large language models (Vid-LLMs) have shown strong capabilities in understanding video content. However, their reliance on dense video token representations introduces substa…