7 citations · 7 across the 5 of their papers we have counts for
1 paper · 1 filter
Yicheng Ji, Jun Zhang, Heming Xia +4
Video large language models (Vid-LLMs) have shown strong capabilities in understanding video content. However, their reliance on dense video token representations introduces substa…