1 paper · 1 filter
Jiayang He, Tianling Xu, Diancheng Kang +4
Video large language models (Video-LLMs) represent videos as dense sequences of visual tokens, whose length grows with the temporal and spatial extent of the input. These tokens of…