1 paper
Jiayang He, Tianling Xu, Diancheng Kang +4
Video large language models (Video-LLMs) represent videos as dense sequences of visual tokens, whose length grows with the temporal and spatial extent of the input. These tokens of…