1 paper
Xiaoyang Guo, Guoping Luo, Jusheng Zhang +2
Adapting video vision-language models (VLMs) is computationally expensive because video inputs produce a large number of visual tokens, making both fine-tuning and inference costly…