1 paper
Shaoguang Wang, Weiyu Guo, Ziyang Chen +3
The practical application of Multimodal Large Language Models (MLLMs) to Video Question Answering (Video-QA) is severely hindered by the high token cost of processing numerous vide…