1 paper · 1 filter
Shruti Singh Baghel, Yash Pratap Singh Rathore, Sushovan Jena +4
Large Vision-Language Models (VLMs) excel at understanding and generating video descriptions but their high memory, computation, and deployment demands hinder practical use particu…