1 paper
Mantek Singh, Jeshwanth Challagundla, Siddharth Raina +1
We present an efficient method to distill reasoning capabilities into compact video-language models (VLMs) for video question answering (VideoQA). Our approach fine-tunes a 2B-para…