1 paper
Shivprasad Sagare, Hemachandran S, Kinshuk Sarabhai +2
Recent advances in multimodal LLMs, have led to several video-text models being proposed for critical video-related tasks. However, most of the previous works support visual input…