1 paper · 2 filters
Zhengjian Kang, Qi Chen, Rui Liu +4
Recent Video Large Language Models (Video-LLMs) have shown strong multimodal reasoning capabilities, yet remain challenged by video understanding tasks that require consistent temp…