5 papers
Learning to Deny: Action Denial in Multimodal Large Language Models
Raiyaan Abdullah, Shehreen Azad, Yogesh Singh Rawat
Multimodal large language models (MLLMs) have rapidly advanced video understanding, achieving strong zero-shot and few-shot recognition across standard benchmarks. Yet their abilit…
StreamReady: Learning What to Answer and When in Long Streaming Videos
Shehreen Azad, Vibhav Vineet, Yogesh Singh Rawat
Streaming video understanding often involves time-sensitive scenarios where models need to answer exactly when the supporting visual evidence appears: answering before the evidence…
DisenQ: Disentangling Q-Former for Activity-Biometrics
Shehreen Azad, Yogesh S Rawat
In this work, we address activity-biometrics, which involves identifying individuals across diverse set of activities. Unlike traditional person identification, this setting introd…
Understanding Depth and Height Perception in Large Visual-Language Models
Shehreen Azad, Yash Jain, Rishit Garg +2
Geometric understanding - including depth and height perception - is fundamental to intelligence and crucial for navigating our environment. Despite the impressive capabilities of…
HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
Shehreen Azad, Vibhav Vineet, Yogesh Singh Rawat
Despite advancements in multimodal large language models (MLLMs), current approaches struggle in medium-to-long video understanding due to frame and context length limitations. As…