collaborators

5 papers

cs.CV2026

Learning to Deny: Action Denial in Multimodal Large Language Models

Raiyaan Abdullah, Shehreen Azad, Yogesh Singh Rawat

Multimodal large language models (MLLMs) have rapidly advanced video understanding, achieving strong zero-shot and few-shot recognition across standard benchmarks. Yet their abilit…

cs.CV2026

StreamReady: Learning What to Answer and When in Long Streaming Videos

Shehreen Azad, Vibhav Vineet, Yogesh Singh Rawat

Streaming video understanding often involves time-sensitive scenarios where models need to answer exactly when the supporting visual evidence appears: answering before the evidence…

cs.CV2025

DisenQ: Disentangling Q-Former for Activity-Biometrics

Shehreen Azad, Yogesh S Rawat

In this work, we address activity-biometrics, which involves identifying individuals across diverse set of activities. Unlike traditional person identification, this setting introd…

cs.CV2025

Understanding Depth and Height Perception in Large Visual-Language Models

Shehreen Azad, Yash Jain, Rishit Garg +2

Geometric understanding - including depth and height perception - is fundamental to intelligence and crucial for navigating our environment. Despite the impressive capabilities of…

cs.CV2025

HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding

Shehreen Azad, Vibhav Vineet, Yogesh Singh Rawat

Despite advancements in multimodal large language models (MLLMs), current approaches struggle in medium-to-long video understanding due to frame and context length limitations. As…