3 papers
cs.CV2026
Learning to Deny: Action Denial in Multimodal Large Language Models
Raiyaan Abdullah, Shehreen Azad, Yogesh Singh Rawat
Multimodal large language models (MLLMs) have rapidly advanced video understanding, achieving strong zero-shot and few-shot recognition across standard benchmarks. Yet their abilit…
cs.CV2025
iSafetyBench: A video-language benchmark for safety in industrial environment
Raiyaan Abdullah, Yogesh Singh Rawat, Shruti Vyas
Recent advances in vision-language models (VLMs) have enabled impressive generalization across diverse video understanding tasks under zero-shot settings. However, their capabiliti…
cs.CV2025
Punching Bag vs. Punching Person: Motion Transferability in Videos
Raiyaan Abdullah, Jared Claypoole, Michael Cogswell +2
Action recognition models demonstrate strong generalization, but can they effectively transfer high-level motion concepts across diverse contexts, even within similar distributions…