4 papers · 1 filter
HumanMoveVQA: Can Video MLLMs reason about human movement in videos?
Pulkit Gera, Faegheh Sardari, Asmar Nadeem +4
Despite the rapid advance of Multimodal Large Language Models (MLLMs) in high-level video understanding, a fundamental bottleneck remains: these models collapse complex human motio…
NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative
Asmar Nadeem, Faegheh Sardari, Robert Dawes +3
Existing video captioning benchmarks and models lack causal-temporal narrative, which is sequences of events linked through cause and effect, unfolding over time and driven by char…
Efficient Audio-Visual Fusion for Video Classification
Mahrukh Awan, Asmar Nadeem, Armin Mustafa
We present Attend-Fusion, a novel and efficient approach for audio-visual fusion in video classification tasks. Our method addresses the challenge of exploiting both audio and visu…
Attend-Fusion: Efficient Audio-Visual Fusion for Video Classification
Mahrukh Awan, Asmar Nadeem, Muhammad Junaid Awan +2
Exploiting both audio and visual modalities for video classification is a challenging task, as the existing methods require large model architectures, leading to high computational…