collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

HumanMoveVQA: Can Video MLLMs reason about human movement in videos?

Pulkit Gera, Faegheh Sardari, Asmar Nadeem +4

Despite the rapid advance of Multimodal Large Language Models (MLLMs) in high-level video understanding, a fundamental bottleneck remains: these models collapse complex human motio…

cs.CV2024

Efficient Audio-Visual Fusion for Video Classification

Mahrukh Awan, Asmar Nadeem, Armin Mustafa

We present Attend-Fusion, a novel and efficient approach for audio-visual fusion in video classification tasks. Our method addresses the challenge of exploiting both audio and visu…

cs.CV2024

Attend-Fusion: Efficient Audio-Visual Fusion for Video Classification

Mahrukh Awan, Asmar Nadeem, Muhammad Junaid Awan +2

Exploiting both audio and visual modalities for video classification is a challenging task, as the existing methods require large model architectures, leading to high computational…

cs.CV2024

NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative

Asmar Nadeem, Faegheh Sardari, Robert Dawes +3

Existing video captioning benchmarks and models lack causal-temporal narrative, which is sequences of events linked through cause and effect, unfolding over time and driven by char…

cs.CV20231 cited

CAD -- Contextual Multi-modal Alignment for Dynamic AVQA

Asmar Nadeem, Adrian Hilton, Robert Dawes +2

In the context of Audio Visual Question Answering (AVQA) tasks, the audio visual modalities could be learnt on three levels: 1) Spatial, 2) Temporal, and 3) Semantic. Existing AVQA…

cs.CV2023

SEM-POS: Grammatically and Semantically Correct Video Captioning

Asmar Nadeem, Adrian Hilton, Robert Dawes +2

Generating grammatically and semantically correct captions in video captioning is a challenging task. The captions generated from the existing methods are either word-by-word that…