3 papers
cs.CV2026
HumanMoveVQA: Can Video MLLMs reason about human movement in videos?
Pulkit Gera, Faegheh Sardari, Asmar Nadeem +4
Despite the rapid advance of Multimodal Large Language Models (MLLMs) in high-level video understanding, a fundamental bottleneck remains: these models collapse complex human motio…
cs.CV2025
Reframing Dense Action Detection (RefDense): A Paradigm Shift in Problem Solving & a Novel Optimization Strategy
Faegheh Sardari, Armin Mustafa, Philip J. B. Jackson +1
Dense action detection involves detecting multiple co-occurring actions while action classes are often ambiguous and represent overlapping concepts. We argue that handling the dual…
cs.CV2025
NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative
Asmar Nadeem, Faegheh Sardari, Robert Dawes +3
Existing video captioning benchmarks and models lack causal-temporal narrative, which is sequences of events linked through cause and effect, unfolding over time and driven by char…