1 citations · 1 across the 4 of their papers we have counts for
4 papers
S3R-Net: A Single-Stage Approach to Self-Supervised Shadow Removal
Nikolina Kubiak, Armin Mustafa, Graeme Phillipson +2
In this paper we present S3R-Net, the Self-Supervised Shadow Removal Network. The two-branch WGAN model achieves self-supervision relying on the unify-and-adaptphenomenon - it unif…
CAD -- Contextual Multi-modal Alignment for Dynamic AVQA
Asmar Nadeem, Adrian Hilton, Robert Dawes +2
In the context of Audio Visual Question Answering (AVQA) tasks, the audio visual modalities could be learnt on three levels: 1) Spatial, 2) Temporal, and 3) Semantic. Existing AVQA…
PAT: Position-Aware Transformer for Dense Multi-Label Action Detection
Faegheh Sardari, Armin Mustafa, Philip J. B. Jackson +1
We present PAT, a transformer-based network that learns complex temporal co-occurrence action dependencies in a video by exploiting multi-scale temporal features. In existing metho…
SEM-POS: Grammatically and Semantically Correct Video Captioning
Asmar Nadeem, Adrian Hilton, Robert Dawes +2
Generating grammatically and semantically correct captions in video captioning is a challenging task. The captions generated from the existing methods are either word-by-word that…