1 citations · 2 across the 4 of their papers we have counts for
4 papers
CAD -- Contextual Multi-modal Alignment for Dynamic AVQA
Asmar Nadeem, Adrian Hilton, Robert Dawes +2
In the context of Audio Visual Question Answering (AVQA) tasks, the audio visual modalities could be learnt on three levels: 1) Spatial, 2) Temporal, and 3) Semantic. Existing AVQA…
PAT: Position-Aware Transformer for Dense Multi-Label Action Detection
Faegheh Sardari, Armin Mustafa, Philip J. B. Jackson +1
We present PAT, a transformer-based network that learns complex temporal co-occurrence action dependencies in a video by exploiting multi-scale temporal features. In existing metho…
SEM-POS: Grammatically and Semantically Correct Video Captioning
Asmar Nadeem, Adrian Hilton, Robert Dawes +2
Generating grammatically and semantically correct captions in video captioning is a challenging task. The captions generated from the existing methods are either word-by-word that…
Super-resolution 3D Human Shape from a Single Low-Resolution Image
Marco Pesavento, Marco Volino, Adrian Hilton
We propose a novel framework to reconstruct super-resolution human shape from a single low-resolution input image. The approach overcomes limitations of existing approaches that re…