1 citations · 2 across the 10 of their papers we have counts for
Showing 2023Show all
2 papers · 1 filter
cs.CV2023
PAT: Position-Aware Transformer for Dense Multi-Label Action Detection
Faegheh Sardari, Armin Mustafa, Philip J. B. Jackson +1
We present PAT, a transformer-based network that learns complex temporal co-occurrence action dependencies in a video by exploiting multi-scale temporal features. In existing metho…
cs.CV2023
SEM-POS: Grammatically and Semantically Correct Video Captioning
Asmar Nadeem, Adrian Hilton, Robert Dawes +2
Generating grammatically and semantically correct captions in video captioning is a challenging task. The captions generated from the existing methods are either word-by-word that…