24 citations · 35 across the 9 of their papers we have counts for
14 papers · 1 filter
What to look at and where: Semantic and Spatial Refined Transformer for detecting human-object interactions
A S M Iftekhar, Hao Chen, Kaustav Kundu +3
We propose a novel one-stage Transformer-based semantic and spatial refined transformer (SSRT) to solve the Human-Object Interaction detection task, which requires to localize huma…
SCVRL: Shuffled Contrastive Video Representation Learning
Michael Dorkenwald, Fanyi Xiao, Biagio Brattoli +2
We propose SCVRL, a novel contrastive-based framework for self-supervised learning for videos. Differently from previous contrast learning based methods that mostly focus on learni…
Hierarchical Self-supervised Representation Learning for Movie Understanding
Fanyi Xiao, Kaustav Kundu, Joseph Tighe +1
Most self-supervised video representation learning approaches focus on action recognition. In contrast, in this paper we focus on self-supervised video learning for movie understan…
Transfer of Representations to Video Label Propagation: Implementation Factors Matter
Daniel McKee, Zitong Zhan, Bing Shuai +3
This work studies feature representations for dense label propagation in video, with a focus on recently proposed methods that learn video correspondence using self-supervised sign…
Multi-Object Tracking with Hallucinated and Unlabeled Videos
Daniel McKee, Bing Shuai, Andrew Berneshawi +4
In this paper, we explore learning end-to-end deep neural trackers without tracking annotations. This is important as large-scale training data is essential for training deep neura…
SiamMOT: Siamese Multi-Object Tracking
Bing Shuai, Andrew Berneshawi, Xinyu Li +2
In this paper, we focus on improving online multi-object tracking (MOT). In particular, we introduce a region-based Siamese Multi-Object Tracking network, which we name SiamMOT. Si…