500 citations · 574 across the 10 of their papers we have counts for
11 papers · 1 filter
RefineVIS: Video Instance Segmentation with Temporal Attention Refinement
Andre Abrantes, Jiang Wang, Peng Chu +2
We introduce a novel framework called RefineVIS for Video Instance Segmentation (VIS) that achieves good object association between frames and accurate segmentation masks by iterat…
SA-VQA: Structured Alignment of Visual and Semantic Representations for Visual Question Answering
Peixi Xiong, Quanzeng You, Pei Yu +2
Visual Question Answering (VQA) attracts much attention from both industry and academia. As a multi-modality task, it is challenging since it requires not only visual and textual u…
TransMOT: Spatial-Temporal Graph Transformer for Multiple Object Tracking
Peng Chu, Jiang Wang, Quanzeng You +2
Tracking multiple objects in videos relies on modeling the spatial-temporal interactions of the objects. In this paper, we propose a solution named TransMOT, which leverages powerf…
Disentanglement-based Cross-Domain Feature Augmentation for Effective Unsupervised Domain Adaptive Person Re-identification
Zhizheng Zhang, Cuiling Lan, Wenjun Zeng +4
Unsupervised domain adaptive (UDA) person re-identification (ReID) aims to transfer the knowledge from the labeled source domain to the unlabeled target domain for person matching.…
Real-time 3D Deep Multi-Camera Tracking
Quanzeng You, Hao Jiang
Tracking a crowd in 3D using multiple RGB cameras is a challenging task. Most previous multi-camera tracking algorithms are designed for offline setting and have high computational…
Real-time Multiple People Hand Localization in 4D Point Clouds
Hao Jiang, Quanzeng You
We propose novel real-time algorithm to localize hands and find their associations with multiple people in the cluttered 4D volumetric data (dynamic 3D volumes). Different from the…