4.4k citations · 4.8k across the 33 of their papers we have counts for
65 papers · 1 filter
3DMODT: Attention-Guided Affinities for Joint Detection & Tracking in 3D Point Clouds
Jyoti Kini, Ajmal Mian, Mubarak Shah
We propose a method for joint detection and tracking of multiple objects in 3D point clouds, a task conventionally treated as a two-step process comprising object detection followe…
Self-Supervised Video Object Segmentation via Cutout Prediction and Tagging
Jyoti Kini, Fahad Shahbaz Khan, Salman Khan +1
We propose a novel self-supervised Video Object Segmentation (VOS) approach that strives to achieve better object-background discriminability for accurate object segmentation. Dist…
Tag-Based Attention Guided Bottom-Up Approach for Video Instance Segmentation
Jyoti Kini, Mubarak Shah
Video Instance Segmentation is a fundamental computer vision task that deals with segmenting and tracking object instances across a video sequence. Most existing methods typically…
Video Action Detection: Analysing Limitations and Challenges
Rajat Modi, Aayush Jung Rana, Akash Kumar +4
Beyond possessing large enough size to feed data hungry machines (eg, transformers), what attributes measure the quality of a dataset? Assuming that the definitions of such attribu…
PSTR: End-to-End One-Step Person Search With Transformers
Jiale Cao, Yanwei Pang, Rao Muhammad Anwer +4
We propose a novel one-step transformer-based person search framework, PSTR, that jointly performs person detection and re-identification (re-id) in a single architecture. PSTR com…
TransGeo: Transformer Is All You Need for Cross-view Image Geo-localization
Sijie Zhu, Mubarak Shah, Chen Chen
The dominant CNN-based methods for cross-view image geo-localization rely on polar transform and fail to model global correlation. We propose a pure transformer-based approach (Tra…