1.8k citations · 1.9k across the 16 of their papers we have counts for
16 papers
SLAck: Semantic, Location, and Appearance Aware Open-Vocabulary Tracking
Siyuan Li, Lei Ke, Yung-Hsu Yang +4
Open-vocabulary Multiple Object Tracking (MOT) aims to generalize trackers to novel categories not in the training set. Currently, the best-performing methods are mainly based on p…
Analyzing Local Representations of Self-supervised Vision Transformers
Ani Vanyan, Alvard Barseghyan, Hakob Tamazyan +3
In this paper, we present a comparative analysis of various self-supervised Vision Transformers (ViTs), focusing on their local representative power. Inspired by large language mod…
R3D3: Dense 3D Reconstruction of Dynamic Scenes from Multiple Cameras
Aron Schmied, Tobias Fischer, Martin Danelljan +2
Dense 3D reconstruction and ego-motion estimation are key challenges in autonomous driving and robotics. Compared to the complex, multi-modal systems deployed today, multi-camera s…
Cascade-DETR: Delving into High-Quality Universal Object Detection
Mingqiao Ye, Lei Ke, Siyuan Li +4
Object localization in general environments is a fundamental part of vision systems. While dominating on the COCO benchmark, recent Transformer-based detection methods are not comp…
StyleGenes: Discrete and Efficient Latent Distributions for GANs
Evangelos Ntavelis, Mohamad Shahbazi, Iason Kastanis +3
We propose a discrete latent distribution for Generative Adversarial Networks (GANs). Instead of drawing latent vectors from a continuous prior, we sample from a finite set of lear…
Mask-Free Video Instance Segmentation
Lei Ke, Martin Danelljan, Henghui Ding +3
The recent advancement in Video Instance Segmentation (VIS) has largely been driven by the use of deeper and increasingly data-hungry transformer-based models. However, video masks…