2 citations · 7 across the 7 of their papers we have counts for
5 papers · 1 filter
Dynamic Token-Pass Transformers for Semantic Segmentation
Yuang Liu, Qiang Zhou, Jing Wang +3
Vision transformers (ViT) usually extract features via forwarding all the tokens in the self-attention layers from top to toe. In this paper, we introduce dynamic token-pass vision…
Entity-Level Text-Guided Image Manipulation
Yikai Wang, Jianan Wang, Guansong Lu +4
Existing text-guided image manipulation methods aim to modify the appearance of the image or to edit a few objects in a virtual or simple scenario, which is far from practical appl…
Weakly Supervised Video Salient Object Detection via Point Supervision
Shuyong Gao, Haozhe Xing, Wei Zhang +3
Video salient object detection models trained on pixel-wise dense annotation have achieved excellent performance, yet obtaining pixel-by-pixel annotated datasets is laborious. Seve…
Point2Seq: Detecting 3D Objects as Sequences
Yujing Xue, Jiageng Mao, Minzhe Niu +5
We present a simple and effective framework, named Point2Seq, for 3D object detection from point clouds. In contrast to previous methods that normally {predict attributes of 3D obj…
LCTR: On Awakening the Local Continuity of Transformer for Weakly Supervised Object Localization
Zhiwei Chen, Changan Wang, Yabiao Wang +6
Weakly supervised object localization (WSOL) aims to learn object localizer solely by using image-level labels. The convolution neural network (CNN) based techniques often result i…