113 citations · 348 across the 32 of their papers we have counts for
63 papers
Tackling Visual Control via Multi-View Exploration Maximization
Mingqi Yuan, Xin Jin, Bo Li +1
We present MEM: Multi-view Exploration Maximization for tackling complex visual control tasks. To the best of our knowledge, MEM is the first approach that combines multi-view repr…
ReSTR: Convolution-free Referring Image Segmentation Using Transformers
Namyup Kim, Dongwon Kim, Cuiling Lan +2
Referring image segmentation is an advanced semantic segmentation task where target is not a predefined class but is described in natural language. Most of existing methods for thi…
Correlation-Aware Deep Tracking
Fei Xie, Chunyu Wang, Guangting Wang +3
Robustness and discrimination power are two fundamental requirements in visual object tracking. In most tracking paradigms, we find that the features extracted by the popular Siame…
Retriever: Learning Content-Style Representation as a Token-Level Bipartite Graph
Dacheng Yin, Xuanchi Ren, Chong Luo +3
This paper addresses the unsupervised learning of content-style decomposed representation. We first give a definition of style and then model the content-style representation as a…
When Shift Operation Meets Vision Transformer: An Extremely Simple Alternative to Attention Mechanism
Guangting Wang, Yucheng Zhao, Chuanxin Tang +2
Attention mechanism has been widely believed as the key to success of vision transformers (ViTs), since it provides a flexible and powerful way to model spatial relationships. Howe…
Multi-Scale Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action Recognition
Pengfei Zhang, Cuiling Lan, Wenjun Zeng +3
Skeleton data is of low dimension. However, there is a trend of using very deep and complicated feedforward neural networks to model the skeleton sequence without considering the c…