activity
20162023
most citedContextual Transformer Networks for Visual Recognition

42 citations · 128 across the 15 of their papers we have counts for

collaborators
Showing cs.CVShow all

11 papers · 1 filter

cs.CV20221 cited

Lightweight and Progressively-Scalable Networks for Semantic Segmentation

Yiheng Zhang, Ting Yao, Zhaofan Qiu +1

Multi-scale learning frameworks have been regarded as a capable class of models to boost semantic segmentation. The problem nevertheless is not trivial especially for the real-worl…

cs.CV20222 cited

Dual Vision Transformer

Ting Yao, Yehao Li, Yingwei Pan +3

Prior works have proposed several strategies to reduce the computational cost of self-attention mechanism. Many of these works consider decomposing the self-attention procedure int…

cs.CV20228 cited

Wave-ViT: Unifying Wavelet and Transformers for Visual Representation Learning

Ting Yao, Yingwei Pan, Yehao Li +2

Multi-scale Vision Transformer (ViT) has emerged as a powerful backbone for computer vision tasks, while the self-attention computation in Transformer scales quadratically w.r.t. t…

cs.CV20212 cited

A Style and Semantic Memory Mechanism for Domain Generalization

Yang Chen, Yu Wang, Yingwei Pan +3

Mainstream state-of-the-art domain generalization algorithms tend to prioritize the assumption on semantic invariance across domains. Meanwhile, the inherent intra-domain style inv…

cs.CV2021

Transferrable Contrastive Learning for Visual Domain Adaptation

Yang Chen, Yingwei Pan, Yu Wang +3

Self-supervised learning (SSL) has recently become the favorite among feature learning methodologies. It is therefore appealing for domain adaptation approaches to consider incorpo…

cs.CV20212 cited

CoCo-BERT: Improving Video-Language Pre-training with Contrastive Cross-modal Matching and Denoising

Jianjie Luo, Yehao Li, Yingwei Pan +3

BERT-type structure has led to the revolution of vision-language pre-training and the achievement of state-of-the-art results on numerous vision-language downstream tasks. Existing…