activity
20132022
most citedCrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention

85 citations · 361 across the 21 of their papers we have counts for

collaborators
Showing cs.CVShow all

12 papers · 1 filter

cs.CV20212 cited

SimulLR: Simultaneous Lip Reading Transducer with Attention-Guided Adaptive Memory

Zhijie Lin, Zhou Zhao, Haoyuan Li +4

Lip reading, aiming to recognize spoken sentences according to the given video of lip movements without relying on the audio stream, has attracted great interest due to its applica…

cs.CV202185 cited

CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention

Wenxiao Wang, Lu Yao, Long Chen +4

Transformers have made great progress in dealing with computer vision tasks. However, existing vision transformers do not yet possess the ability of building the interactions among…

cs.CV20213 cited

Salient Object Ranking with Position-Preserved Attention

Hao Fang, Daoxin Zhang, Yi Zhang +5

Instance segmentation can detect where the objects are in an image, but hard to understand the relationship between them. We pay attention to a typical relationship, relative salie…

cs.CV2021

Discriminative-Generative Dual Memory Video Anomaly Detection

Xin Guo, Zhongming Jin, Chong Chen +5

Recently, people tried to use a few anomalies for video anomaly detection (VAD) instead of only normal data during the training process. A side effect of data imbalance occurs when…

cs.CV202111 cited

OCM3D: Object-Centric Monocular 3D Object Detection

Liang Peng, Fei Liu, Senbo Yan +2

Image-only and pseudo-LiDAR representations are commonly used for monocular 3D object detection. However, methods based on them have shortcomings of either not well capturing the s…

cs.CV2021

X-view: Non-egocentric Multi-View 3D Object Detector

Liang Xie, Guodong Xu, Deng Cai +1

3D object detection algorithms for autonomous driving reason about 3D obstacles either from 3D birds-eye view or perspective view or both. Recent works attempt to improve the detec…