87 citations · 254 across the 10 of their papers we have counts for
11 papers · 1 filter
Frame-wise Action Representations for Long Videos via Sequence Contrastive Learning
Minghao Chen, Fangyun Wei, Chong Li +1
Prior works on action representation learning mainly focus on designing various architectures to extract the global representations for short video clips. In contrast, many practic…
Rethinking and Improving Relative Position Encoding for Vision Transformer
Kan Wu, Houwen Peng, Minghao Chen +2
Relative position encoding (RPE) is important for transformer to capture sequence ordering of input tokens. General efficacy has been proven in natural language processing. However…
AutoFormer: Searching Transformers for Visual Recognition
Minghao Chen, Houwen Peng, Jianlong Fu +1
Recently, pure transformer-based models have shown great potentials for vision tasks such as image classification and detection. However, the design of transformer networks is chal…
Salient Object Ranking with Position-Preserved Attention
Hao Fang, Daoxin Zhang, Yi Zhang +5
Instance segmentation can detect where the objects are in an image, but hard to understand the relationship between them. We pay attention to a typical relationship, relative salie…
One-Shot Neural Ensemble Architecture Search by Diversity-Guided Search Space Shrinking
Minghao Chen, Houwen Peng, Jianlong Fu +1
Despite remarkable progress achieved, most neural architecture search (NAS) methods focus on searching for one single accurate and robust architecture. To further build models with…
Suppress-and-Refine Framework for End-to-End 3D Object Detection
Zili Liu, Guodong Xu, Honghui Yang +5
3D object detector based on Hough voting achieves great success and derives many follow-up works. Despite constantly refreshing the detection accuracy, these works suffer from hand…