activity
20202022
most citedUniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

59 citations · 275 across the 30 of their papers we have counts for

collaborators

34 papers

cs.CL202214 cited

Toward Efficient Language Model Pretraining and Downstream Adaptation via Self-Evolution: A Case Study on SuperGLUE

Qihuang Zhong, Liang Ding, Yibing Zhan +11

This technical report briefly describes our JDExplore d-team's Vega v2 submission on the SuperGLUE leaderboard. SuperGLUE is more challenging than the widely used general language…

cs.CV20221 cited

Improving Training and Inference of Face Recognition Models via Random Temperature Scaling

Lei Shang, Mouxiao Huang, Wu Shi +6

Data uncertainty is commonly observed in the images for face recognition (FR). However, deep learning algorithms often make predictions with high confidence even for uncertain or i…

cs.CV20223 cited

Towards All-in-one Pre-training via Maximizing Multi-modal Mutual Information

Weijie Su, Xizhou Zhu, Chenxin Tao +7

To effectively exploit the potential of large-scale models, various pre-training strategies supported by massive data from different sources are proposed, including supervised pre-…

cs.CV20226 cited

BEVFormer v2: Adapting Modern Image Backbones to Bird's-Eye-View Recognition via Perspective Supervision

Chenyu Yang, Yuntao Chen, Hao Tian +9

We present a novel bird's-eye-view (BEV) detector with perspective supervision, which converges faster and better suits modern image backbones. Existing state-of-the-art BEV detect…

cs.CV20223 cited

Uni-Perceiver v2: A Generalist Model for Large-Scale Vision and Vision-Language Tasks

Hao Li, Jinguo Zhu, Xiaohu Jiang +8

Despite the remarkable success of foundation models, their task-specific fine-tuning paradigm makes them inconsistent with the goal of general perception modeling. The key to elimi…

cs.CV202259 cited

UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Kunchang Li, Yali Wang, Yinan He +4

Learning discriminative spatiotemporal representation is the key problem of video understanding. Recently, Vision Transformers (ViTs) have shown their power in learning long-term v…