activity
20182023
most citedUniFormer: Unified Transformer for Efficient Spatiotemporal Representation Learning

108 citations · 297 across the 19 of their papers we have counts for

collaborators

24 papers

cs.CV202320 cited

LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models

Peng Xu, Wenqi Shao, Kaipeng Zhang +7

Large Vision-Language Models (LVLMs) have recently played a dominant role in multimodal vision-language learning. Despite the great success, it lacks a holistic evaluation of their…

cs.CV202252 cited

ConvMAE: Masked Convolution Meets Masked Autoencoders

Peng Gao, Teli Ma, Hongsheng Li +3

Vision Transformers (ViT) become widely-adopted architectures for various vision tasks. Masked auto-encoding for feature pretraining and multi-scale hybrid convolution-transformer…

cs.CV20223 cited

POS-BERT: Point Cloud One-Stage BERT Pre-Training

Kexue Fu, Peng Gao, ShaoLei Liu +3

Recently, the pre-training paradigm combining Transformer and masked language modeling has achieved tremendous success in NLP, images, and point clouds, such as BERT. However, dire…

cs.CV202213 cited

Distillation with Contrast is All You Need for Self-Supervised Point Cloud Representation Learning

Kexue Fu, Peng Gao, Renrui Zhang +3

In this paper, we propose a simple and general framework for self-supervised point cloud representation learning. Human beings understand the 3D world by extracting two levels of i…

cs.CV2022108 cited

UniFormer: Unified Transformer for Efficient Spatiotemporal Representation Learning

Kunchang Li, Yali Wang, Peng Gao +4

It is a challenging task to learn rich and multi-scale spatiotemporal semantics from high-dimensional videos, due to large local redundancy and complex global dependency between vi…

cs.CV20225 cited

TerViT: An Efficient Ternary Vision Transformer

Sheng Xu, Yanjing Li, Teli Ma +4

Vision transformers (ViTs) have demonstrated great potential in various visual tasks, but suffer from expensive computational and memory cost problems when deployed on resource-con…