activity
20202025
most citedA Survey on Visual Transformer

4.4k citations · 4.5k across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV2025

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs

Yiman Zhang, Ziheng Luo, Qiangyu Yan +4

In this paper, we introduce OmniEval, a benchmark for evaluating omni-modality models like MiniCPM-O 2.6, which encompasses visual, auditory, and textual inputs. Compared with exis…

cs.CV2023★ 1 cited

TinySAM: Pushing the Envelope for Efficient Segment Anything Model

Han Shu, Wenshuo Li, Yehui Tang +5

Recently segment anything model (SAM) has shown powerful segmentation capability and has drawn great attention in computer vision fields. Massive following works have developed var…

cs.CV2020★ 4.4k cited

A Survey on Visual Transformer

Kai Han, Yunhe Wang, Hanting Chen +10

Transformer, first applied to the field of natural language processing, is a type of deep neural network mainly based on the self-attention mechanism. Thanks to its strong represen…

cs.CV2020★ 44 cited

Optical Flow Distillation: Towards Efficient and Stable Video Style Transfer

Xinghao Chen, Yiman Zhang, Yunhe Wang +3

Video style transfer techniques inspire many exciting applications on mobile devices. However, their efficiency and stability are still far from satisfactory. To boost the transfer…

cs.CV2020★ 15 cited

MTP: Multi-Task Pruning for Efficient Semantic Segmentation Networks

Xinghao Chen, Yiman Zhang, Yunhe Wang

This paper focuses on channel pruning for semantic segmentation networks. Previous methods to compress and accelerate deep neural networks in the classification task cannot be stra…