4.4k citations · 4.5k across the 5 of their papers we have counts for
5 papers
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
Yiman Zhang, Ziheng Luo, Qiangyu Yan +4
In this paper, we introduce OmniEval, a benchmark for evaluating omni-modality models like MiniCPM-O 2.6, which encompasses visual, auditory, and textual inputs. Compared with exis…
TinySAM: Pushing the Envelope for Efficient Segment Anything Model
Han Shu, Wenshuo Li, Yehui Tang +5
Recently segment anything model (SAM) has shown powerful segmentation capability and has drawn great attention in computer vision fields. Massive following works have developed var…
A Survey on Visual Transformer
Kai Han, Yunhe Wang, Hanting Chen +10
Transformer, first applied to the field of natural language processing, is a type of deep neural network mainly based on the self-attention mechanism. Thanks to its strong represen…
Optical Flow Distillation: Towards Efficient and Stable Video Style Transfer
Xinghao Chen, Yiman Zhang, Yunhe Wang +3
Video style transfer techniques inspire many exciting applications on mobile devices. However, their efficiency and stability are still far from satisfactory. To boost the transfer…
MTP: Multi-Task Pruning for Efficient Semantic Segmentation Networks
Xinghao Chen, Yiman Zhang, Yunhe Wang
This paper focuses on channel pruning for semantic segmentation networks. Previous methods to compress and accelerate deep neural networks in the classification task cannot be stra…