167 citations · 282 across the 10 of their papers we have counts for
12 papers · 1 filter
Benchmarking Large and Small MLLMs
Xuelu Feng, Yunsheng Li, Dongdong Chen +4
Large multimodal language models (MLLMs) such as GPT-4V and GPT-4o have achieved remarkable advancements in understanding and generating multimodal content, showcasing superior qua…
Self-Supervised Learning based on Heat Equation
Yinpeng Chen, Xiyang Dai, Dongdong Chen +4
This paper presents a new perspective of self-supervised learning based on extending heat equation into high dimensional feature space. In particular, we remove time dependence by…
Reduce Information Loss in Transformers for Pluralistic Image Inpainting
Qiankun Liu, Zhentao Tan, Dongdong Chen +6
Transformers have achieved great success in pluralistic image inpainting recently. However, we find existing transformer based solutions regard each pixel as a token, thus suffer f…
MiniViT: Compressing Vision Transformers with Weight Multiplexing
Jinnian Zhang, Houwen Peng, Kan Wu +4
Vision Transformer (ViT) models have recently drawn much attention in computer vision due to their high model capability. However, ViT models suffer from huge number of parameters,…
MicroNet: Improving Image Recognition with Extremely Low FLOPs
Yunsheng Li, Yinpeng Chen, Xiyang Dai +6
This paper aims at addressing the problem of substantial performance degradation at extremely low computational cost (e.g. 5M FLOPs on ImageNet classification). We found that two f…
Dynamic Head: Unifying Object Detection Heads with Attentions
Xiyang Dai, Yinpeng Chen, Bin Xiao +4
The complex nature of combining localization and classification in object detection has resulted in the flourished development of methods. Previous works tried to improve the perfo…