35 citations · 35 across the 4 of their papers we have counts for
4 papers · 1 filter
SemanticMIM: Marring Masked Image Modeling with Semantics Compression for General Visual Representation
Yike Yuan, Huanzhang Dou, Fengjun Guo +1
This paper represents a neat yet effective framework, named SemanticMIM, to integrate the advantages of masked image modeling (MIM) and contrastive learning (CL) for general visual…
MMBench: Is Your Multi-modal Model an All-around Player?
Yuan Liu, Haodong Duan, Yuanhan Zhang +9
Large vision-language models (VLMs) have recently achieved remarkable progress, exhibiting impressive multimodal perception and reasoning abilities. However, effectively evaluating…
DenseDINO: Boosting Dense Self-Supervised Learning with Token-Based Point-Level Consistency
Yike Yuan, Xinghe Fu, Yunlong Yu +1
In this paper, we propose a simple yet effective transformer framework for self-supervised learning called DenseDINO to learn dense visual representations. To exploit the spatial i…
Disassembling Object Representations without Labels
Zunlei Feng, Xinchao Wang, Yongming He +3
In this paper, we study a new representation-learning task, which we termed as disassembling object representations. Given an image featuring multiple objects, the goal of disassem…