5 papers · 1 filter
Studying Image Tokenizers as Visual Languages in Unified Multimodal Models
Siting Li, Zhengyang Wang, Simon Shaolei Du +2
Image tokenizers define the ``visual language'' of unified multimodal models, yet are commonly studied through isolated metrics or generation-/understanding-only evaluations. These…
Unveiling the Backbone-Optimizer Coupling Bias in Visual Representation Learning
Siyuan Li, Juanxi Tian, Zedong Wang +6
This paper delves into the interplay between vision backbones and optimizers, unvealing an inter-dependent phenomenon termed \textit{\textbf{b}ackbone-\textbf{o}ptimizer \textbf{c}…
TopoFR: A Closer Look at Topology Alignment on Face Recognition
Jun Dan, Yang Liu, Jiankang Deng +4
The field of face recognition (FR) has undergone significant advancements with the rise of deep learning. Recently, the success of unsupervised learning and graph neural networks h…
PRCL: Probabilistic Representation Contrastive Learning for Semi-Supervised Semantic Segmentation
Haoyu Xie, Changqi Wang, Jian Zhao +4
Tremendous breakthroughs have been developed in Semi-Supervised Semantic Segmentation (S4) through contrastive learning. However, due to limited annotations, the guidance on unlabe…
Masked Modeling for Self-supervised Representation Learning on Vision and Beyond
Siyuan Li, Luyuan Zhang, Zedong Wang +8
As the deep learning revolution marches on, self-supervised learning has garnered increasing attention in recent years thanks to its remarkable representation learning ability and…