3 papers
cs.CV2026
Studying Image Tokenizers as Visual Languages in Unified Multimodal Models
Siting Li, Zhengyang Wang, Simon Shaolei Du +2
Image tokenizers define the ``visual language'' of unified multimodal models, yet are commonly studied through isolated metrics or generation-/understanding-only evaluations. These…
cs.CV2024
Unveiling the Backbone-Optimizer Coupling Bias in Visual Representation Learning
Siyuan Li, Juanxi Tian, Zedong Wang +6
This paper delves into the interplay between vision backbones and optimizers, unvealing an inter-dependent phenomenon termed \textit{\textbf{b}ackbone-\textbf{o}ptimizer \textbf{c}…
cs.CV2024
TopoFR: A Closer Look at Topology Alignment on Face Recognition
Jun Dan, Yang Liu, Jiankang Deng +4
The field of face recognition (FR) has undergone significant advancements with the rise of deep learning. Recently, the success of unsupervised learning and graph neural networks h…