2 papers
cs.CV2026
Hyperbolic Hierarchical Clustering for Visual Representation Learning
Jianan Wei, Guikun Chen, Zhiyuan Weng +3
We investigate the token mixer in vision backbones by revisiting clustering, one of the most classic approaches in machine learning. An effective token mixer is a fundamental compo…
cs.CV2025
Seeing What Matters: Empowering CLIP with Patch Generation-to-Selection
Gensheng Pei, Tao Chen, Yujia Wang +4
The CLIP model has demonstrated significant advancements in aligning visual and language modalities through large-scale pre-training on image-text pairs, enabling strong zero-shot…