2 papers
cs.CV2026
Finding Distributed Object-Centric Properties in Self-Supervised Transformers
Samyak Rawlekar, Amitabh Swain, Yujun Cai +3
Self-supervised Vision Transformers (ViTs) like DINO show an emergent ability to discover objects, typically observed in [CLS] token attention maps of the final layer. However, the…
cs.CV2025
Efficiently Disentangling CLIP for Multi-Object Perception
Samyak Rawlekar, Yujun Cai, Yiwei Wang +2
Vision-language models like CLIP excel at recognizing the single, prominent object in a scene. However, they struggle in complex scenes containing multiple objects. We identify a f…