Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Simplifying DINO via Coding Rate Regularization
Ziyang Wu, Jingyuan Zhang, Druv Pai +5
DINO and DINOv2 are two model families being widely used to learn representations from unlabeled imagery data at large scales. Their learned representations often enable state-of-t…
cs.CV2025
Scaling White-Box Transformers for Vision
Jinrui Yang, Xianhang Li, Druv Pai +4
CRATE, a white-box transformer architecture designed to learn compressed and sparse representations, offers an intriguing alternative to standard vision transformers (ViTs) due to…