3 papers
cs.LG2026
QUEST: A robust attention formulation using query-modulated spherical attention
Hariprasath Govindarajan, Per Sidén, Jacob Roll +1
The Transformer model architecture has become one of the most widely used in deep learning and the attention mechanism is at its core. The standard attention formulation uses a sof…
cs.LG2024
On Partial Prototype Collapse in the DINO Family of Self-Supervised Methods
Hariprasath Govindarajan, Per Sidén, Jacob Roll +1
A prominent self-supervised learning paradigm is to model the representations as clusters, or more generally as a mixture model. Learning to map the data samples to compact represe…
cs.LG2024
DINO as a von Mises-Fisher mixture model
Hariprasath Govindarajan, Per Sidén, Jacob Roll +1
Self-distillation methods using Siamese networks are popular for self-supervised pre-training. DINO is one such method based on a cross-entropy loss between -dimensional probabi…