7 citations · 7 across the 3 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
QUEST: A robust attention formulation using query-modulated spherical attention
Hariprasath Govindarajan, Per Sidén, Jacob Roll +1
The Transformer model architecture has become one of the most widely used in deep learning and the attention mechanism is at its core. The standard attention formulation uses a sof…
cs.LG2024
On Partial Prototype Collapse in the DINO Family of Self-Supervised Methods
Hariprasath Govindarajan, Per Sidén, Jacob Roll +1
A prominent self-supervised learning paradigm is to model the representations as clusters, or more generally as a mixture model. Learning to map the data samples to compact represe…