2 papers
cs.LG2026
QUEST: A robust attention formulation using query-modulated spherical attention
Hariprasath Govindarajan, Per Sidén, Jacob Roll +1
The Transformer model architecture has become one of the most widely used in deep learning and the attention mechanism is at its core. The standard attention formulation uses a sof…
cs.LG2025
Revisiting Likelihood-Based Out-of-Distribution Detection by Modeling Representations
Yifan Ding, Arturas Aleksandraus, Amirhossein Ahmadian +3
Out-of-distribution (OOD) detection is critical for ensuring the reliability of deep learning systems, particularly in safety-critical applications. Likelihood-based deep generativ…