1 paper · 1 filter
Pierre Marion, Raphaël Berthier, Gérard Biau +1
Attention-based models, such as Transformer, excel across various tasks but lack a comprehensive theoretical understanding, especially regarding token-wise sparsity and internal li…