1 citations · 2 across the 8 of their papers we have counts for
Showing 2024Show all
2 papers · 1 filter
stat.ML2024
Attention layers provably solve single-location regression
Pierre Marion, Raphaël Berthier, Gérard Biau +1
Attention-based models, such as Transformer, excel across various tasks but lack a comprehensive theoretical understanding, especially regarding token-wise sparsity and internal li…
cs.LG2024
On the Minimal Degree Bias in Generalization on the Unseen for non-Boolean Functions
Denys Pushkin, Raphaël Berthier, Emmanuel Abbe
We investigate the out-of-domain generalization of random feature (RF) models and Transformers. We first prove that in the `generalization on the unseen (GOTU)' setting, where trai…