2 papers
stat.ML2025
Convergence of Shallow ReLU Networks on Weakly Interacting Data
Léo Dana, Francis Bach, Loucas Pillaud-Vivien
We analyse the convergence of one-hidden-layer ReLU networks trained by gradient flow on data points. Our main contribution leverages the high dimensionality of the ambient spa…
cs.AI2025
Memorization in Attention-only Transformers
Léo Dana, Muni Sreenivas Pydi, Yann Chevaleyre
Recent research has explored the memorization capacity of multi-head attention, but these findings are constrained by unrealistic limitations on the context size. We present a nove…