Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Sinkhorn doubly stochastic attention rank decay analysis
Michela Lapenna, Rita Fioresi, Bahman Gharesifard
The self-attention mechanism is central to the success of Transformer architectures. However, standard row-stochastic attention has been shown to suffer from significant signal deg…
cs.LG2023
Geometric Deep Learning: a Temperature Based Analysis of Graph Neural Networks
M. Lapenna, F. Faglioni, F. Zanchetta +1
We examine a Geometric Deep Learning model as a thermodynamic system treating the weights as non-quantum and non-relativistic particles. We employ the notion of temperature previou…