Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Contribution Weights: A Geometrical Analysis of Self-Attention Transformers
Harry Jake Cunningham, Nicola Muca Cirone
Analyzing attention weights has become a standard approach for interpreting the information flow of Large Language Models (LLMs). However, this approach has significant limitations…
cs.LG2024
Reparameterized Multi-Resolution Convolutions for Long Sequence Modelling
Harry Jake Cunningham, Giorgio Giannone, Mingtian Zhang +1
Global convolutions have shown increasing promise as powerful general-purpose sequence models. However, training long convolutions is challenging, and kernel parameterizations must…