1 paper · 1 filter
Debabrata Mahapatra, Shubham Agarwal, Apoorv Saxena +1
Prior work on input-token importance in auto-regressive transformers has relied on Softmax-normalized attention weights, which obscure the richer structure of pre-Softmax query-key…