1 paper
S M Rafiuddin, Muntaha Nujat Khan
Transformer attention scales quadratically with sequence length O(n^2), limiting long-context use. We propose Adaptive Retention, a probabilistic, layer-wise token selection mechan…