1 paper
Megan Flynn, Alexander Wang, Dean Edward Alvarez +2
We present STAT: a simple algorithm to prune transformer models without any fine-tuning. STAT eliminates both attention heads and neurons from the network, while preserving accurac…