1 paper · 1 filter
Reza Sedghi, Robin Schiewer, Anand Subramoney +1
At typical context lengths, the feed-forward MLP block accounts for a large share of a transformer's compute budget, motivating sparse alternatives to dense MLP blocks. We study sp…