1 paper · 1 filter
Benjamin L. Badger, Ethan Roland
Transformer-based large language models are in some respects limited by the quadratic time and space computational complexity of attention. We introduce the Toeplitz MLP Mixer (TMM…