1 paper · 1 filter
Hee Seung Hwang, Xindi Wu, Sanghyuk Chun +1
Fast weight architectures offer a promising alternative to attention-based transformers for long-context modeling by maintaining constant memory overhead regardless of context leng…