1 paper · 1 filter
Huy Nguyen, Pedram Akbarian, Fanqi Yan +1
Top-K sparse softmax gating mixture of experts has been widely used for scaling up massive deep-learning architectures without increasing the computational cost. Despite its popula…