1 paper · 1 filter
Kang Chen, Minshen Yu, Junjie Nian +3
In sparse Mixture-of-Experts language models, does the same token id imply the same router state and the same experts producing it? Holding the emitted token id fixed at repeated a…