2 citations · 2 across the 4 of their papers we have counts for
1 paper · 1 filter
Tobias Falke, Nicolas Anastassacos, Samson Tan +6
Sparse Mixture-of-Experts (MoE) architectures are increasingly popular for frontier large language models (LLM) but they introduce training challenges due to routing complexity. Fu…