efficient training 1large language models 1residual stream expansion 1sparse computation 1transformer architectures 1
From the 1 of 8 linked papers with an AI index.
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
When Does Sparse MoE Help in Vision? The Role of Backbone Compute Leverage in Sparse Routing
Libo Sun, Po-wei Harn, Peixiong He +1
Mixture-of-Experts (MoE) networks promise favorable accuracy-compute trade-offs, yet practical vision deployments are hindered by expert collapse and limited end-to-end efficiency…
cs.CV2026
FineRMoE: Dimension Expansion for Finer-Grained Expert with Its Upcycling Approach
Ning Liao, Xiaoxing Wang, Xiaohan Qin +1
As revealed by the scaling law of fine-grained MoE, model performance ceases to be improved once the granularity of the intermediate dimension exceeds the optimal threshold, limiti…