2 papers
cs.CL2026
Meta-Learning Where to Allocate Experts: Task-Conditioned Layer-Wise Compression for MoEs
Rongfeng Wang, Shichao Weng, Zhiqiang Wang +4
Mixture-of-Experts (MoE) models route each token to a subset of expert networks, increasing capacity while keeping per-token computation sparse. In many deployed MoEs, the number o…
cs.LG2026
Resource-Efficient Pruning for Transformer via Low-Rank Importance Estimation
Peng Liu, Huibing Zeng, Yiqun Zhang +2
With the rapid development of large-scale pre-trained language models based on Transformer architectures, their high computational and memory costs have become a major obstacle to…