2 papers
cs.LG2025
MoE Pathfinder: Trajectory-driven Expert Pruning
Xican Yang, Yuanhe Tian, Yan Song
Mixture-of-experts (MoE) architectures used in large language models (LLMs) achieve state-of-the-art performance across diverse tasks yet face practical challenges such as deployme…
cs.CL2025
Frustratingly Easy Task-aware Pruning for Large Language Models
Yuanhe Tian, Junjie Liu, Xican Yang +2
Pruning provides a practical solution to reduce the resources required to run large language models (LLMs) to benefit from their effective capabilities as well as control their cos…