Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource
Houyi Li, Ka Man Lo, Shijie Xuyang +7
Mixture-of-Experts (MoE) language models dramatically expand model capacity and achieve remarkable performance without increasing per-token compute. However, can MoEs surpass dense…
cs.CL2025
Is Compression Really Linear with Code Intelligence?
Shijie Xuyang, Xianzhen Luo, Zheng Chu +6
Understanding the relationship between data compression and the capabilities of Large Language Models (LLMs) is crucial, especially in specialized domains like code intelligence. P…