2 papers
cs.LG2025
Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off
Mingkuan Zhao, Wentao Hu, Jiayin Wang +5
The design of Large Language Models (LLMs) has long been hampered by a fundamental conflict within their core attention mechanism: its remarkable expressivity is built upon a compu…
cs.LG2025
Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts Models
Wentao Hu, Mingkuan Zhao, Shuangyong Song +3
Sparse Mixture-of-Experts (SMoE) architectures have enabled a new frontier in scaling Large Language Models (LLMs), offering superior performance by activating only a fraction of t…