2 papers
cs.CL2025
FlatQuant: Flatness Matters for LLM Quantization
Yuxuan Sun, Ruikang Liu, Haoli Bai +10
Recently, quantization has been widely used for the compression and acceleration of large language models (LLMs). Due to the outliers in LLMs, it is crucial to flatten weights and…
cs.CL2025
Cluster-Driven Expert Pruning for Mixture-of-Experts Large Language Models
Hongcheng Guo, Juntao Yao, Boyang Wang +5
Mixture-of-Experts (MoE) architectures have emerged as a promising paradigm for scaling large language models (LLMs) with sparse activation of task-specific experts. Despite their…