5 papers
Scalable Physics-Inspired Transformers for Spin Glasses
Lu Zhong, Wenli Duan, Jing Liu +2
Efficient sampling of the Boltzmann distribution in frustrated spin glasses is central to statistical mechanics and combinatorial optimization. Despite advances in machine-learning…
MODE: Modality-Decomposed Expert-Level Mixed-Precision Quantization for MoE Multimodal LLMs
Yuanteng Chen, Peisong Wang, Zhilei Liu +9
Mixture-of-Experts Multimodal Large Language Models (MoE-MLLMs) offer remarkable performance but incur prohibitive GPU memory costs, making compression essential. Among PTQ methods…
Distributed Specialization: Rare-Token Neurons in Large Language Models
Jing Liu, Haozheng Wang, Yueheng Li
Large language models (LLMs) struggle with representing and generating rare tokens despite their importance in specialized domains. We investigate whether LLMs develop internal spe…
No Clustering, No Routing: How Transformers Actually Process Rare Tokens
Jing Liu
Large language models struggle with rare token prediction, yet the mechanisms driving their specialization remain unclear. Prior work identified specialized ``plateau'' neurons for…
Emergent Specialization: Rare Token Neurons in Language Models
Jing Liu, Haozheng Wang, Yueheng Li
Large language models struggle with representing and generating rare tokens despite their importance in specialized domains. In this study, we identify neuron structures with excep…