3 papers
cs.CL2026
PCoMoE: Shifting MoE Inference from Monolithic Expert Selection to Fine-Grained Path Composition
Ziyan Gan, Fangxin Liu, Chenyang Guan +10
Mixture-of-Experts (MoE) architectures scale Large Language Model (LLM) capacity efficiently by activating a sparse subset of experts per token. However, modern MoE inference remai…
cs.LG2025
LCD: Advancing Extreme Low-Bit Clustering for Large Language Models via Knowledge Distillation
Fangxin Liu, Ning Yang, Junping Zhao +3
Large language models (LLMs) have achieved significant progress in natural language processing but face challenges in deployment due to high memory and computational requirements.…
cs.CL2025
DASH: Input-Aware Dynamic Layer Skipping for Efficient LLM Inference with Markov Decision Policies
Ning Yang, Fangxin Liu, Junjie Wang +4
Large language models (LLMs) have achieved remarkable performance across a wide range of NLP tasks. However, their substantial inference cost poses a major barrier to real-world de…