3 papers
cs.CL2026
PCoMoE: Shifting MoE Inference from Monolithic Expert Selection to Fine-Grained Path Composition
Ziyan Gan, Fangxin Liu, Chenyang Guan +10
Mixture-of-Experts (MoE) architectures scale Large Language Model (LLM) capacity efficiently by activating a sparse subset of experts per token. However, modern MoE inference remai…
cs.CR2025
LaMoS: Enabling Efficient Large Number Modular Multiplication through SRAM-based CiM Acceleration
Haomin Li, Fangxin Liu, Chenyang Guan +3
Barrett's algorithm is one of the most widely used methods for performing modular multiplication, a critical nonlinear operation in modern privacy computing techniques such as homo…
cs.LG2025
SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
Zhixiong Zhao, Fangxin Liu, Junjie Wang +4
The emergence of accurate open large language models (LLMs) has sparked a push for advanced quantization techniques to enable efficient deployment on end-user devices. In this pape…