3 citations · 3 across the 6 of their papers we have counts for
Showing 2026Show all
3 papers · 1 filter
cs.CL2026
Fitting Is Not Enough: Smoothness in Extremely Quantized LLMs
Yuzhuang Xu, Xu Han, Yuxuan Li +2
Large language models (LLMs) achieve strong performance but incur high deployment costs, motivating extremely low-bit but lossy quantization. Existing quantization algorithms mainl…
cs.CL2026
HUOZIIME: An On-Device LLM-enhanced Input Method for Deep Personalization
Baocai Shan, Yuzhuang Xu, Wanxiang Che
Mobile input method editors (IMEs) are the primary interface for text input, yet they remain constrained to manual typing and struggle to produce personalized text. While lightweig…
cs.DC2026
ArcLight: A Lightweight LLM Inference Architecture for Many-Core CPUs
Yuzhuang Xu, Xu Han, Yuxuan Li +1
Although existing frameworks for large language model (LLM) inference on CPUs are mature, they fail to fully exploit the computation potential of many-core CPU platforms. Many-core…