2 papers
cs.LG2023
Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference
Ranggi Hwang, Jianyu Wei, Shijie Cao +4
Large language models (LLMs) based on transformers have made significant strides in recent years, the success of which is driven by scaling up their model size. Despite their high…
cs.LG2023
LUT-NN: Empower Efficient Neural Network Inference with Centroid Learning and Table Lookup
Xiaohu Tang, Yang Wang, Ting Cao +5
On-device Deep Neural Network (DNN) inference consumes significant computing resources and development efforts. To alleviate that, we propose LUT-NN, the first system to empower in…