31 citations · 35 across the 13 of their papers we have counts for
4 papers · 1 filter
NanoSpec: Accelerating Speculative Decoding using Minimalist In-Context Vocabularies
Zhiyang Chen, Daliang Xu, Yinyuan Zhang +3
The massive vocabulary sizes of large language models, often exceeding 100k tokens, impose a computational bottleneck on the final linear projection layer during speculative decodi…
Accelerating Mobile Language Model via Speculative Decoding and NPU-Coordinated Execution
Zhiyang Chen, Daliang Xu, Haiyang Shen +5
Performing Retrieval-Augmented Generation (RAG) directly on mobile devices is promising for data privacy and responsiveness but is hindered by the architectural constraints of mobi…
PhoneLM:an Efficient and Capable Small Language Model Family through Principled Pre-training
Rongjie Yi, Xiang Li, Weikai Xie +6
The interest in developing small language models (SLM) for on-device deployment is fast growing. However, the existing SLM design hardly considers the device hardware characteristi…
Small Language Models: Survey, Measurements, and Insights
Zhenyan Lu, Xiang Li, Dongqi Cai +5
Small language models (SLMs), despite their widespread adoption in modern smart devices, have received significantly less academic attention compared to their large language model…