1 paper · 1 filter
Yonggan Fu, Xin Dong, Shizhe Diao +12
Efficient deployment of small language models (SLMs) is essential for numerous real-world applications with stringent latency constraints. While previous work on SLM design has pri…