23 citations · 23 across the 2 of their papers we have counts for
2 papers
cs.CL2026
PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference
Niqi Lyu, Pengtao Shi, Wei Qiu +4
Large language models (LLMs) provide strong reasoning capabilities but are expensive to serve at scale, whereas small language models (SLMs) are cheaper but less reliable on diffic…
cs.LG2021★ 23 cited
MQBench: Towards Reproducible and Deployable Model Quantization Benchmark
Yuhang Li, Mingzhu Shen, Jian Ma +6
Model quantization has emerged as an indispensable technique to accelerate deep learning inference. While researchers continue to push the frontier of quantization algorithms, exis…