1 paper
Haotian Zheng, Zhanwei Wang, Mingyao Cui +3
Speculative inference (SPIN) was originally developed as an efficient architecture to accelerate Large Language Models (LLMs). In this work, we propose its distributed deployment t…