1 citations · 1 across the 1 of their papers we have counts for
1 paper
Adnan Hoque, Mudhakar Srivatsa, Chih-Chieh Yang +1
In this paper, we present a novel method that reduces model inference latency during distributed deployment of Large Language Models (LLMs). Our contribution is an optimized infere…