1 paper · 1 filter
Han-Byul Kim, Duc Hoang, Arnav Kundu +2
With the rapid expansion in the scale of large language models (LLMs), enabling efficient distributed inference across multiple computing units has become increasingly critical. Ho…