1 paper
Shuming Shi, Enbo Zhao, Deng Cai +3
We present Inferflow, an efficient and highly configurable inference engine for large language models (LLMs). With Inferflow, users can serve most of the common transformer models…