1 paper
Yifei Wang, Tianlin Li, Xiaohan Zhang +3
Inference optimization is a vital technique for deploying LLMs at scale. Compilation is the most widely adopted optimization technique for LLMs. While it assumes semantic equivalen…