1 paper
Lu Ye, Ze Tao, Yong Huang +1
Self-attention is an essential component of large language models (LLM) but a significant source of inference latency for long sequences. In multi-tenant LLM serving scenarios, the…