1 paper
Hyun-rae Jo, Dongkun Shin
Recently, large language models (LLM) based on transformers are facing memory bottleneck issues due to KV cache, especially in long sequence handling. Previous researches proposed…