1 citations · 1 across the 1 of their papers we have counts for
1 paper
June Yong Yang, Byeongwook Kim, Jeongin Bae +5
Key-Value (KV) Caching has become an essential technique for accelerating the inference speed and throughput of generative Large Language Models~(LLMs). However, the memory footpri…