7 citations · 7 across the 2 of their papers we have counts for
1 paper · 1 filter
Gopi Krishna Jha, Sameh Gobriel, Liubov Talamanova +1
Key-value (KV) caching has emerged as a crucial optimization technique for accelerating inference in large language models (LLMs). By allowing the attention operation to scale line…