182 citations · 194 across the 4 of their papers we have counts for
1 paper · 1 filter
Yunho Jin, Chun-Feng Wu, David Brooks +1
Generating texts with a large language model (LLM) consumes massive amounts of memory. Apart from the already-large model parameters, the key/value (KV) cache that holds informatio…