1 paper
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh +4
LLMs are seeing growing use for applications which require large context windows, and with these large context windows KV cache activations surface as the dominant contributor to m…