1 paper
Hangyeol Kim, Sanghyun Lee, Teokkyu Suh +1
Processing-in-Memory (PIM) is promising for accelerating memory-bound decode attention, but attention acceleration alone is insufficient for production LLM serving, where engines d…