2 papers
cs.AR2026
PATTON: Enabling Commodity PIM for Production LLM Serving
Hangyeol Kim, Sanghyun Lee, Teokkyu Suh +1
Processing-in-Memory (PIM) is promising for accelerating memory-bound decode attention, but attention acceleration alone is insufficient for production LLM serving, where engines d…
cs.AR2025
RED: Energy Optimization Framework for eDRAM-based PIM with Reconfigurable Voltage Swing and Retention-aware Scheduling
Jae-Young Kim, Donghyuk Kim, Seungjae Yoo +3
In the era of artificial intelligence (AI), Transformer demonstrates its performance across various applications. The excessive amount of parameters incurs high latency and energy…