3 papers
cs.AR2026
Hot-Cold Tiering of HBM and High Bandwidth Flash for Agentic LLM Serving
Jongjin Baek, Won Ji, Seungjae Yoo +1
Large language model (LLM) serving is increasingly agentic, with multi-turn sessions that idle between actions yet must retain their full context. Limited GPU memory capacity force…
cs.LG2026
CoX-MoE: Coalesced Expert Execution for High-Throughput MoE Inference with AMX-Enabled CPU-GPU Co-Execution
Muyoung Son, Yi Chen, Seungjae Yoo +2
The Mixture-of-Experts (MoE) architecture improves computational efficiency via sparse expert activation, but throughput-oriented inference faces substantial GPU memory pressure du…
cs.AR2025
RED: Energy Optimization Framework for eDRAM-based PIM with Reconfigurable Voltage Swing and Retention-aware Scheduling
Jae-Young Kim, Donghyuk Kim, Seungjae Yoo +3
In the era of artificial intelligence (AI), Transformer demonstrates its performance across various applications. The excessive amount of parameters incurs high latency and energy…