9 papers
Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving
Yuqi Xue, Jichuan Chang, Jian Huang
To meet the ever-increasing computing demands of large language model (LLM) services, modern cloud platforms have widely deployed neural processing units (NPUs). These NPU chips ha…
Enabling Spatially Fine-Grained DVFS in Neural Processing Units for Energy-Efficient LLM Serving
Yuqi Xue, Jerry Wu, Corey Yu +1
As neural processing units (NPUs) evolve rapidly to accommodate the ever-increasing compute demand of large language models (LLMs), their power consumption is becoming a limiting f…
Vegas: Self-Speculative Decoding with Verification-Guided Sparse Attention
Yikang Yue, Yuqi Xue, Jian Huang
Long-context large language model (LLM) inference has become the norm for today's AI applications. However, it is severely bottlenecked by the increasing memory demands of its KV c…
Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with Voxel
Yiqi Liu, Noelle Crawford, Michael Wang +2
To overcome the well-known memory bottleneck of AI chips, 3D stacked architectures that employ advanced packaging technology with high-density through-silicon vias (TSVs) pins have…
ReGate: Enabling Power Gating in Neural Processing Units
Yuqi Xue, Jian Huang
The energy efficiency of neural processing units (NPU) is playing a critical role in developing sustainable data centers. Our study with different generations of NPU chips reveals…
ELK: Exploring the Efficiency of Inter-core Connected AI Chips with Deep Learning Compiler Techniques
Yiqi Liu, Yuqi Xue, Noelle Crawford +2
To meet the increasing demand of deep learning (DL) models, AI chips are employing both off-chip memory (e.g., HBM) and high-bandwidth low-latency interconnect for direct inter-cor…