activity
20242026
collaborators

9 papers

cs.AR2026

Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving

Yuqi Xue, Jichuan Chang, Jian Huang

To meet the ever-increasing computing demands of large language model (LLM) services, modern cloud platforms have widely deployed neural processing units (NPUs). These NPU chips ha…

cs.AR2026

Enabling Spatially Fine-Grained DVFS in Neural Processing Units for Energy-Efficient LLM Serving

Yuqi Xue, Jerry Wu, Corey Yu +1

As neural processing units (NPUs) evolve rapidly to accommodate the ever-increasing compute demand of large language models (LLMs), their power consumption is becoming a limiting f…

cs.LG2026

Vegas: Self-Speculative Decoding with Verification-Guided Sparse Attention

Yikang Yue, Yuqi Xue, Jian Huang

Long-context large language model (LLM) inference has become the norm for today's AI applications. However, it is severely bottlenecked by the increasing memory demands of its KV c…

cs.AR2026

Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with Voxel

Yiqi Liu, Noelle Crawford, Michael Wang +2

To overcome the well-known memory bottleneck of AI chips, 3D stacked architectures that employ advanced packaging technology with high-density through-silicon vias (TSVs) pins have…

cs.AR2025

ReGate: Enabling Power Gating in Neural Processing Units

Yuqi Xue, Jian Huang

The energy efficiency of neural processing units (NPU) is playing a critical role in developing sustainable data centers. Our study with different generations of NPU chips reveals…

cs.AR2025

ELK: Exploring the Efficiency of Inter-core Connected AI Chips with Deep Learning Compiler Techniques

Yiqi Liu, Yuqi Xue, Noelle Crawford +2

To meet the increasing demand of deep learning (DL) models, AI chips are employing both off-chip memory (e.g., HBM) and high-bandwidth low-latency interconnect for direct inter-cor…