8 papers
Benchmarking LLMs on File System Design and Implementation
Yuqi Xue, Daixuan Li, Jian Huang
Large Language Models (LLMs) are fundamentally transforming computer system research and development. As we employ LLMs in file system (fs) development, it is essential to understa…
Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving
Yuqi Xue, Jichuan Chang, Jian Huang
To meet the ever-increasing computing demands of large language model (LLM) services, modern cloud platforms have widely deployed neural processing units (NPUs). These NPU chips ha…
Enabling Spatially Fine-Grained DVFS in Neural Processing Units for Energy-Efficient LLM Serving
Yuqi Xue, Jerry Wu, Corey Yu +1
As neural processing units (NPUs) evolve rapidly to accommodate the ever-increasing compute demand of large language models (LLMs), their power consumption is becoming a limiting f…
Vegas: Self-Speculative Decoding with Verification-Guided Sparse Attention
Yikang Yue, Yuqi Xue, Jian Huang
Long-context large language model (LLM) inference has become the norm for today's AI applications. However, it is severely bottlenecked by the increasing memory demands of its KV c…
Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs
Xingang Guo, Yaxin Li, Xiangyi Kong +62
Modern engineering, spanning electrical, mechanical, aerospace, civil, and computer disciplines, stands as a cornerstone of human civilization and the foundation of our society. Ho…
ReGate: Enabling Power Gating in Neural Processing Units
Yuqi Xue, Jian Huang
The energy efficiency of neural processing units (NPU) is playing a critical role in developing sustainable data centers. Our study with different generations of NPU chips reveals…