collaborators

8 papers

cs.OS2026

Benchmarking LLMs on File System Design and Implementation

Yuqi Xue, Daixuan Li, Jian Huang

Large Language Models (LLMs) are fundamentally transforming computer system research and development. As we employ LLMs in file system (fs) development, it is essential to understa…

cs.AR2026

Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving

Yuqi Xue, Jichuan Chang, Jian Huang

To meet the ever-increasing computing demands of large language model (LLM) services, modern cloud platforms have widely deployed neural processing units (NPUs). These NPU chips ha…

cs.AR2026

Enabling Spatially Fine-Grained DVFS in Neural Processing Units for Energy-Efficient LLM Serving

Yuqi Xue, Jerry Wu, Corey Yu +1

As neural processing units (NPUs) evolve rapidly to accommodate the ever-increasing compute demand of large language models (LLMs), their power consumption is becoming a limiting f…

cs.LG2026

Vegas: Self-Speculative Decoding with Verification-Guided Sparse Attention

Yikang Yue, Yuqi Xue, Jian Huang

Long-context large language model (LLM) inference has become the norm for today's AI applications. However, it is severely bottlenecked by the increasing memory demands of its KV c…

cs.CE2025

Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs

Xingang Guo, Yaxin Li, Xiangyi Kong +62

Modern engineering, spanning electrical, mechanical, aerospace, civil, and computer disciplines, stands as a cornerstone of human civilization and the foundation of our society. Ho…

cs.AR2025

ReGate: Enabling Power Gating in Neural Processing Units

Yuqi Xue, Jian Huang

The energy efficiency of neural processing units (NPU) is playing a critical role in developing sustainable data centers. Our study with different generations of NPU chips reveals…