7 papers
HBF Sucks? A Full-Stack Characterization of High-Bandwidth Flash for KV-Centric LLM Serving
Zhuoran Li, Zhuohang Bian, Xin Huang +3
A faster storage device should make serving faster. We find the opposite. High-Bandwidth Flash (HBF) stacks NAND behind a wide, package-local interface, promising flash-scale capac…
TokenStack: A Heterogeneous HBM-PIM Architecture and Runtime for Efficient LLM Inference
Zhuoran Li, Zhuohang Bian, Zihao Huang +5
Large language model (LLM) serving is now limited by the key-value (KV) cache. During decode, each new token rereads prior KV state, so attention becomes a bandwidth- and capacity-…
TAMI-MPC:Trusted Acceleration of Minimal-Interaction MPC for Efficient Nonlinear Inference
Zhuoran Li, Hanieh Totonchi Asl, Yifei Cai +2
Secure multi-party computation (MPC) offers a practical foundation for privacy-preserving machine learning at the edge. However, current MPC systems rely heavily on communication a…
SecDTD: Dynamic Token Drop for Secure Transformers Inference
Yifei Cai, Zhuoran Li, Yizhou Feng +4
The rapid adoption of Transformer-based AI has been driven by accessible models such as ChatGPT, which provide API-based services for developers and businesses. However, as these o…
SLOTH: Lightweight Detection and Localization of On-Chip Fail-Slow Failures for DNN Accelerators
Junchi Wu, Xinfei Wan, Zhuoran Li +5
Spatial DNN accelerators are essential for high-performance inference, but their performance is undermined by widespread fail-slow failures. Detecting such failures on-chip is chal…
Large Model Enabled Embodied Intelligence for 6G Integrated Perception, Communication, and Computation Network
Zhuoran Li, Zhen Gao, Xinhua Liu +6
The advent of sixth-generation (6G) places intelligence at the core of wireless architecture, fusing perception, communication, and computation into a single closed-loop. This pape…