7 papers
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
Xiangchen Li, Jiakun Fan, Qingyuan Wang +7
As Large Language Models (LLMs) become increasingly accessible to end users, an ever-growing number of inference requests are initiated from edge devices and computed on centralize…
MoCrop: Training Free Motion Guided Cropping for Efficient Video Action Recognition
Binhua Huang, Wendong Yao, Shaowu Chen +3
Standard video action recognition models often process typically resized full frames, suffering from spatial redundancy and high computational costs. To address this, we introduce…
TinyDrop: Tiny Model Guided Token Dropping for Vision Transformers
Guoxin Wang, Qingyuan Wang, Binhua Huang +2
Vision Transformers (ViTs) achieve strong performance in image classification but incur high computational costs from processing all image tokens. To reduce inference costs in larg…
Optimal Brain Connection: Towards Efficient Structural Pruning
Shaowu Chen, Wei Ma, Binhua Huang +5
Structural pruning has been widely studied for its effectiveness in compressing neural networks. However, existing methods often neglect the interconnections among parameters. To a…
ORXE: Orchestrating Experts for Dynamically Configurable Efficiency
Qingyuan Wang, Guoxin Wang, Barry Cardiff +1
This paper presents ORXE, a modular and adaptable framework for achieving real-time configurable efficiency in AI models. By leveraging a collection of pre-trained experts with div…
DyCE: Dynamically Configurable Exiting for Deep Learning Compression and Real-time Scaling
Qingyuan Wang, Barry Cardiff, Antoine Frappé +2
Conventional deep learning (DL) model compression and scaling methods focus on altering the model's components, impacting the results across all samples uniformly. However, since s…