collaborators

7 papers

cs.DC2026

WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching

Xiangchen Li, Jiakun Fan, Qingyuan Wang +7

As Large Language Models (LLMs) become increasingly accessible to end users, an ever-growing number of inference requests are initiated from edge devices and computed on centralize…

cs.CV2026

MoCrop: Training Free Motion Guided Cropping for Efficient Video Action Recognition

Binhua Huang, Wendong Yao, Shaowu Chen +3

Standard video action recognition models often process typically resized full frames, suffering from spatial redundancy and high computational costs. To address this, we introduce…

cs.CV2025

TinyDrop: Tiny Model Guided Token Dropping for Vision Transformers

Guoxin Wang, Qingyuan Wang, Binhua Huang +2

Vision Transformers (ViTs) achieve strong performance in image classification but incur high computational costs from processing all image tokens. To reduce inference costs in larg…

cs.CV2025

Optimal Brain Connection: Towards Efficient Structural Pruning

Shaowu Chen, Wei Ma, Binhua Huang +5

Structural pruning has been widely studied for its effectiveness in compressing neural networks. However, existing methods often neglect the interconnections among parameters. To a…

cs.CV2025

ORXE: Orchestrating Experts for Dynamically Configurable Efficiency

Qingyuan Wang, Guoxin Wang, Barry Cardiff +1

This paper presents ORXE, a modular and adaptable framework for achieving real-time configurable efficiency in AI models. By leveraging a collection of pre-trained experts with div…

cs.LG2025

DyCE: Dynamically Configurable Exiting for Deep Learning Compression and Real-time Scaling

Qingyuan Wang, Barry Cardiff, Antoine Frappé +2

Conventional deep learning (DL) model compression and scaling methods focus on altering the model's components, impacting the results across all samples uniformly. However, since s…