12 papers
Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models?
Guoheng Sun, Kaixi Feng, Shwai He +8
Vision-Language-Action (VLA) models enable instruction-driven robotic manipulation, but they inherit oversized language backbones from pretrained VLMs whose capacity far exceeds wh…
Demystifying When Pruning Works via Representation Hierarchies
Shwai He, Guoheng Sun, Haichao Zhang +2
Network pruning, which removes less important parameters or architectures, is often expected to improve efficiency while preserving performance. However, this expectation does not…
Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts
Shwai He, Weilin Cai, Jiayi Huang +1
The Mixture of Experts (MoE) is an effective architecture for scaling large language models by leveraging sparse expert activation to balance performance and efficiency. However, u…
MoEless: Efficient MoE LLM Serving via Serverless Computing
Hanfei Yu, Bei Ouyang, Shwai He +2
Large Language Models (LLMs) have become a cornerstone of AI, driving progress across diverse domains such as content creation, search and recommendation systems, and AI-assisted w…
Making Large Language Models Efficient Dense Retrievers
Yibin Lei, Shwai He, Ang Li +1
Recent work has shown that directly fine-tuning large language models (LLMs) for dense retrieval yields strong performance, but their substantial parameter counts make them computa…
Understanding and Harnessing Sparsity in Unified Multimodal Models
Shwai He, Chaorui Deng, Ang Li +1
Large multimodal models have achieved remarkable progress in both understanding and generation. Recent efforts pursue unified multimodal models that integrate heterogeneous compone…