3 papers
cs.CV2025
ShaRP: SHAllow-LayeR Pruning for Efficient Video Large Language Models
Yingjie Xia, Tao Liu, Jinglei Shi +4
Video Large Language Models (VLLMs) incur substantial prefilling cost due to the large number of visual tokens. While attention-based token pruning offers a promising acceleration…
cs.CV2025
AndesVL Technical Report: An Efficient Mobile-side Multimodal Large Language Model
Zhiwei Jin, Xiaohui Song, Nan Wang +36
In recent years, while cloud-based MLLMs such as QwenVL, InternVL, GPT-4o, Gemini, and Claude Sonnet have demonstrated outstanding performance with enormous model sizes reaching hu…
cs.LG2025
HMI: Hierarchical Knowledge Management for Efficient Multi-Tenant Inference in Pretrained Language Models
Jun Zhang, Jue Wang, Huan Li +6
The significant computational demands of pretrained language models (PLMs), which often require dedicated hardware, present a substantial challenge in serving them efficiently, esp…