2 papers
cs.DC2026
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
Fan Bai, Pai Peng, Zhengzhi Tang +8
With the widespread adoption of large multimodal models, efficient inference across text, image, audio, and video modalities has become critical. However, existing multimodal infer…
cs.CV2025
Multi-Cue Adaptive Visual Token Pruning for Large Vision-Language Models
Bozhi Luan, Wengang Zhou, Hao Feng +3
As the computational needs of Large Vision-Language Models (LVLMs) increase, visual token pruning has proven effective in improving inference speed and memory efficiency. Tradition…