10 papers
AMIEOD: Adaptive Multi-Experts Image Enhancement for Object Detection in Low-Illumination Scenes
Xiaochen Huang, Honggang Chen, Weicheng Zhang +4
In multimedia application scenarios, images captured under low-illumination conditions often lead to lower accuracy in visual perception tasks compared to those taken in well-lit e…
Where and How to Prune: An Empirical Study of Visual Token Pruning for GUI Agent Navigation
Daiqiang Li, Zihao Pan, Zeyu Zhang +8
In recent years, GUI agents have demonstrated strong potential in navigation tasks. However, preserving complete historical screenshots introduces substantial computational overhea…
Variation-aware Vision Token Dropping for Faster Large Vision-Language Models
Junjie Chen, Xuyang Liu, Zichen Wen +3
Large vision-language models (LVLMs) have demonstrated remarkable capabilities in multimodal understanding tasks. However, the increasing demand for high-resolution image and long-…
Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models
Xuyang Liu, Ziming Wang, Junjie Chen +6
Large vision-language models (LVLMs) excel at visual understanding, but face efficiency challenges due to quadratic complexity in processing long multi-modal contexts. While token…
LEMUR: Large scale End-to-end MUltimodal Recommendation
Xintian Han, Honggang Chen, Quan Lin +14
Traditional ID-based recommender systems often struggle with cold-start and generalization challenges. Multimodal recommendation systems, which leverage textual and visual data, of…
Filter, Correlate, Compress: Training-Free Token Reduction for MLLM Acceleration
Yuhang Han, Xuyang Liu, Zihan Zhang +6
The quadratic complexity of Multimodal Large Language Models (MLLMs) with respect to context length poses significant computational and memory challenges, hindering their real-worl…