3 papers
cs.LG2026
Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs
Kibum Kim, Jiwan Kim, Kyle Min +4
Video Large Language Models (Video LLMs) incur high inference latency due to a large number of visual tokens provided to LLMs. To address this, training-free visual token pruning h…
cs.CV2025
Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization
Shuo Xing, Peiran Li, Yuping Wang +6
The emergence of large Vision Language Models (VLMs) has broadened the scope and capabilities of single-modal Large Language Models (LLMs) by integrating visual modalities, thereby…
cs.AI2025
mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation
Chan-Wei Hu, Yueqi Wang, Shuo Xing +4
Large Vision-Language Models (LVLMs) have made remarkable strides in multimodal tasks such as visual question answering, visual grounding, and complex reasoning. However, they rema…