3 papers
cs.CV2026
LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models
Zeyu Xu, Xingzhong Hou, Pengkai Guo +6
Vision-Language Models (VLMs) have achieved strong progress in multimodal understanding. However, scaling dense or sparse Mixture-of-Experts (MoE) models to improve performance lim…
cs.CV2025
E-VRAG: Enhancing Long Video Understanding with Resource-Efficient Retrieval Augmented Generation
Zeyu Xu, Junkang Zhang, Qiang Wang +1
Vision-Language Models (VLMs) have enabled substantial progress in video understanding by leveraging cross-modal reasoning capabilities. However, their effectiveness is limited by…
cs.CV2025
MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning
Yi Liu, Xiao Xu, Zeyu Xu +10
Vision-Language Models (VLMs) have achieved remarkable breakthroughs in recent years, enabling a diverse array of applications in everyday life. However, the substantial computatio…