5 papers
LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models
Zeyu Xu, Xingzhong Hou, Pengkai Guo +6
Vision-Language Models (VLMs) have achieved strong progress in multimodal understanding. However, scaling dense or sparse Mixture-of-Experts (MoE) models to improve performance lim…
Test-Time Perturbation Learning with Delayed Feedback for Vision-Language-Action Models
Zehua Zang, Xi Wang, Fuchun Sun +4
Vision-Language-Action models (VLAs) achieve remarkable performance in sequential decision-making but remain fragile to subtle environmental shifts, such as small changes in object…
MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning
Yi Liu, Xiao Xu, Zeyu Xu +10
Vision-Language Models (VLMs) have achieved remarkable breakthroughs in recent years, enabling a diverse array of applications in everyday life. However, the substantial computatio…
V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization
Yuxi Xie, Guanzhen Li, Xiao Xu +1
Large vision-language models (LVLMs) suffer from hallucination, resulting in misalignment between the output textual response and the input visual content. Recent research indicate…
COrAL: Order-Agnostic Language Modeling for Efficient Iterative Refinement
Yuxi Xie, Anirudh Goyal, Xiaobao Wu +5
Iterative refinement has emerged as an effective paradigm for enhancing the capabilities of large language models (LLMs) on complex tasks. However, existing approaches typically im…