Showing 2026Show all
2 papers · 1 filter
cs.CV2026
LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models
Zeyu Xu, Xingzhong Hou, Pengkai Guo +6
Vision-Language Models (VLMs) have achieved strong progress in multimodal understanding. However, scaling dense or sparse Mixture-of-Experts (MoE) models to improve performance lim…
cs.CV2026
Test-Time Perturbation Learning with Delayed Feedback for Vision-Language-Action Models
Zehua Zang, Xi Wang, Fuchun Sun +4
Vision-Language-Action models (VLAs) achieve remarkable performance in sequential decision-making but remain fragile to subtle environmental shifts, such as small changes in object…