7 papers
The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning
Wencheng Ye, Yi Bin, Yujuan Ding +7
Vision-language models increasingly succeed on multimodal reasoning benchmarks, yet their visual evidence often becomes unstable once it enters the language stack, weakening eviden…
Multimodal Learning on Low-Quality Data with Conformal Predictive Self-Calibration
Xun Jiang, Yufan Gu, Disen Hu +5
Multimodal learning often grapples with the challenge of low-quality data, which predominantly manifests as two facets: modality imbalance and noisy corruption. While these issues…
Language-Grounded Decoupled Action Representation for Robotic Manipulation
Wuding Weng, Tongshu Wu, Liucheng Chen +5
The heterogeneity between high-level vision-language understanding and low-level action control remains a fundamental challenge in robotic manipulation. Although recent methods hav…
Truth in the Few: High-Value Data Selection for Efficient Multi-Modal Reasoning
Shenshen Li, Xing Xu, Kaiyuan Deng +3
While multi-modal large language models (MLLMs) have made significant progress in complex reasoning tasks via reinforcement learning, it is commonly believed that extensive trainin…
MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
Zhenhan Yin, Xuanhan Wang, Jiahao Jiang +8
While leveraging abundant human videos and simulated robot data poses a scalable solution to the scarcity of real-world robot data, the generalization capability of existing vision…
HarmoCLIP: Harmonizing Global and Regional Representations in Contrastive Vision-Language Models
Haoxi Zeng, Haoxuan Li, Yi Bin +4
Contrastive Language-Image Pre-training (CLIP) has demonstrated remarkable generalization ability and strong performance across a wide range of vision-language tasks. However, due…