collaborators

7 papers

cs.AI2026

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning

Wencheng Ye, Yi Bin, Yujuan Ding +7

Vision-language models increasingly succeed on multimodal reasoning benchmarks, yet their visual evidence often becomes unstable once it enters the language stack, weakening eviden…

cs.CV2026

Multimodal Learning on Low-Quality Data with Conformal Predictive Self-Calibration

Xun Jiang, Yufan Gu, Disen Hu +5

Multimodal learning often grapples with the challenge of low-quality data, which predominantly manifests as two facets: modality imbalance and noisy corruption. While these issues…

cs.RO2026

Language-Grounded Decoupled Action Representation for Robotic Manipulation

Wuding Weng, Tongshu Wu, Liucheng Chen +5

The heterogeneity between high-level vision-language understanding and low-level action control remains a fundamental challenge in robotic manipulation. Although recent methods hav…

cs.CV2026

Truth in the Few: High-Value Data Selection for Efficient Multi-Modal Reasoning

Shenshen Li, Xing Xu, Kaiyuan Deng +3

While multi-modal large language models (MLLMs) have made significant progress in complex reasoning tasks via reinforcement learning, it is commonly believed that extensive trainin…

cs.RO2025

MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training

Zhenhan Yin, Xuanhan Wang, Jiahao Jiang +8

While leveraging abundant human videos and simulated robot data poses a scalable solution to the scarcity of real-world robot data, the generalization capability of existing vision…

cs.CV2025

HarmoCLIP: Harmonizing Global and Regional Representations in Contrastive Vision-Language Models

Haoxi Zeng, Haoxuan Li, Yi Bin +4

Contrastive Language-Image Pre-training (CLIP) has demonstrated remarkable generalization ability and strong performance across a wide range of vision-language tasks. However, due…