3 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.RO2026
CoINS: Counterfactual Interactive Navigation via Skill-Aware VLM
Kangjie Zhou, Zhejia Wen, Zhiyong Zhuo +9
Recent Vision-Language Models (VLMs) have demonstrated significant potential in robotic planning. However, they typically function as semantic reasoners, lacking an intrinsic under…
cs.CV2024★ 3 cited
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
Qizhe Zhang, Aosong Cheng, Ming Lu +6
Large vision-language models (LVLMs) generally contain significantly more visual tokens than their textual counterparts, resulting in a considerable computational burden. Recent ef…