2 papers
cs.RO2026
Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models
Yiran Ling, Qing Lian, Jinghang Li +6
In this paper, we propose GTA-VLA(Guide, Think, Act), an interactive Vision-Language-Action (VLA) framework that enables spatially steerable embodied reasoning by allowing users to…
cs.CV2024
The Instinctive Bias: Spurious Images lead to Illusion in MLLMs
Tianyang Han, Qing Lian, Rui Pan +5
Large language models (LLMs) have recently experienced remarkable progress, where the advent of multi-modal large language models (MLLMs) has endowed LLMs with visual capabilities,…