Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
EmbRACE-3K: Embodied Reasoning and Action in Complex Environments
Mingxian Lin, Wei Huang, Yitang Li +6
Recent advanced vision-language models(VLMs) have demonstrated strong performance on passive, offline image and video understanding tasks. However, their effectiveness in embodied…
cs.CV2025
VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models
Kui Wu, Shuhang Xu, Hao Chen +4
We introduce a novel self-improving framework that enhances Embodied Visual Tracking (EVT) with Vision-Language Models (VLMs) to address the limitations of current active visual tr…
cs.CV2025
Hierarchical Instruction-aware Embodied Visual Tracking
Kui Wu, Hao Chen, Churan Wang +4
User-Centric Embodied Visual Tracking (UC-EVT) presents a novel challenge for reinforcement learning-based models due to the substantial gap between high-level user instructions an…