2 papers
cs.CV2026
GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
Tong Wei, Yijun Yang, Changhao Zhang +4
Multi-turn reinforcement learning (RL) for multi-modal agents built upon vision-language models (VLMs) is hampered by sparse rewards and long-horizon credit assignment. Recent meth…
cs.AI2026
Hi-Agent: Hierarchical Vision-Language Agents for Mobile Device Control
Zhe Wu, Hongjin Lu, Junliang Xing +10
Building agents that autonomously operate mobile devices has attracted increasing attention. While Vision-Language Models (VLMs) show promise, most existing approaches rely on dire…