2 papers
cs.AI2026
Trust the Right Teacher: Quality-Aware Self-Distillation for GUI Grounding
Jingyuan Huang, Zuming Huang, Yucheng Shi +4
Graphical user interface (GUI) grounding requires vision-language models (VLMs) to identify small target elements in high-resolution screenshots and predict precise screen coordina…
cs.AI2026
Online Skill Learning for Web Agents via State-Grounded Dynamic Retrieval
Jiaxi Li, Ke Deng, Yun Wang +5
Language agents increasingly rely on reusable skills to improve multi-step web automation across related tasks. A growing line of work studies online skill learning, where agents c…