15 papers
StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents
Haojie Hao, Longkun Hao, Yihang Lou +8
Reinforcement Learning (RL) has become a promising approach for improving GUI Agents in long-horizon, stochastic digital environments, but trajectory-level success feedback is too…
Speculative Rollback Correction for Quality-Diverse Web Agent Imitation
Longkun Hao, Hongyu Lin, Hao Li +11
Training interactive web agents through imitation learning from expert trajectories has emerged as a highly effective approach. However, determining the optimal timing for expert i…
MIRAGE: Mobile Agents with Implicit Reasoning and Generative World Models
Zhichao Yang, Yuanze Hu, Haojie Hao +7
Mobile agents are increasingly expected to operate everyday applications from screenshots and language goals, where reliable control requires reasoning over screen affordances, mul…
Mechanistic Insights into Functional Sparsity in Multimodal LLMs via CoRe Heads
Ruoxi Sun, Quantong Qiu, Juntao Li +3
While Multimodal Large Language Models (MLLMs) demonstrate remarkable proficiency on complex vision-language tasks, the mechanisms by which they extract query-relevant visual featu…
ToolFG: Towards Well-Grounded Fine-Grained Image Classification
Yu Xue, Haoxuan Qu, Zhuoling Li +4
Fine-grained image classification (FGIC) has broad applications and has attracted significant research attention. In this paper, we explore a novel paradigm for solving FGIC by pro…
HTMLCure: Turning Browser Experience into State Guided Repair for Interactive HTML
Jiajun Wu, Jian Yang, Tuney Zheng +4
LLMs can now produce full HTML pages, but many of those pages are only superficially correct: they render once, then fail under scroll, hover, click, resize, or gameplay. Evaluatio…