4 papers
A2Eval: Agentic and Automated Evaluation for Embodied Brain
Shuai Zhang, Jiayu Hu, Zijie Chen +9
Current embodied VLM evaluation relies on static, expert-defined, manually annotated benchmarks that exhibit severe redundancy and coverage imbalance. This labor intensive paradigm…
ProBench: Benchmarking GUI Agents with Accurate Process Information
Leyang Yang, Ziwei Wang, Xiaoxuan Tang +4
With the deep integration of artificial intelligence and interactive technology, Graphical User Interface (GUI) Agent, as the carrier connecting goal-oriented natural language and…
History-Aware Reasoning for GUI Agents
Ziwei Wang, Leyang Yang, Xiaoxuan Tang +4
Advances in Multimodal Large Language Models have significantly enhanced Graphical User Interface (GUI) automation. Equipping GUI agents with reliable episodic reasoning capabiliti…
PG-Agent: An Agent Powered by Page Graph
Weizhi Chen, Ziwei Wang, Leyang Yang +5
Graphical User Interface (GUI) agents possess significant commercial and social value, and GUI agents powered by advanced multimodal large language models (MLLMs) have demonstrated…