8 papers
RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
Yang Shi, Yuhao Dong, Yue Ding +22
The integration of visual understanding and generation into unified multimodal models represents a significant stride toward general-purpose AI. However, a fundamental question rem…
CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
Songqin Nong, Xiaoxuan Tang, Jingxuan Xu +4
As autonomous agents become adept at understanding and interacting with graphical user interface (GUI) environments, a new era of automated task execution is emerging. Recent studi…
Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents
Haiyang Xu, Xi Zhang, Haowei Liu +16
The paper introduces GUI-Owl-1.5, the latest native GUI agent model that features instruct/thinking variants in multiple sizes (2B/4B/8B/32B/235B) and supports a range of platforms…
SVRepair: Structured Visual Reasoning for Automated Program Repair
Xiaoxuan Tang, Jincheng Wang, Liwei Luo +5
Large language models (LLMs) have recently been applied to Automated Program Repair (APR), yet most existing approaches remain unimodal and fail to use diagnostic signals contained…
A2Eval: Agentic and Automated Evaluation for Embodied Brain
Shuai Zhang, Jiayu Hu, Zijie Chen +9
Current embodied VLM evaluation relies on static, expert-defined, manually annotated benchmarks that exhibit severe redundancy and coverage imbalance. This labor intensive paradigm…
ProBench: Benchmarking GUI Agents with Accurate Process Information
Leyang Yang, Ziwei Wang, Xiaoxuan Tang +4
With the deep integration of artificial intelligence and interactive technology, Graphical User Interface (GUI) Agent, as the carrier connecting goal-oriented natural language and…