collaborators

8 papers

cs.AI2026

RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark

Yang Shi, Yuhao Dong, Yue Ding +22

The integration of visual understanding and generation into unified multimodal models represents a significant stride toward general-purpose AI. However, a fundamental question rem…

cs.AI2026

CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks

Songqin Nong, Xiaoxuan Tang, Jingxuan Xu +4

As autonomous agents become adept at understanding and interacting with graphical user interface (GUI) environments, a new era of automated task execution is emerging. Recent studi…

cs.AI2026

Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents

Haiyang Xu, Xi Zhang, Haowei Liu +16

The paper introduces GUI-Owl-1.5, the latest native GUI agent model that features instruct/thinking variants in multiple sizes (2B/4B/8B/32B/235B) and supports a range of platforms…

cs.SE2026

SVRepair: Structured Visual Reasoning for Automated Program Repair

Xiaoxuan Tang, Jincheng Wang, Liwei Luo +5

Large language models (LLMs) have recently been applied to Automated Program Repair (APR), yet most existing approaches remain unimodal and fail to use diagnostic signals contained…

cs.CL2026

A2Eval: Agentic and Automated Evaluation for Embodied Brain

Shuai Zhang, Jiayu Hu, Zijie Chen +9

Current embodied VLM evaluation relies on static, expert-defined, manually annotated benchmarks that exhibit severe redundancy and coverage imbalance. This labor intensive paradigm…

cs.AI2025

ProBench: Benchmarking GUI Agents with Accurate Process Information

Leyang Yang, Ziwei Wang, Xiaoxuan Tang +4

With the deep integration of artificial intelligence and interactive technology, Graphical User Interface (GUI) Agent, as the carrier connecting goal-oriented natural language and…