Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
MobileBench-OL: A Comprehensive Chinese Benchmark for Evaluating Mobile GUI Agents in Real-World Environment
Qinzhuo Wu, Zhizhuo Yang, Hanhao Li +3
Recent advances in mobile Graphical User Interface (GUI) agents highlight the growing need for comprehensive evaluation benchmarks. While new online benchmarks offer more realistic…
cs.CL2025
JoyAgent-JDGenie: Technical Report on the GAIA
Jiarun Liu, Shiyue Xu, Shangkun Liu +13
Large Language Models are increasingly deployed as autonomous agents for complex real-world tasks, yet existing systems often focus on isolated improvements without a unifying desi…