Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests
Gyuhyeong Kim, Hyojung Gwon, Jeonghyeon Kim +2
Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated GitHub issues: long, structured, and information-rich. Real user…
cs.AI2025
MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents
Youngmin Im, Byeongung Jo, Jaeyoung Wi +6
Mobile GUI Agents, AI agents capable of interacting with mobile applications on behalf of users, have the potential to transform human computer interaction. However, current evalua…