2 papers
cs.SE2026
AmbiBench: Benchmarking Mobile GUI Agents Beyond One-Shot Instructions in the Wild
Jiazheng Sun, Mingxuan Li, Yingying Zhang +11
Benchmarks are paramount for gauging progress in the domain of Mobile GUI Agents. In practical scenarios, users frequently fail to articulate precise directives containing full tas…
cs.SE2025
LogiAgent: Automated Logical Testing for REST Systems with LLM-Based Multi-Agents
Ke Zhang, Chenxi Zhang, Chong Wang +6
Automated testing for REST APIs has become essential for ensuring the correctness and reliability of modern web services. While existing approaches primarily focus on detecting ser…