3 papers
cs.MA2025
ColorAgent: Building A Robust, Personalized, and Interactive OS Agent
Ning Li, Qiqiang Lin, Zheng Wu +19
With the advancements in hardware, software, and large language model technologies, the interaction between humans and operating systems has evolved from the command-line interface…
cs.AI2025
ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
Yuanyi Song, Heyuan Huang, Qiqiang Lin +9
The rapid advancement of multimodal large language models has enabled agents to operate mobile devices by directly interacting with graphical user interfaces, opening new possibili…
cs.CL2024
HammerBench: Fine-Grained Function-Calling Evaluation in Real Mobile Device Scenarios
Jun Wang, Jiamu Zhou, Muning Wen +7
Evaluating the performance of LLMs in multi-turn human-agent interactions presents significant challenges, particularly due to the complexity and variability of user behavior. In t…