3 papers
cs.HC2026
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
Felix Henry, Xiaochen Lin, Jiangyou Zhu +4
Current benchmarks for graphical user interface (GUI) agents predominantly rely on static screenshots. However, real-world smartphone interaction routinely requires agents to proce…
cs.AI2025
OmniPlay: Benchmarking Omni-Modal Models on Omni-Modal Game Playing
Fuqing Bie, Shiyu Huang, Xijia Tao +6
While generalist foundation models like Gemini and GPT-4o demonstrate impressive multi-modal competence, existing evaluations fail to test their intelligence in dynamic, interactiv…
cs.MA2025
Generalizable Agent Modeling for Agent Collaboration-Competition Adaptation with Multi-Retrieval and Dynamic Generation
Chenxu Wang, Yonggang Jin, Cheng Hu +7
Adapting a single agent to a new multi-agent system brings challenges, necessitating adjustments across various tasks, environments, and interactions with unknown teammates and opp…