5 papers
OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
Felix Henry, Xiaochen Lin, Jiangyou Zhu +4
Current benchmarks for graphical user interface (GUI) agents predominantly rely on static screenshots. However, real-world smartphone interaction routinely requires agents to proce…
Table-as-Search: Formulate Long-Horizon Agentic Information Seeking as Table Completion
Tian Lan, Felix Henry, Bin Zhu +7
Current Information Seeking (InfoSeeking) agents struggle to maintain focus and coherence during long-horizon exploration, as tracking search states, including planning procedure a…
Can Large Language Models Master Complex Card Games?
Wei Wang, Fuqing Bie, Junzhe Chen +4
Complex games have long been an important benchmark for testing the progress of artificial intelligence algorithms. AlphaGo, AlphaZero, and MuZero have defeated top human players i…
Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
Leyi Pan, Zheyu Fu, Yunpeng Zhai +9
Omni-modal Large Language Models (OLLMs) that integrate visual, auditory, and textual processing face severe safety risks. They exhibit fragile defenses against audio-visual joint…
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
Chonghan Liu, Haoran Wang, Felix Henry +4
Spatial perception and reasoning are core components of human cognition, encompassing object recognition, spatial relational understanding, and dynamic reasoning. Despite progress…