From the 1 of 12 linked papers with an AI index.
12 papers
Beyond Sequential Interaction: Benchmarking Parallel Execution and Coordination for GUI Agents
Zedong Yu, Qianxing Li, Zhi Gao +8
Graphical user interface (GUI) agents are systems powered by large multimodal models (LMMs). They perceive screen state and execute user instructions through GUI actions such as cl…
SAGA: Scene-Aware, Goal-Evolving Agents for Long-Horizon Strategy Game Planning
Tianyu Jin, Shuo Chen, Yida Wang +6
SAGA is a multi-agent framework that uses large language models to plan long‑term strategies in complex games by representing the game world as a scene graph, retrieving state on d…
Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis
Zicheng Kong, Dehua Ma, Zhenbo Xu +9
Multimodal large language models (MLLMs) struggle with alignment due to the limitations of existing reward models (RMs), which are predominantly vision-centric, dependent on costly…
AirGroundBench: Probing Spatial Intelligence in Multimodal Large Models under Heterogeneous Multi-View Embodied Collaboration
Haotian Li, Yida Wang, Leyuan Wang +7
In recent years, multimodal large language models (MLLMs) have shown strong potential for embodied intelligence, yet their ability to maintain geometrically consistent spatial unde…
GUI Knowledge Bench: Revealing the Knowledge Gap of VLMs in GUI Tasks
Chenrui Shi, Zedong Yu, Zhi Gao +7
Vision language models (VLMs) have advanced graphical user interface (GUI) task automation but still lag behind humans. We hypothesize this gap stems from missing core GUI knowledg…
Simulation-Free PSRO: Removing Game Simulation from Policy Space Response Oracles
Yingzhuo Liu, Shuodi Liu, Weijun Luo +2
Policy Space Response Oracles (PSRO) combines game-theoretic equilibrium computation with learning and is effective in approximating Nash Equilibrium in zero-sum games. However, th…