From the 1 of 15 linked papers with an AI index.
15 papers
HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses
Xiao Zhang, Yusheng Wang, Yuhao Fei +5
Modern agent harnesses persist state across tasks and sessions through persistent carriers like memory, skills, tools, and shared artifacts. However, this capability creates delaye…
Beyond Sequential Interaction: Benchmarking Parallel Execution and Coordination for GUI Agents
Zedong Yu, Qianxing Li, Zhi Gao +8
Graphical user interface (GUI) agents are systems powered by large multimodal models (LMMs). They perceive screen state and execute user instructions through GUI actions such as cl…
SAGA: Scene-Aware, Goal-Evolving Agents for Long-Horizon Strategy Game Planning
Tianyu Jin, Shuo Chen, Yida Wang +6
SAGA is a multi-agent framework that uses large language models to plan long‑term strategies in complex games by representing the game world as a scene graph, retrieving state on d…
Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis
Zicheng Kong, Dehua Ma, Zhenbo Xu +9
Multimodal large language models (MLLMs) struggle with alignment due to the limitations of existing reward models (RMs), which are predominantly vision-centric, dependent on costly…
AirGroundBench: Probing Spatial Intelligence in Multimodal Large Models under Heterogeneous Multi-View Embodied Collaboration
Haotian Li, Yida Wang, Leyuan Wang +7
In recent years, multimodal large language models (MLLMs) have shown strong potential for embodied intelligence, yet their ability to maintain geometrically consistent spatial unde…
GUI Knowledge Bench: Revealing the Knowledge Gap of VLMs in GUI Tasks
Chenrui Shi, Zedong Yu, Zhi Gao +7
Vision language models (VLMs) have advanced graphical user interface (GUI) task automation but still lag behind humans. We hypothesize this gap stems from missing core GUI knowledg…