works on

From the 1 of 15 linked papers with an AI index.

collaborators

15 papers

cs.CR2026

HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses

Xiao Zhang, Yusheng Wang, Yuhao Fei +5

Modern agent harnesses persist state across tasks and sessions through persistent carriers like memory, skills, tools, and shared artifacts. However, this capability creates delaye…

cs.AI2026

Beyond Sequential Interaction: Benchmarking Parallel Execution and Coordination for GUI Agents

Zedong Yu, Qianxing Li, Zhi Gao +8

Graphical user interface (GUI) agents are systems powered by large multimodal models (LMMs). They perceive screen state and execute user instructions through GUI actions such as cl…

cs.AI2026

SAGA: Scene-Aware, Goal-Evolving Agents for Long-Horizon Strategy Game Planning

Tianyu Jin, Shuo Chen, Yida Wang +6

SAGA is a multi-agent framework that uses large language models to plan long‑term strategies in complex games by representing the game world as a scene graph, retrieving state on d…

cs.CL2026

Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis

Zicheng Kong, Dehua Ma, Zhenbo Xu +9

Multimodal large language models (MLLMs) struggle with alignment due to the limitations of existing reward models (RMs), which are predominantly vision-centric, dependent on costly…

cs.CV2026

AirGroundBench: Probing Spatial Intelligence in Multimodal Large Models under Heterogeneous Multi-View Embodied Collaboration

Haotian Li, Yida Wang, Leyuan Wang +7

In recent years, multimodal large language models (MLLMs) have shown strong potential for embodied intelligence, yet their ability to maintain geometrically consistent spatial unde…

cs.AI2026

GUI Knowledge Bench: Revealing the Knowledge Gap of VLMs in GUI Tasks

Chenrui Shi, Zedong Yu, Zhi Gao +7

Vision language models (VLMs) have advanced graphical user interface (GUI) task automation but still lag behind humans. We hypothesize this gap stems from missing core GUI knowledg…