benchmark 1environment state verification 1gui automation 1reinforcement learning 1task evaluation 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.AI2026
Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
Chenrui Shi, Yuwei Wu, Yang Liu +5
The paper introduces an Interactive Reward Agent that evaluates GUI task completion by proposing conditions and verifying them using system, application, and GUI tools, and demonst…
cs.AI2026
GUI Knowledge Bench: Revealing the Knowledge Gap of VLMs in GUI Tasks
Chenrui Shi, Zedong Yu, Zhi Gao +7
Vision language models (VLMs) have advanced graphical user interface (GUI) task automation but still lag behind humans. We hypothesize this gap stems from missing core GUI knowledg…
cs.CV2025
R^3-VQA: "Read the Room" by Video Social Reasoning
Lixing Niu, Jiapeng Li, Xingping Yu +6
"Read the room" is a significant social reasoning capability in human daily life. Humans can infer others' mental states from subtle social cues. Previous social reasoning tasks an…