anticipatory planning 1benchmark evaluation 1computer-use agents 1error analysis 1gui automation 1latency reduction 1multimodal models 1policy trees 1reliability 1task scoring 1
From the 2 of 2 linked papers with an AI index.
2 papers
cs.LG2026
Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees
Zihan Dong, Rui Qian, Qishi Zhan +3
The paper introduces Adaptive Anticipatory Policy Trees (AAPT), a method that pre‑computes conditional action trees during idle screen time so GUI agents can react instantly to eve…
cs.AI2026
How Benchmarks Mis-Score Computer-Use Agents
Zihan Dong, Zhiyuan Ma, Zekun Wang +5
The paper examines how current benchmarks for computer-use agents often give inaccurate scores due to issues in task design, trajectory observation, scoring, and reporting, and pro…