1 paper · 1 filter
Jingbo Zhou, Yusai Zhao, Qi Bao +12
The paper presents OmegaUse-OfficeVal, a benchmark that evaluates large language model agents on long‑horizon office‑suite tasks while providing economic signals (human labor time…