3 citations · 3 across the 2 of their papers we have counts for
3 papers · 1 filter
OSExpert: Computer-Use Agents Learning Professional Skills via Exploration
Jiateng Liu, Zhenhailong Wang, Rushi Wang +6
General-purpose computer-use agents have shown impressive performance across diverse digital environments. However, our new benchmark, OSExpert-Eval, indicates they remain far less…
Current Agents Fail to Leverage World Model as Tool for Foresight
Cheng Qian, Emre Can Acikgoz, Bingxuan Li +8
Agents built on vision-language models increasingly face tasks that demand anticipating future states rather than relying on short-horizon reasoning. Generative world models offer…
Where LLM Agents Fail and How They can Learn From Failures
Kunlun Zhu, Zijia Liu, Bingxuan Li +15
Large Language Model (LLM) agents, which integrate planning, memory, reflection, and tool-use modules, have shown promise in solving complex, multi-step tasks. Yet their sophistica…