From the 1 of 4 linked papers with an AI index.
4 papers
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
Jingbo Zhou, Yusai Zhao, Qi Bao +12
The paper presents OmegaUse-OfficeVal, a benchmark that evaluates large language model agents on long‑horizon office‑suite tasks while providing economic signals (human labor time…
From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs
Le Zhang, Jihan Yang, Soundarya Krishnan +11
Human-level agentic intelligence extends beyond low-level geometric perception, evolving from recognizing where things are to understanding what they are for. While existing benchm…
Unrewarded Exploration in Large Language Models Reveals Latent Learning from Psychology
Jian Xiong, Jingbo Zhou, Zihan Zhou +6
Latent learning, classically theorized by Tolman, shows that biological agents (e.g., rats) can acquire internal representations of their environment without rewards, enabling rapi…
OmegaUse: Building a General-Purpose GUI Agent for Autonomous Task Execution
Le Zhang, Yixiong Xiao, Xinjiang Lu +12
Graphical User Interface (GUI) agents show great potential for enabling foundation models to complete real-world tasks, revolutionizing human-computer interaction and improving hum…