From the 1 of 5 linked papers with an AI index.
1 citations · 1 across the 2 of their papers we have counts for
5 papers
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
Mengqi Yuan, Zilong Zhou, Xinzhuang Xiong +33
The paper presents OSWorld 2.0, a benchmark consisting of 108 long‑horizon, real‑world computer‑use workflows designed to evaluate how well AI agents can handle complex, multi‑step…
JumpStarter: Human-AI Planning with Task-Structured Context Curation
Xuanming Zhang, Sitong Wang, Jenny Ma +3
Human-AI collaboration on complex planning goals is bottlenecked by how LLM interfaces handle context: users must manually curate and re-surface relevant information across long an…
Budget-Aware Anytime Reasoning with LLM-Synthesized Preference Data
Xuanming Zhang, Shwan Ashrafi, Aziza Mirsaidova +5
We study the reasoning behavior of large language models (LLMs) under limited computation budgets. In such settings, producing useful partial solutions quickly is often more practi…
DoubleAgents: Human-Agent Alignment in a Socially Embedded Workflow
Tao Long, Xuanming Zhang, Sitong Wang +2
Aligning agentic AI with user intent is critical for delegating complex, socially embedded tasks, yet user preferences are often implicit, evolving, and difficult to specify upfron…
Bringing Pedagogy into Focus: Evaluating Virtual Teaching Assistants' Question-Answering in Asynchronous Learning Environments
Li Siyan, Zhen Xu, Vethavikashini Chithrra Raghuram +3
Asynchronous learning environments (ALEs) are widely adopted for formal and informal learning, but timely and personalized support is often limited. In this context, Virtual Teachi…