works on

From the 1 of 12 linked papers with an AI index.

collaborators
Showing cs.AIShow all

7 papers · 1 filter

cs.AI2026

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

Mengqi Yuan, Zilong Zhou, Xinzhuang Xiong +33

The paper presents OSWorld 2.0, a benchmark consisting of 108 long‑horizon, real‑world computer‑use workflows designed to evaluate how well AI agents can handle complex, multi‑step…

cs.AI2026

Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments

Parth Asawa, Christopher M. Glaze, Gabriel Orlanski +7

Continual learning, the ability of AI systems to improve through sequential experience, has attracted substantial interest, but no high-quality benchmark exists to evaluate it. We…

cs.AI2026

Can Generalist Agents Automate Data Curation?

Feiyang Kang, Hanze Li, Adam Nguyen +5

Curating training data is among the most consequential yet labor-intensive parts of modern AI development: practitioners iteratively propose, implement, evaluate, and revise data p…

cs.AI2026

RIFT: A RubrIc Failure Mode Taxonomy and Automated Diagnostics

Zhengyang Qi, Charles Dickens, Derek Pham +4

Rubric-based evaluation is widely used in LLM benchmarks and training pipelines for open-ended, less verifiable tasks. While prior work has demonstrated the effectiveness of rubric…

cs.AI2026

LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild

Jiayu Wang, Yifei Ming, Riya Dulepet +7

Deep research -- producing comprehensive, citation-grounded reports by searching and synthesizing information from hundreds of live web sources -- marks an important frontier for a…

cs.AI2026

SkillOrchestra: Learning to Route Agents via Skill Transfer

Jiayu Wang, Yifei Ming, Zixuan Ke +3

Compound AI systems promise capabilities beyond those of individual models, yet their success depends critically on effective orchestration. Existing routing approaches face two li…