2 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.AI2026
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
Mengqi Yuan, Zilong Zhou, Xinzhuang Xiong +33
Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to reveal the limitations of f…
cs.CV2026
3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code
Yipeng Gao, Lei Shu, Genzhi Ye +5
Procedural 3D modeling through code is emerging as a versatile paradigm, offering deterministic, engine-ready, and precisely editable assets that neural 3D generators inherently la…
cs.AI2024★ 2 cited
Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?
Ruisheng Cao, Fangyu Lei, Haoyuan Wu +20
Data science and engineering workflows often span multiple stages, from warehousing to orchestration, using tools like BigQuery, dbt, and Airbyte. As vision language models (VLMs)…