works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.AI2026

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

Mengqi Yuan, Zilong Zhou, Xinzhuang Xiong +33

The paper presents OSWorld 2.0, a benchmark consisting of 108 long‑horizon, real‑world computer‑use workflows designed to evaluate how well AI agents can handle complex, multi‑step…

cs.AI2026

Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use

Song-Lin Lv, Weiming Wu, Rui Zhu +2

While Large Language Model (LLM) agents demonstrate proficiency in static benchmarks, their deployment in real-world scenarios is hindered by the dynamic nature of user queries, to…

cs.LG2026

MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs

Ziqiao Shang, Lingyue Ge, Ling-Yue Ge +12

Systematically evaluating Multimodal Large Language Models (MLLMs) is essential for advancing Artificial General Intelligence (AGI). Yet existing benchmarks remain inadequate for r…

cs.AI2026

A Progressive Visual-Logic-Aligned Framework for Ride-Hailing Adjudication

Weiming Wu, Zi-Jian Cheng, Jie Meng +5

The efficient adjudication of responsibility disputes is pivotal for maintaining marketplace fairness. However, the exponential surge in ride-hailing volume renders manual review i…

cs.CL2025

NeSyGeo: A Neuro-Symbolic Framework for Multimodal Geometric Reasoning Data Generation

Weiming Wu, Jin Ye, Zi-kang Wang +3

Obtaining large-scale, high-quality reasoning data is crucial for improving the geometric reasoning capabilities of multi-modal large language models (MLLMs). However, existing dat…

cs.CV2025

On Path to Multimodal Generalist: General-Level and General-Bench

Hao Fei, Yuan Zhou, Juncheng Li +29

The Multimodal Large Language Model (MLLM) is currently experiencing rapid growth, driven by the advanced capabilities of LLMs. Unlike earlier specialists, existing MLLMs are evolv…