From the 1 of 7 linked papers with an AI index.
7 papers
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
Mengqi Yuan, Zilong Zhou, Xinzhuang Xiong +33
The paper presents OSWorld 2.0, a benchmark consisting of 108 long‑horizon, real‑world computer‑use workflows designed to evaluate how well AI agents can handle complex, multi‑step…
Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use
Song-Lin Lv, Weiming Wu, Rui Zhu +2
While Large Language Model (LLM) agents demonstrate proficiency in static benchmarks, their deployment in real-world scenarios is hindered by the dynamic nature of user queries, to…
MapTab: A Diagnostic Benchmark for Long-Horizon Multi-Criteria Multimodal Reasoning on Heterogeneous Topological Graphs
Ziqiao Shang, Lingyue Ge, Ling-Yue Ge +12
Systematically evaluating Multimodal Large Language Models (MLLMs) is essential for advancing Artificial General Intelligence (AGI). Yet existing benchmarks remain inadequate for r…
A Progressive Visual-Logic-Aligned Framework for Ride-Hailing Adjudication
Weiming Wu, Zi-Jian Cheng, Jie Meng +5
The efficient adjudication of responsibility disputes is pivotal for maintaining marketplace fairness. However, the exponential surge in ride-hailing volume renders manual review i…
NeSyGeo: A Neuro-Symbolic Framework for Multimodal Geometric Reasoning Data Generation
Weiming Wu, Jin Ye, Zi-kang Wang +3
Obtaining large-scale, high-quality reasoning data is crucial for improving the geometric reasoning capabilities of multi-modal large language models (MLLMs). However, existing dat…
On Path to Multimodal Generalist: General-Level and General-Bench
Hao Fei, Yuan Zhou, Juncheng Li +29
The Multimodal Large Language Model (MLLM) is currently experiencing rapid growth, driven by the advanced capabilities of LLMs. Unlike earlier specialists, existing MLLMs are evolv…