3 papers
cs.CL2026
HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam
Weiqi Zhai, Zhihai Wang, Jinghang Wang +35
Humanity's Last Exam (HLE) has become a widely used benchmark for evaluating frontier large language models on challenging, multi-domain questions. However, community-led analyses…
cs.LG2026
DeepVision-103K: A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning
Haoxiang Sun, Lizhen Xu, Bing Zhao +5
Reinforcement Learning with Verifiable Rewards (RLVR) has been shown effective in enhancing the visual reflection and reasoning capabilities of Large Multimodal Models (LMMs). Howe…
cs.SE2026
AmbiBench: Benchmarking Mobile GUI Agents Beyond One-Shot Instructions in the Wild
Jiazheng Sun, Mingxuan Li, Yingying Zhang +11
Benchmarks are paramount for gauging progress in the domain of Mobile GUI Agents. In practical scenarios, users frequently fail to articulate precise directives containing full tas…