2 papers
cs.AI2026
Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation
Xue Liu, Xin Ma, Yuxin Ma +36
As Large Language Models (LLMs) exhibit plateauing performance on conventional benchmarks, a pivotal challenge persists: evaluating their proficiency in complex, open-ended tasks c…
cs.RO2026
Leveraging VR Robot Games to Facilitate Data Collection for Embodied Intelligence Tasks
Yihan Zhang, Ziyun Huang, Linqi Ye
Collecting embodied interaction data at scale remains costly and difficult due to the limited accessibility of conventional interfaces. We present a gamified data collection framew…