24 papers
H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
Dingyi Rong, Yue Shi, Chaofan Ma +6
Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human ma…
ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models
Yilin Jiang, Xiaorong Zhu, Fei Tan +9
Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does…
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research
Wanghan Xu, Shuo Li, Tianlin Ye +48
AI coding agents are increasingly used for scientific work, but their end-to-end autonomous research capability remains difficult to verify. We present ResearchClawBench, a benchma…
Beyond Single Character: Evaluating MLLMs for Sentence-Level Oracle Bone Inscription Understanding
Ziqi Li, Zijian Chen, Tingzhu Chen +1
Existing AI-assisted oracle bone inscription (OBI) visual recognition and understanding studies mainly focus on character-level, ignoring the long-form textual coherence and contex…
LatentRevise: Learning from Zero-Hit Reasoning
Yiqiu Guo, Xueting Han, Qi Jia +2
Reinforcement learning with verifiable rewards (RLVR) is bottlenecked by hard prompts on which correct trajectories have low probability, so sampling misses them within a practical…
RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation
Dayu Xia, Yue Shi, Yao Mu +7
Vision-language models (VLMs) are increasingly explored as visual critics, reward generators, and failure detectors in robotic manipulation. These roles implicitly require models t…