activity
20242026
most citedOBI-Bench: Can LMMs Aid in Study of Ancient Script on Oracle Bones?

3 citations · 3 across the 23 of their papers we have counts for

collaborators

23 papers

cs.AI2026

SciMIF: Understanding Multimodal Instruction Following in Scientific Domains

Ye Shen, Yuting Zheng, Dun Pei +4

Understanding instruction-following capabilities in scientific domains is essential for effectively leveraging Multimodal Large Language Models (MLLMs) to advance the development o…

cs.RO2026

H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models

Dingyi Rong, Yue Shi, Chaofan Ma +6

Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human ma…

cs.CL2026

ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models

Yilin Jiang, Xiaorong Zhu, Fei Tan +9

Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does…

cs.LG2026

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research

Wanghan Xu, Shuo Li, Tianlin Ye +48

AI coding agents are increasingly used for scientific work, but their end-to-end autonomous research capability remains difficult to verify. We present ResearchClawBench, a benchma…

cs.CV2026

Beyond Single Character: Evaluating MLLMs for Sentence-Level Oracle Bone Inscription Understanding

Ziqi Li, Zijian Chen, Tingzhu Chen +1

Existing AI-assisted oracle bone inscription (OBI) visual recognition and understanding studies mainly focus on character-level, ignoring the long-form textual coherence and contex…

cs.CL2026

LatentRevise: Learning from Zero-Hit Reasoning

Yiqiu Guo, Xueting Han, Qi Jia +2

Reinforcement learning with verifiable rewards (RLVR) is bottlenecked by hard prompts on which correct trajectories have low probability, so sampling misses them within a practical…