Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
MDK12-Bench: A Comprehensive Evaluation of Multimodal Large Language Models on Multidisciplinary Exams
Pengfei Zhou, Xiaopeng Peng, Fanrui Zhang +18
Multimodal large language models (MLLMs), which integrate language and visual cues for problem-solving, are crucial for advancing artificial general intelligence (AGI). However, cu…
cs.AI2025
MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification
Zhaopan Xu, Pengfei Zhou, Jiaxin Ai +6
Reasoning is an essential capacity for large language models (LLMs) to address complex tasks, where the identification of process errors is vital for improving this ability. Recent…