3 papers
cs.AI2025
TIM-PRM: Verifying multimodal reasoning with Tool-Integrated PRM
Peng Kuang, Xiangxiang Wang, Wentao Liu +2
Multimodal Large Language Models (MLLMs) have achieved impressive performances in mathematical reasoning, yet they remain vulnerable to visual hallucinations and logical inconsiste…
cs.MA2025
S-DAG: A Subject-Based Directed Acyclic Graph for Multi-Agent Heterogeneous Reasoning
Jiangwen Dong, Zehui Lin, Wanyu Lin +1
Large Language Models (LLMs) have achieved impressive performance in complex reasoning problems. Their effectiveness highly depends on the specific nature of the task, especially t…
cs.CL2025
Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory
Hongli Zhou, Hui Huang, Ziqing Zhao +10
The evaluation of large language models (LLMs) via benchmarks is widespread, yet inconsistencies between different leaderboards and poor separability among top models raise concern…