14 papers · 1 filter
The Illusion of : Evaluating the Breakdown of Counterfactual Reasoning in LLMs
Yucheng Wang, Yuetian Du, Zhengyi Liu +8
Counterfactual reasoning requires models to reason beyond the observed world and explain how altered conditions propagate through downstream consequences. Existing benchmarks large…
OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses
Guangzheng Hu, Ziyue Jiang, Weixu Qiao +14
Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and text-to-speech (TTS) generation are increasingly used as "OmniJudges" for evalua…
Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks
Xuan Ren, Weiqi Zhai, Tianle Pu +3
Scientific reasoning benchmarks typically evaluate large language models (LLMs) using final-answer accuracy. However, a correct answer does not necessarily demonstrate the reasonin…
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
Zeyu Wang, Jingye Xu, Xiaogang Li +7
Current multimodal benchmarks for scientific reasoning primarily evaluate local information extraction -- models recognize symbols and values and then perform textual inference. Th…
BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs
Ben Wang, Xiaogang Li, Ruochen Gao +6
Current multimodal models handle static image recognition well, but intuitive physical reasoning remains a weakness. Predicting how objects will move and interact from a single ima…
Benchmark Health Index: A Systematic Framework for Benchmarking the Benchmarks of LLMs
Longyuan Zhu, Hairan Hua, Linlin Miao +1
Large Language Models (LLMs) are advancing rapidly, yet the benchmarks used to measure this progress are becoming increasingly unreliable. Score inflation and selective reporting h…