2 papers
cs.CV2025
VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning
Jingkun Ma, Runzhe Zhan, Yang Li +4
A hallmark of advanced artificial intelligence is the capacity to progress from passive visual perception to the strategic modification of visual information to facilitate complex…
cs.CL2025
Do Large Language Models Judge Error Severity Like Humans?
Diege Sun, Guanyi Chen, Zhao Fan +2
Large Language Models (LLMs) are increasingly used as automated evaluators in natural language generation, yet it remains unclear whether they can accurately replicate human judgme…