2 papers
cs.CV2025
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark
Ziming Cheng, Binrui Xu, Lisheng Gong +14
With enhanced capabilities and widespread applications, Multimodal Large Language Models (MLLMs) are increasingly required to process and reason over multiple images simultaneously…
cs.CL2025
MIST: Towards Multi-dimensional Implicit BiaS Evaluation of LLMs for Theory of Mind
Yanlin Li, Hao Liu, Huimin Liu +3
Theory of Mind (ToM) in Large Language Models (LLMs) refers to the model's ability to infer the mental states of others, with failures in this ability often manifesting as systemic…