3 papers
cs.CV2026
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
Kelaiti Xiao, Liang Yang, Dongyu Zhang +2
We introduce VisualQuest, a novel dataset designed to rigorously evaluate multimodal large language models (MLLMs) on abstract visual reasoning tasks that require the integration o…
cs.CL2025
Visual Puns from Idioms: An Iterative LLM-T2IM-MLLM Framework
Kelaiti Xiao, Liang Yang, Dongyu Zhang +2
We study idiom-based visual puns--images that align an idiom's literal and figurative meanings--and present an iterative framework that coordinates a large language model (LLM), a…
cs.CL2025
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
Junyu Lu, Kai Ma, Kaichun Wang +5
Large Language Models (LLMs) have become essential for offensive language detection, yet their ability to handle annotation disagreement remains underexplored. Disagreement samples…