6 papers
Visual Room 2.0: Seeing is Not Understanding for MLLMs
Haokun Li, Yazhou Zhang, Jizhi Ding +2
Can multi-modal large language models (MLLMs) truly understand what they can see? Extending Searle's Chinese Room into the multi-modal domain, this paper proposes the Visual Room a…
Seeing is Not Understanding: A Benchmark on Perception-Cognition Disparities in Large Language Models
Haokun Li, Yazhou Zhang, Jizhi Ding +2
With the rapid advancement of Multimodal Large Language Models (MLLMs), they have demonstrated exceptional capabilities across a variety of vision-language tasks. However, current…
Large Language Models for Subjective Language Understanding: A Survey
Changhao Song, Yazhou Zhang, Hui Gao +2
Subjective language understanding refers to a broad set of natural language processing tasks where the goal is to interpret or generate content that conveys personal feelings, opin…
Are MLMs Trapped in the Visual Room?
Yazhou Zhang, Chunwang Zou, Qimeng Liu +6
Can multi-modal large models (MLMs) that can ``see'' an image be said to ``understand'' it? Drawing inspiration from Searle's Chinese Room, we propose the \textbf{Visual Room} argu…
Emotion-o1: Adaptive Long Reasoning for Emotion Understanding in LLMs
Changhao Song, Yazhou Zhang, Hui Gao +2
Long chain-of-thought (CoT) reasoning has shown great promise in enhancing the emotion understanding performance of large language models (LLMs). However, current fixed-length CoT…
Beyond Single-Sentence Prompts: Upgrading Value Alignment Benchmarks with Dialogues and Stories
Yazhou Zhang, Qimeng Liu, Qiuchi Li +2
Evaluating the value alignment of large language models (LLMs) has traditionally relied on single-sentence adversarial prompts, which directly probe models with ethically sensitive…