1 paper · 1 filter
Yazhou Zhang, Chunwang Zou, Qimeng Liu +6
Can multi-modal large models (MLMs) that can ``see'' an image be said to ``understand'' it? Drawing inspiration from Searle's Chinese Room, we propose the \textbf{Visual Room} argu…