paper

What Color Is the Text? A Benchmark for Hallucination Induced by Image-Embedded Prompt

arXiv:2511.13400

Abstract

We introduce Embedded Stroop, a controlled diagnostic paradigm for measuring image-embedded prompt interference in Multimodal Large Language Models (MLLMs), where the query is rendered directly inside the visual input. Using the What-Color-Is-the-Text (WCIT) benchmark, which covers 59 fine-grained colors under Standard, Flipped, and Masked variants, we evaluate 16 proprietary and open-source models. To distinguish semantic capture from general color-naming failure, we decompose model responses into Accuracy, Stroop Hallucination Rate (SHR; answering the embedded word rather than the true text color), and Other Error Rate, and validate the effect with permutation tests, with 56 of 64 conditions remaining significant after FDR correction. Although exact color accuracy is low (6.3%) under the 59-color vocabulary, mapping predictions to 11 basic color families shows that models retain coarse color perception (38.4%) while still exhibiting a substantial SHR (21.6%). A conditional analysis restricted to colors correctly named in the Standard setting further confirms the effect, with pooled conditional SHR exceeding 50\%. Masking or flipping the embedded text reduces Stroop hallucinations, suggesting that semantic legibility can dominate visual color perception in MLLMs.

What Color Is the Text? A Benchmark for Hallucination Induced by Image-Embedded Prompt · wovepaper