1 paper · 1 filter
Shixin Yi, Lin Shang
Multimodal reasoning with vision-language models (VLMs) often suffers from hallucinations, as models tend to generate explanations after only a superficial inspection of the image.…