2 citations · 2 across the 1 of their papers we have counts for
1 paper
Tao Huang, Zhekun Liu, Rui Wang +2
Despite the remarkable multimodal capabilities of Large Vision-Language Models (LVLMs), discrepancies often occur between visual inputs and textual outputs--a phenomenon we term vi…