1 paper
Dongwan Kim, Viresh Ranjan, Takashi Nagata +2
Despite the remarkable success of the LLaVA architecture for vision-language tasks, its design inherently struggles to effectively integrate visual features due to the inherent mis…