1 paper
Qianlong Yang, Bowen Ye, Xianda Guo +4
Despite the progress of multimodal large language models (MLLMs), they continue to exhibit deficiencies in visual perception. Following visual instruction tuning, internal MLLM rep…