1 paper
Ziwei Xiang, Fanhu Zeng, Hongjian Fang +6
Large Vision Language Models (LVLMs) have achieved remarkable success in a range of downstream tasks that require multimodal interaction, but their capabilities come with substanti…