1 paper
Bo Zhao, Boya Wu, Muyang He +1
Thanks to the emerging of foundation models, the large language and vision models are integrated to acquire the multimodal ability of visual captioning, question answering, etc. Al…