1 paper · 1 filter
Peng Xie, Yequan Bie, Jianda Mao +4
Pre-trained vision-language models (VLMs) have showcased remarkable performance in image and natural language understanding, such as image captioning and response generation. As th…