1 paper
Ziheng Wu, Zhenghao Chen, Ruipu Luo +6
Recently, vision-language models have made remarkable progress, demonstrating outstanding capabilities in various tasks such as image captioning and video understanding. We introdu…