1 paper
Jinyi Hu, Yuan Yao, Chongyi Wang +13
Recently there has been a significant surge in multimodal learning in terms of both image-to-text and text-to-image generation. However, the success is typically limited to English…