1 paper · 1 filter
Yuchi Wang, Shuhuai Ren, Rundong Gao +5
Diffusion models have exhibited remarkable capabilities in text-to-image generation. However, their performance in image-to-text generation, specifically image captioning, has lagg…