1 citations · 1 across the 1 of their papers we have counts for
1 paper
Yasumasa Onoe, Sunayana Rane, Zachary Berger +9
Vision-language datasets are vital for both text-to-image (T2I) and image-to-text (I2T) research. However, current datasets lack descriptions with fine-grained detail that would al…