5 citations · 5 across the 1 of their papers we have counts for
1 paper
Junhyeong Cho, Gilhyun Nam, Sungyeon Kim +2
In a joint vision-language space, a text feature (e.g., from "a photo of a dog") could effectively represent its relevant image features (e.g., from dog photos). Also, a recent stu…