3 citations · 4 across the 12 of their papers we have counts for
1 paper · 1 filter
Junha Song, Byeongho Heo, Geonmo Gu +3
When humans describe a visual scene, they do not process the entire image uniformly; instead, they selectively fixate on regions relevant to their intended description. In contrast…