7 citations · 7 across the 1 of their papers we have counts for
1 paper · 1 filter
Nan Ding, Sebastian Goodman, Fei Sha +1
We introduce a new multi-modal task for computer systems, posed as a combined vision-language comprehension challenge: identifying the most suitable text describing a scene, given…