207 citations · 212 across the 6 of their papers we have counts for
1 paper · 1 filter
Pranav Aggarwal, Zhe Lin, Baldo Faieta +1
Text-visual (or called semantic-visual) embedding is a central problem in vision-language research. It typically involves mapping of an image and a text description to a common fea…