1 citations · 1 across the 11 of their papers we have counts for
1 paper · 2 filters
Leanne Nortje, Dan Oneata, Herman Kamper
We propose a visually grounded speech model that learns new words and their visual depictions from just a few word-image example pairs. Given a set of test images and a spoken quer…