1 paper · 1 filter
David Harwath, Adrià Recasens, Dídac Surís +3
In this paper, we explore neural network models that learn to associate segments of spoken audio captions with the semantically relevant portions of natural images that they refer…