1 paper
Gabriel Ilharco, Yuan Zhang, Jason Baldridge
Systems that can associate images with their spoken audio captions are an important step towards visually grounded language learning. We describe a scalable method to automatically…