66 citations · 84 across the 13 of their papers we have counts for
1 paper · 2 filters
Andrew Rouditchenko, Angie Boggust, David Harwath +11
Current methods for learning visually grounded language from videos often rely on text annotation, such as human generated captions or machine generated automatic speech recognitio…