7 citations · 10 across the 8 of their papers we have counts for
3 papers · 1 filter
Grounded Video Caption Generation
Evangelos Kazakos, Cordelia Schmid, Josef Sivic
We propose a new task, dataset and model for grounded video caption generation. This task unifies captioning and object grounding in video, where the objects in the caption are gro…
TIM: A Time Interval Machine for Audio-Visual Action Recognition
Jacob Chalk, Jaesung Huh, Evangelos Kazakos +2
Diverse actions give rise to rich audio-visual signals in long videos. Recent works showcase that the two modalities of audio and video exhibit different temporal extents of events…
Graph Guided Question Answer Generation for Procedural Question-Answering
Hai X. Pham, Isma Hadji, Xinnuo Xu +6
In this paper, we focus on task-specific question answering (QA). To this end, we introduce a method for generating exhaustive and high-quality training data, which allows us to tr…