37 citations · 83 across the 8 of their papers we have counts for
1 paper · 1 filter
Hung Le, Chinnadhurai Sankar, Seungwhan Moon +3
A video-grounded dialogue system is required to understand both dialogue, which contains semantic dependencies from turn to turn, and video, which contains visual cues of spatial a…