200 citations · 233 across the 6 of their papers we have counts for
1 paper · 1 filter
Hung Le, Chinnadhurai Sankar, Seungwhan Moon +3
A video-grounded dialogue system is required to understand both dialogue, which contains semantic dependencies from turn to turn, and video, which contains visual cues of spatial a…