5 citations · 6 across the 4 of their papers we have counts for
1 paper · 1 filter
Hongcheng Liu, Zhe Chen, Hui Li +3
Generating dialogue grounded in videos requires a high level of understanding and reasoning about the visual scenes in the videos. However, existing large visual-language models ar…