23 citations · 24 across the 4 of their papers we have counts for
Showing 2023 · cs.CVShow all
2 papers · 2 filters
cs.CV2023
Uncovering Hidden Connections: Iterative Search and Reasoning for Video-grounded Dialog
Haoyu Zhang, Meng Liu, Yisen Feng +3
In contrast to conventional visual question answering, video-grounded dialog necessitates a profound understanding of both dialog history and video content for accurate response ge…
cs.CV2023★ 23 cited
Learnable Pillar-based Re-ranking for Image-Text Retrieval
Leigang Qu, Meng Liu, Wenjie Wang +3
Image-text retrieval aims to bridge the modality gap and retrieve cross-modal content based on semantic similarities. Prior work usually focuses on the pairwise relations (i.e., wh…