4 citations · 4 across the 3 of their papers we have counts for
3 papers
cs.CV2023
Text-Video Retrieval with Disentangled Conceptualization and Set-to-Set Alignment
Peng Jin, Hao Li, Zesen Cheng +5
Text-video retrieval is a challenging cross-modal task, which aims to align visual entities with natural language descriptions. Current methods either fail to leverage the local de…
cs.CV2023
TG-VQA: Ternary Game of Video Question Answering
Hao Li, Peng Jin, Zesen Cheng +5
Video question answering aims at answering a question about the video content by reasoning the alignment semantics within them. However, since relying heavily on human instructions…
cs.CV2022★ 4 cited
Locality Guidance for Improving Vision Transformers on Tiny Datasets
Kehan Li, Runyi Yu, Zhennan Wang +3
While the Vision Transformer (VT) architecture is becoming trendy in computer vision, pure VT models perform poorly on tiny datasets. To address this issue, this paper proposes the…