2 citations · 2 across the 1 of their papers we have counts for
3 papers
cs.CV2023
Multi-event Video-Text Retrieval
Gengyuan Zhang, Jisen Ren, Jindong Gu +1
Video-Text Retrieval (VTR) is a crucial multi-modal task in an era of massive video-text data on the Internet. A plethora of work characterized by using a two-stream Vision-Languag…
cs.CV2023
Can Vision-Language Models be a Good Guesser? Exploring VLMs for Times and Location Reasoning
Gengyuan Zhang, Yurui Zhang, Kerui Zhang +1
Vision-Language Models (VLMs) are expected to be capable of reasoning with commonsense knowledge as human beings. One example is that humans can reason where and when an image is t…
cs.CV2022★ 2 cited
CL-CrossVQA: A Continual Learning Benchmark for Cross-Domain Visual Question Answering
Yao Zhang, Haokun Chen, Ahmed Frikha +5
Visual Question Answering (VQA) is a multi-discipline research task. To produce the right answer, it requires an understanding of the visual content of images, the natural language…