4 citations · 4 across the 3 of their papers we have counts for
3 papers
cs.CV2023
Improving Vision-and-Language Reasoning via Spatial Relations Modeling
Cheng Yang, Rui Xu, Ye Guo +5
Visual commonsense reasoning (VCR) is a challenging multi-modal task, which requires high-level cognition and commonsense reasoning ability about the real world. In recent years, l…
cs.CV2023
A Large Cross-Modal Video Retrieval Dataset with Reading Comprehension
Weijia Wu, Yuzhong Zhao, Zhuang Li +4
Most existing cross-modal language-to-video retrieval (VR) research focuses on single-modal input from video, i.e., visual representation, while the text is omnipresent in human en…
cs.CV2022★ 4 cited
Real-time End-to-End Video Text Spotter with Contrastive Representation Learning
Wejia Wu, Zhuang Li, Jiahong Li +5
Video text spotting(VTS) is the task that requires simultaneously detecting, tracking and recognizing text in the video. Existing video text spotting methods typically develop soph…