1 citations · 1 across the 1 of their papers we have counts for
1 paper
Yan Zhang, Gangyan Zeng, Huawen Shen +3
Video text-based visual question answering (Video TextVQA) is a practical task that aims to answer questions by jointly reasoning textual and visual information in a given video. I…