1 paper
Yan Zhang, Gangyan Zeng, Daiqing Wu +5
Video text-based visual question answering (Video TextVQA) aims to answer questions by explicitly reading and reasoning about the text involved in a video. Most works in this field…