3 citations · 3 across the 2 of their papers we have counts for
1 paper · 1 filter
Yixuan Weng, Bin Li
The goal of visual answering localization (VAL) in the video is to obtain a relevant and concise time clip from a video as the answer to the given natural language question. Early…