7 citations · 7 across the 1 of their papers we have counts for
1 paper
Yan Xia, Zhou Zhao, Shangwei Ye +3
In this paper, we introduce a new task, spoken video grounding (SVG), which aims to localize the desired video fragments from spoken language descriptions. Compared with using text…