1 paper
Jinwoo Nam, Daechul Ahn, Dongyeop Kang +2
Understanding videos to localize moments with natural language often requires large expensive annotated video regions paired with language queries. To eliminate the annotation cost…