3 citations · 4 across the 4 of their papers we have counts for
4 papers · 1 filter
Dense but Efficient VideoQA for Intricate Compositional Reasoning
Jihyeon Lee, Wooyoung Kang, Eun-Sol Kim
It is well known that most of the conventional video question answering (VideoQA) datasets consist of easy questions requiring simple reasoning processes. However, long videos inev…
Video-Text Representation Learning via Differentiable Weak Temporal Alignment
Dohwan Ko, Joonmyung Choi, Juyeon Ko +4
Learning generic joint representations for video and text by a supervised method requires a prohibitively substantial amount of manually annotated video datasets. As a practical al…
Boundary-aware Self-supervised Learning for Video Scene Segmentation
Jonghwan Mun, Minchul Shin, Gunsoo Han +4
Self-supervised learning has drawn attention through its effectiveness in learning in-domain representations with no ground-truth annotations; in particular, it is shown that prope…
Image-to-Image Retrieval by Learning Similarity between Scene Graphs
Sangwoong Yoon, Woo Young Kang, Sungwook Jeon +4
As a scene graph compactly summarizes the high-level content of an image in a structured and symbolic manner, the similarity between scene graphs of two images reflects the relevan…