94 citations · 138 across the 8 of their papers we have counts for
9 papers · 1 filter
Panoramic Vision Transformer for Saliency Detection in 360° Videos
Heeseung Yun, Sehun Lee, Gunhee Kim
360 video saliency detection is one of the challenging benchmarks for 360 video understanding since non-negligible distortion and discontinuity occur in the project…
Parameter Efficient Multimodal Transformers for Video Representation Learning
Sangho Lee, Youngjae Yu, Gunhee Kim +3
The recent success of Transformers in the language domain has motivated adapting it to a multimodal setting, where a new visual model is trained in tandem with an already pretraine…
CurlingNet: Compositional Learning between Images and Text for Fashion IQ Data
Youngjae Yu, Seunghwan Lee, Yuncheol Choi +1
We present an approach named CurlingNet that can measure the semantic distance of composition of image-text embedding. In order to learn an effective image-text composition for the…
A Joint Sequence Fusion Model for Video Question Answering and Retrieval
Youngjae Yu, Jongseok Kim, Gunhee Kim
We present an approach named JSFusion (Joint Sequence Fusion) that can measure semantic similarity between any pairs of multimodal sequence data (e.g. a video clip and a language s…
Video Prediction with Appearance and Motion Conditions
Yunseok Jang, Gunhee Kim, Yale Song
Video prediction aims to generate realistic future frames by learning dynamic visual patterns. One fundamental challenge is to deal with future uncertainty: How should a model beha…
A Memory Network Approach for Story-based Temporal Summarization of 360° Videos
Sangho Lee, Jinyoung Sung, Youngjae Yu +1
We address the problem of story-based temporal summarization of long 360° videos. We propose a novel memory network model named Past-Future Memory Network (PFMN), in which we first…