7 citations · 19 across the 4 of their papers we have counts for
7 papers · 1 filter
Parameter Efficient Multimodal Transformers for Video Representation Learning
Sangho Lee, Youngjae Yu, Gunhee Kim +3
The recent success of Transformers in the language domain has motivated adapting it to a multimodal setting, where a new visual model is trained in tandem with an already pretraine…
CurlingNet: Compositional Learning between Images and Text for Fashion IQ Data
Youngjae Yu, Seunghwan Lee, Yuncheol Choi +1
We present an approach named CurlingNet that can measure the semantic distance of composition of image-text embedding. In order to learn an effective image-text composition for the…
A Joint Sequence Fusion Model for Video Question Answering and Retrieval
Youngjae Yu, Jongseok Kim, Gunhee Kim
We present an approach named JSFusion (Joint Sequence Fusion) that can measure semantic similarity between any pairs of multimodal sequence data (e.g. a video clip and a language s…
A Memory Network Approach for Story-based Temporal Summarization of 360° Videos
Sangho Lee, Jinyoung Sung, Youngjae Yu +1
We address the problem of story-based temporal summarization of long 360° videos. We propose a novel memory network model named Past-Future Memory Network (PFMN), in which we first…
A Deep Ranking Model for Spatio-Temporal Highlight Detection from a 360 Video
Youngjae Yu, Sangho Lee, Joonil Na +2
We address the problem of highlight detection from a 360 degree video by summarizing it both spatially and temporally. Given a long 360 degree video, we spatially select pleasantly…
Supervising Neural Attention Models for Video Captioning by Human Gaze Data
Youngjae Yu, Jongwook Choi, Yeonhwa Kim +3
The attention mechanisms in deep neural networks are inspired by human's attention that sequentially focuses on the most relevant parts of the information over time to generate pre…