30 citations · 96 across the 11 of their papers we have counts for
8 papers · 1 filter
Structured Co-reference Graph Attention for Video-grounded Dialogue
Junyeong Kim, Sunjae Yoon, Dahyun Kim +1
A video-grounded dialogue system referred to as the Structured Co-reference Graph Attention (SCGA) is presented for decoding the answer sequence to a question regarding a given vid…
Semantic Grouping Network for Video Captioning
Hobin Ryu, Sunghun Kang, Haeyong Kang +1
This paper considers a video caption generating network referred to as Semantic Grouping Network (SGN) that attempts (1) to group video frames with discriminating word phrases of p…
SCNet: Training Inference Sample Consistency for Instance Segmentation
Thang Vu, Haeyong Kang, Chang D. Yoo
Cascaded architectures have brought significant performance improvement in object detection and instance segmentation. However, there are lingering issues regarding the disparity i…
Modality Shifting Attention Network for Multi-modal Video Question Answering
Junyeong Kim, Minuk Ma, Trung Pham +2
This paper considers a network referred to as Modality Shifting Attention Network (MSAN) for Multimodal Video Question Answering (MVQA) task. MSAN decomposes the task into two sub-…
Cascade RPN: Delving into High-Quality Region Proposal Network with Adaptive Convolution
Thang Vu, Hyunjun Jang, Trung X. Pham +1
This paper considers an architecture referred to as Cascade Region Proposal Network (Cascade RPN) for improving the region-proposal quality and detection performance by \textit{sys…
Gaining Extra Supervision via Multi-task learning for Multi-Modal Video Question Answering
Junyeong Kim, Minuk Ma, Kyungsu Kim +2
This paper proposes a method to gain extra supervision via multi-task learning for multi-modal video question answering. Multi-modal video question answering is an important task t…