34 citations · 39 across the 4 of their papers we have counts for
3 papers · 1 filter
Mounting Video Metadata on Transformer-based Language Model for Open-ended Video Question Answering
Donggeon Lee, Seongho Choi, Youwon Jang +1
Video question answering has recently received a lot of attention from multimodal video researchers. Most video question answering datasets are usually in the form of multiple-choi…
Multimodal Dual Attention Memory for Video Story Question Answering
Kyung-Min Kim, Seong-Ho Choi, Jin-Hwa Kim +1
We propose a video story question-answering (QA) architecture, Multimodal Dual Attention Memory (MDAM). The key idea is to use a dual attention mechanism with late fusion. MDAM use…
DeepStory: Video Story QA by Deep Embedded Memory Networks
Kyung-Min Kim, Min-Oh Heo, Seong-Ho Choi +1
Question-answering (QA) on video contents is a significant challenge for achieving human-level intelligence as it involves both vision and language in real-world settings. Here we…