99 citations · 160 across the 9 of their papers we have counts for
Showing 2019Show all
2 papers · 1 filter
cs.CV2019
Multimodal Unified Attention Networks for Vision-and-Language Interactions
Zhou Yu, Yuhao Cui, Jun Yu +2
Learning an effective attention mechanism for multimodal data is important in many vision-and-language tasks that require a synergic understanding of both the visual and textual co…
cs.CV2019★ 99 cited
Deep Modular Co-Attention Networks for Visual Question Answering
Zhou Yu, Jun Yu, Yuhao Cui +2
Visual Question Answering (VQA) requires a fine-grained and simultaneous understanding of both the visual content of images and the textual content of questions. Therefore, designi…