99 citations · 103 across the 2 of their papers we have counts for
4 papers · 1 filter
ROSITA: Enhancing Vision-and-Language Semantic Alignments via Cross- and Intra-modal Knowledge Integration
Yuhao Cui, Zhou Yu, Chunqi Wang +4
Vision-and-language pretraining (VLP) aims to learn generic multimodal representations from massive image-text pairs. While various successful attempts have been proposed, learning…
Deep Multimodal Neural Architecture Search
Zhou Yu, Yuhao Cui, Jun Yu +3
Designing effective neural networks is fundamentally important in deep multimodal learning. Most existing works focus on a single task and design neural architectures manually, whi…
Multimodal Unified Attention Networks for Vision-and-Language Interactions
Zhou Yu, Yuhao Cui, Jun Yu +2
Learning an effective attention mechanism for multimodal data is important in many vision-and-language tasks that require a synergic understanding of both the visual and textual co…
Deep Modular Co-Attention Networks for Visual Question Answering
Zhou Yu, Jun Yu, Yuhao Cui +2
Visual Question Answering (VQA) requires a fine-grained and simultaneous understanding of both the visual content of images and the textual content of questions. Therefore, designi…