90 citations · 285 across the 29 of their papers we have counts for
5 papers · 1 filter
ROSITA: Enhancing Vision-and-Language Semantic Alignments via Cross- and Intra-modal Knowledge Integration
Yuhao Cui, Zhou Yu, Chunqi Wang +4
Vision-and-language pretraining (VLP) aims to learn generic multimodal representations from massive image-text pairs. While various successful attempts have been proposed, learning…
Deep Multimodal Neural Architecture Search
Zhou Yu, Yuhao Cui, Jun Yu +3
Designing effective neural networks is fundamentally important in deep multimodal learning. Most existing works focus on a single task and design neural architectures manually, whi…
Weakly-Supervised Multi-Level Attentional Reconstruction Network for Grounding Textual Queries in Videos
Yijun Song, Jingwen Wang, Lin Ma +2
The task of temporally grounding textual queries in videos is to localize one video segment that semantically corresponds to the given query. Most of the existing approaches rely o…
Multimodal Unified Attention Networks for Vision-and-Language Interactions
Zhou Yu, Yuhao Cui, Jun Yu +2
Learning an effective attention mechanism for multimodal data is important in many vision-and-language tasks that require a synergic understanding of both the visual and textual co…
ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering
Zhou Yu, Dejing Xu, Jun Yu +4
Recent developments in modeling language and vision have been successfully applied to image question answering. It is both crucial and natural to extend this research direction to…