102 citations · 241 across the 4 of their papers we have counts for
4 papers
Deep Modular Co-Attention Networks for Visual Question Answering
Zhou Yu, Jun Yu, Yuhao Cui +2
Visual Question Answering (VQA) requires a fine-grained and simultaneous understanding of both the visual content of images and the textual content of questions. Therefore, designi…
Multimodal Transformer with Multi-View Visual Representation for Image Captioning
Jun Yu, Jing Li, Zhou Yu +1
Image captioning aims to automatically generate a natural language description of a given image, and most state-of-the-art models have adopted an encoder-decoder framework. The fra…
Single Pixel Reconstruction for One-stage Instance Segmentation
Jun Yu, Jinghan Yao, Jian Zhang +2
Object instance segmentation is one of the most fundamental but challenging tasks in computer vision, and it requires the pixel-level image understanding. Most existing approaches…
Multi-modal Factorized Bilinear Pooling with Co-Attention Learning for Visual Question Answering
Zhou Yu, Jun Yu, Jianping Fan +1
Visual question answering (VQA) is challenging because it requires a simultaneous understanding of both the visual content of images and the textual content of questions. The appro…