activity
20172020
most citedMulti-modal Factorized Bilinear Pooling with Co-Attention Learning for Visual Question Answering

102 citations · 242 across the 14 of their papers we have counts for

collaborators
Showing 2019Show all

8 papers · 1 filter

cs.CV201911 cited

ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering

Zhou Yu, Dejing Xu, Jun Yu +4

Recent developments in modeling language and vision have been successfully applied to image question answering. It is both crucial and natural to extend this research direction to…

cs.CV201930 cited

Multimodal Transformer with Multi-View Visual Representation for Image Captioning

Jun Yu, Jing Li, Zhou Yu +1

Image captioning aims to automatically generate a natural language description of a given image, and most state-of-the-art models have adopted an encoder-decoder framework. The fra…

cs.CV201910 cited

Single Pixel Reconstruction for One-stage Instance Segmentation

Jun Yu, Jinghan Yao, Jian Zhang +2

Object instance segmentation is one of the most fundamental but challenging tasks in computer vision, and it requires the pixel-level image understanding. Most existing approaches…

cs.CV201913 cited

On Exploring Undetermined Relationships for Visual Relationship Detection

Yibing Zhan, Jun Yu, Ting Yu +1

In visual relationship detection, human-notated relationships can be regarded as determinate relationships. However, there are still large amount of unlabeled data, such as object…

cs.CV2019

Local Deep-Feature Alignment for Unsupervised Dimension Reduction

Jian Zhang, Jun Yu, Dacheng Tao

This paper presents an unsupervised deep-learning framework named Local Deep-Feature Alignment (LDFA) for dimension reduction. We construct neighbourhood for each data sample and l…

cs.SI201925 cited

Adapting Stochastic Block Models to Power-Law Degree Distributions

Maoying Qiao, Jun Yu, Wei Bian +2

Stochastic block models (SBMs) have been playing an important role in modeling clusters or community structures of network data. But, it is incapable of handling several complex fe…