102 citations · 242 across the 14 of their papers we have counts for
8 papers · 1 filter
ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering
Zhou Yu, Dejing Xu, Jun Yu +4
Recent developments in modeling language and vision have been successfully applied to image question answering. It is both crucial and natural to extend this research direction to…
Multimodal Transformer with Multi-View Visual Representation for Image Captioning
Jun Yu, Jing Li, Zhou Yu +1
Image captioning aims to automatically generate a natural language description of a given image, and most state-of-the-art models have adopted an encoder-decoder framework. The fra…
Single Pixel Reconstruction for One-stage Instance Segmentation
Jun Yu, Jinghan Yao, Jian Zhang +2
Object instance segmentation is one of the most fundamental but challenging tasks in computer vision, and it requires the pixel-level image understanding. Most existing approaches…
On Exploring Undetermined Relationships for Visual Relationship Detection
Yibing Zhan, Jun Yu, Ting Yu +1
In visual relationship detection, human-notated relationships can be regarded as determinate relationships. However, there are still large amount of unlabeled data, such as object…
Local Deep-Feature Alignment for Unsupervised Dimension Reduction
Jian Zhang, Jun Yu, Dacheng Tao
This paper presents an unsupervised deep-learning framework named Local Deep-Feature Alignment (LDFA) for dimension reduction. We construct neighbourhood for each data sample and l…
Adapting Stochastic Block Models to Power-Law Degree Distributions
Maoying Qiao, Jun Yu, Wei Bian +2
Stochastic block models (SBMs) have been playing an important role in modeling clusters or community structures of network data. But, it is incapable of handling several complex fe…