activity
20192022
most citedSemVLP: Vision-Language Pre-training by Aligning Semantics at Multiple Levels

20 citations · 46 across the 5 of their papers we have counts for

collaborators

10 papers

cs.CV20226 cited

Image Captioning In the Transformer Age

Yang Xu, Li Li, Haiyang Xu +3

Image Captioning (IC) has achieved astonishing developments by incorporating various techniques into the CNN-RNN encoder-decoder architecture. However, since CNN and RNN do not sha…

cs.MM20216 cited

Grid-VLP: Revisiting Grid Features for Vision-Language Pre-training

Ming Yan, Haiyang Xu, Chenliang Li +4

Existing approaches to vision-language pre-training (VLP) heavily rely on an object detector based on bounding boxes (regions), where salient objects are first detected from images…

cs.CV20218 cited

E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual Learning

Haiyang Xu, Ming Yan, Chenliang Li +4

Vision-language pre-training (VLP) on large-scale image-text pairs has achieved huge success for the cross-modal downstream tasks. The most existing pre-training methods mainly ado…

cs.CL202120 cited

SemVLP: Vision-Language Pre-training by Aligning Semantics at Multiple Levels

Chenliang Li, Ming Yan, Haiyang Xu +4

Vision-language pre-training (VLP) on large-scale image-text pairs has recently witnessed rapid progress for learning cross-modal representations. Existing pre-training methods eit…

cs.CL20206 cited

Neural Topic Modeling with Bidirectional Adversarial Training

Rui Wang, Xuemeng Hu, Deyu Zhou +4

Recent years have witnessed a surge of interests of using neural topic models for automatic topic extraction from text, since they avoid the complicated mathematical derivations fo…

cs.CL2020

Adversarial Multi-Binary Neural Network for Multi-class Classification

Haiyang Xu, Junwen Chen, Kun Han +1

Multi-class text classification is one of the key problems in machine learning and natural language processing. Emerging neural networks deal with the problem using a multi-output…