20 citations · 46 across the 5 of their papers we have counts for
10 papers
Image Captioning In the Transformer Age
Yang Xu, Li Li, Haiyang Xu +3
Image Captioning (IC) has achieved astonishing developments by incorporating various techniques into the CNN-RNN encoder-decoder architecture. However, since CNN and RNN do not sha…
Grid-VLP: Revisiting Grid Features for Vision-Language Pre-training
Ming Yan, Haiyang Xu, Chenliang Li +4
Existing approaches to vision-language pre-training (VLP) heavily rely on an object detector based on bounding boxes (regions), where salient objects are first detected from images…
E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual Learning
Haiyang Xu, Ming Yan, Chenliang Li +4
Vision-language pre-training (VLP) on large-scale image-text pairs has achieved huge success for the cross-modal downstream tasks. The most existing pre-training methods mainly ado…
SemVLP: Vision-Language Pre-training by Aligning Semantics at Multiple Levels
Chenliang Li, Ming Yan, Haiyang Xu +4
Vision-language pre-training (VLP) on large-scale image-text pairs has recently witnessed rapid progress for learning cross-modal representations. Existing pre-training methods eit…
Neural Topic Modeling with Bidirectional Adversarial Training
Rui Wang, Xuemeng Hu, Deyu Zhou +4
Recent years have witnessed a surge of interests of using neural topic models for automatic topic extraction from text, since they avoid the complicated mathematical derivations fo…
Adversarial Multi-Binary Neural Network for Multi-class Classification
Haiyang Xu, Junwen Chen, Kun Han +1
Multi-class text classification is one of the key problems in machine learning and natural language processing. Emerging neural networks deal with the problem using a multi-output…