most citedImproving Referring Expression Grounding with Cross-modal Attention-guided Erasing

24 citations · 31 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CV20207 cited

PV-NAS: Practical Neural Architecture Search for Video Recognition

Zihao Wang, Chen Lin, Lu Sheng +2

Recently, deep learning has been utilized to solve video recognition problem due to its prominent representation ability. Deep neural networks for video tasks is highly customized…

cs.LG2020

End-To-End Graph-based Deep Semi-Supervised Learning

Zihao Wang, Enmei Tu, Zhou Meng

The quality of a graph is determined jointly by three key factors of the graph: nodes, edges and similarity measure (or edge weights), and is very crucial to the success of graph-b…

cs.LG2019

Semi-Supervised Deep Learning Using Improved Unsupervised Discriminant Projection

Xiao Han, Zihao Wang, Enmei Tu +2

Deep learning demands a huge amount of well-labeled data to train the network parameters. How to use the least amount of labeled data to obtain the desired classification accuracy…

cs.CV2019

CAMP: Cross-Modal Adaptive Message Passing for Text-Image Retrieval

Zihao Wang, Xihui Liu, Hongsheng Li +4

Text-image cross-modal retrieval is a challenging task in the field of language and vision. Most previous approaches independently embed images and sentences into a joint embedding…

cs.CV201924 cited

Improving Referring Expression Grounding with Cross-modal Attention-guided Erasing

Xihui Liu, Zihao Wang, Jing Shao +2

Referring expression grounding aims at locating certain objects or persons in an image with a referring expression, where the key challenge is to comprehend and align various types…