31 citations · 62 across the 6 of their papers we have counts for
4 papers · 1 filter
CLIP-GEN: Language-Free Training of a Text-to-Image Generator with CLIP
Zihao Wang, Wei Liu, Qian He +2
Training a text-to-image generator in the general domain (e.g., Dall.e, CogView) requires huge amounts of paired text-image data, which is too expensive to collect. In this paper,…
PV-NAS: Practical Neural Architecture Search for Video Recognition
Zihao Wang, Chen Lin, Lu Sheng +2
Recently, deep learning has been utilized to solve video recognition problem due to its prominent representation ability. Deep neural networks for video tasks is highly customized…
CAMP: Cross-Modal Adaptive Message Passing for Text-Image Retrieval
Zihao Wang, Xihui Liu, Hongsheng Li +4
Text-image cross-modal retrieval is a challenging task in the field of language and vision. Most previous approaches independently embed images and sentences into a joint embedding…
Improving Referring Expression Grounding with Cross-modal Attention-guided Erasing
Xihui Liu, Zihao Wang, Jing Shao +2
Referring expression grounding aims at locating certain objects or persons in an image with a referring expression, where the key challenge is to comprehend and align various types…