487 citations · 493 across the 5 of their papers we have counts for
5 papers
SLAN: Self-Locator Aided Network for Cross-Modal Understanding
Jiang-Tian Zhai, Qi Zhang, Tong Wu +4
Learning fine-grained interplay between vision and language allows to a more accurate understanding for VisionLanguage tasks. However, it remains challenging to extract key image r…
ClipCrop: Conditioned Cropping Driven by Vision-Language Model
Zhihang Zhong, Mingxi Cheng, Zhirong Wu +7
Image cropping has progressed tremendously under the data-driven paradigm. However, current approaches do not account for the intentions of the user, which is an issue especially w…
Dual Pyramid Generative Adversarial Networks for Semantic Image Synthesis
Shijie Li, Ming-Ming Cheng, Juergen Gall
The goal of semantic image synthesis is to generate photo-realistic images from semantic label maps. It is highly relevant for tasks like content generation and image editing. Curr…
SegNeXt: Rethinking Convolutional Attention Design for Semantic Segmentation
Meng-Hao Guo, Cheng-Ze Lu, Qibin Hou +3
We present SegNeXt, a simple convolutional network architecture for semantic segmentation. Recent transformer-based models have dominated the field of semantic segmentation due to…
Interactive Style Transfer: All is Your Palette
Zheng Lin, Zhao Zhang, Kang-Rui Zhang +2
Neural style transfer (NST) can create impressive artworks by transferring reference style to content image. Current image-to-image NST methods are short of fine-grained controls,…