98 citations · 254 across the 7 of their papers we have counts for
14 papers · 1 filter
FAME-ViL: Multi-Tasking Vision-Language Model for Heterogeneous Fashion Tasks
Xiao Han, Xiatian Zhu, Licheng Yu +3
In the fashion domain, there exists a variety of vision-and-language (V+L) tasks, including cross-modal retrieval, text-guided image retrieval, multi-modal classification, and imag…
UIGR: Unified Interactive Garment Retrieval
Xiao Han, Sen He, Li Zhang +2
Interactive garment retrieval (IGR) aims to retrieve a target garment image based on a reference garment image along with user feedback on what to change on the reference garment.…
Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers
Sixiao Zheng, Jiachen Lu, Hengshuang Zhao +8
Most recent semantic segmentation methods adopt a fully-convolutional network (FCN) with an encoder-decoder architecture. The encoder progressively reduces the spatial resolution a…
Improving Semantic Segmentation via Decoupled Body and Edge Supervision
Xiangtai Li, Xia Li, Li Zhang +5
Existing semantic segmentation approaches either aim to improve the object's inner consistency by modeling the global context, or refine objects detail along their boundaries by mu…
XingGAN for Person Image Generation
Hao Tang, Song Bai, Li Zhang +2
We propose a novel Generative Adversarial Network (XingGAN or CrossingGAN) for person image generation tasks, i.e., translating the pose of a given person to a desired one. The pro…
Self-supervised Video Object Segmentation
Fangrui Zhu, Li Zhang, Yanwei Fu +2
The objective of this paper is self-supervised representation learning, with the goal of solving semi-supervised video object segmentation (a.k.a. dense tracking). We make the foll…