1 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.CV2023★ 1 cited
VGSG: Vision-Guided Semantic-Group Network for Text-based Person Search
Shuting He, Hao Luo, Wei Jiang +2
Text-based Person Search (TBPS) aims to retrieve images of target pedestrian indicated by textual descriptions. It is essential for TBPS to extract fine-grained local features and…
cs.CV2023★ 1 cited
Dual-view Curricular Optimal Transport for Cross-lingual Cross-modal Retrieval
Yabing Wang, Shuhui Wang, Hao Luo +5
Current research on cross-modal retrieval is mostly English-oriented, as the availability of a large number of English-oriented human-labeled vision-language corpora. In order to b…
cs.CV2023★ 1 cited
Revisiting Vision Transformer from the View of Path Ensemble
Shuning Chang, Pichao Wang, Hao Luo +2
Vision Transformers (ViTs) are normally regarded as a stack of transformer layers. In this work, we propose a novel view of ViTs showing that they can be seen as ensemble networks…