15 citations · 26 across the 4 of their papers we have counts for
5 papers · 1 filter
LOIS: Looking Out of Instance Semantics for Visual Question Answering
Siyu Zhang, Yeming Chen, Yaoru Sun +3
Visual question answering (VQA) has been intensively studied as a multimodal task that requires effort in bridging vision and language to infer answers correctly. Recent attempts h…
Hierarchical Matching and Reasoning for Multi-Query Image Retrieval
Zhong Ji, Zhihao Li, Yan Zhang +3
As a promising field, Multi-Query Image Retrieval (MQIR) aims at searching for the semantically relevant image given multiple region-specific text queries. Existing works mainly fo…
Structural and Statistical Texture Knowledge Distillation for Semantic Segmentation
Deyi Ji, Haoran Wang, Mingyuan Tao +3
Existing knowledge distillation works for semantic segmentation mainly focus on transferring high-level contextual knowledge from teacher to student. However, low-level texture kno…
Step-Wise Hierarchical Alignment Network for Image-Text Matching
Zhong Ji, Kexin Chen, Haoran Wang
Image-text matching plays a central role in bridging the semantic gap between vision and language. The key point to achieve precise visual-semantic alignment lies in capturing the…
Consensus-Aware Visual-Semantic Embedding for Image-Text Matching
Haoran Wang, Ying Zhang, Zhong Ji +2
Image-text matching plays a central role in bridging vision and language. Most existing approaches only rely on the image-text instance pair to learn their representations, thereby…