5 papers
Toward Visual Grounding: A Survey
Linhui Xiao, Xiaoshan Yang, Xiangyuan Lan +2
Visual Grounding, also known as Referring Expression Comprehension and Phrase Grounding, aims to ground the specific region(s) within the image(s) based on the given expression tex…
SelaVPR++: Towards Seamless Adaptation of Foundation Models for Efficient Place Recognition
Feng Lu, Tong Jin, Xiangyuan Lan +4
Recent studies show that the visual place recognition (VPR) method using pre-trained visual foundation models can achieve promising performance. In our previous work, we propose a…
Limb-Aware Virtual Try-On Network with Progressive Clothing Warping
Shengping Zhang, Xiaoyu Han, Weigang Zhang +3
Image-based virtual try-on aims to transfer an in-shop clothing image to a person image. Most existing methods adopt a single global deformation to perform clothing warping directl…
EMMA: Empowering Multi-modal Mamba with Structural and Hierarchical Alignment
Yifei Xing, Xiangyuan Lan, Ruiping Wang +4
Mamba-based architectures have shown to be a promising new direction for deep learning models owing to their competitive performance and sub-quadratic deployment speed. However, cu…
OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion
Hao Wang, Pengzhen Ren, Zequn Jie +8
Open-vocabulary detection is a challenging task due to the requirement of detecting objects based on class names, including those not encountered during training. Existing methods…