89 citations · 198 across the 24 of their papers we have counts for
6 papers · 2 filters
CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation
Lin Sun, Jiale Cao, Jin Xie +2
Contrastive Language-Image Pre-training (CLIP) exhibits strong zero-shot classification ability on various image-level tasks, leading to the research to adapt CLIP for pixel-level…
iSeg: An Iterative Refinement-based Framework for Training-free Segmentation
Lin Sun, Jiale Cao, Jin Xie +2
Stable diffusion has demonstrated strong image synthesis ability to given text descriptions, suggesting it to contain strong semantic clue for grouping objects. The researchers hav…
Multi-Granularity Language-Guided Training for Multi-Object Tracking
Yuhao Li, Jiale Cao, Muzammal Naseer +4
Most existing multi-object tracking methods typically learn visual tracking features via maximizing dis-similarities of different instances and minimizing similarities of the same…
Vision Foundation Model Driven Foreground-Aware Pseudo-LiDAR Generation for Monocular 3D Object Detection
Bonan Ding, Jin Xie, Jing Nie +2
Pseudo-LiDAR has become a promising paradigm for monocular 3D object detection by transforming monocular images into point cloud representations that can be processed by LiDAR-base…
Implicit and Explicit Language Guidance for Diffusion-based Visual Perception
Hefeng Wang, Jiale Cao, Jin Xie +2
Text-to-image diffusion models have shown powerful ability on conditional image synthesis. With large-scale vision-language pre-training, diffusion models are able to generate high…
CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance Segmentation
Wenqi Zhu, Jiale Cao, Jin Xie +2
Open-vocabulary video instance segmentation strives to segment and track instances belonging to an open set of categories in a videos. The vision-language model Contrastive Languag…