activity
20212024
most citedCiteTracker: Correlating Image and Text for Visual Tracking

5 citations · 14 across the 10 of their papers we have counts for

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2024

1st Place Solution for MOSE Track in CVPR 2024 PVUW Workshop: Complex Video Object Segmentation

Deshui Miao, Xin Li, Zhenyu He +2

Tracking and segmenting multiple objects in complex scenes has always been a challenge in the field of video object segmentation, especially in scenarios where objects are occluded…

cs.CV2024

Spatial-Temporal Multi-level Association for Video Object Segmentation

Deshui Miao, Xin Li, Zhenyu He +2

Existing semi-supervised video object segmentation methods either focus on temporal feature matching or spatial-temporal feature modeling. However, they do not address the issues o…

cs.CV20232 cited

Channel and Spatial Relation-Propagation Network for RGB-Thermal Semantic Segmentation

Zikun Zhou, Shukun Wu, Guoqing Zhu +2

RGB-Thermal (RGB-T) semantic segmentation has shown great potential in handling low-light conditions where RGB-based segmentation is hindered by poor RGB imaging quality. The key t…

cs.CV20231 cited

Cross-Modality Proposal-guided Feature Mining for Unregistered RGB-Thermal Pedestrian Detection

Chao Tian, Zikun Zhou, Yuqing Huang +2

RGB-Thermal (RGB-T) pedestrian detection aims to locate the pedestrians in RGB-T image pairs to exploit the complementation between the two modalities for improving detection robus…

cs.CV20235 cited

CiteTracker: Correlating Image and Text for Visual Tracking

Xin Li, Yuqing Huang, Zhenyu He +3

Existing visual tracking methods typically take an image patch as the reference of the target to perform tracking. However, a single image patch cannot provide a complete and preci…

cs.CV20232 cited

Transferable Decoding with Visual Entities for Zero-Shot Image Captioning

Junjie Fei, Teng Wang, Jinrui Zhang +3

Image-to-text generation aims to describe images using natural language. Recently, zero-shot image captioning based on pre-trained vision-language models (VLMs) and large language…