22 citations · 22 across the 4 of their papers we have counts for
4 papers
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
X. Feng, D. Zhang, S. Hu +5
Vision-Language Tracking (VLT) aims to localize a target in video sequences using a visual template and language description. While textual cues enhance tracking potential, current…
Finger in Camera Speaks Everything: Unconstrained Air-Writing for Real-World
Meiqi Wu, Kaiqi Huang, Yuanqiang Cai +3
Air-writing is a challenging task that combines the fields of computer vision and natural language processing, offering an intuitive and natural approach for human-computer interac…
Visual Language Tracking with Multi-modal Interaction: A Robust Benchmark
Xuchen Li, Shiyu Hu, Xiaokun Feng +4
Visual Language Tracking (VLT) enhances tracking by mitigating the limitations of relying solely on the visual modality, utilizing high-level semantic information through language.…
Global Instance Tracking: Locating Target More Like Humans
Shiyu Hu, Xin Zhao, Lianghua Huang +1
Target tracking, the essential ability of the human visual system, has been simulated by computer vision tasks. However, existing trackers perform well in austere experimental envi…