9 papers
Dual-branch Distilled Transformer for Efficient Asymmetric UAV Tracking
Hongtao Yang, Bineng Zhong, Qihua Liang +4
Given the real-time demands of UAV tracking, many methods simplify the backbone to reduce computation, but this often weakens feature representation and degrades performance in com…
Learning to Track Instance from Single Nature Language Description
Yaozong Zheng, Bineng Zhong, Qihua Liang +3
How to achieve vision-language (VL) tracking using natural language descriptions from a video sequence \textbf{without relying on any bounding-box ground truth}? In this work, we a…
Boosting Self-Supervised Tracking with Contextual Prompts and Noise Learning
Yaozong Zheng, Qihua Liang, Bineng Zhong +4
Learning robust contextual knowledge from unlabeled videos is essential for advancing self-supervised tracking. However, conventional self-supervised trackers lack effective contex…
Difference Feedback: Generating Multimodal Process-Level Supervision for VLM Reinforcement Learning
Feiding, Yongkang Zhang, Yuhao Liao +10
Vision--language models (VLMs) are increasingly aligned via Group Relative Policy Optimization (GRPO)-style training. However, relying solely on terminal outcome rewards yields spa…
UBATrack: Spatio-Temporal State Space Model for General Multi-Modal Tracking
Qihua Liang, Liang Chen, Yaozong Zheng +3
Multi-modal object tracking has attracted considerable attention by integrating multiple complementary inputs (e.g., thermal, depth, and event data) to achieve outstanding performa…
Robust RGB-T Tracking via Learnable Visual Fourier Prompt Fine-tuning and Modality Fusion Prompt Generation
Hongtao Yang, Bineng Zhong, Qihua Liang +3
Recently, visual prompt tuning is introduced to RGB-Thermal (RGB-T) tracking as a parameter-efficient finetuning (PEFT) method. However, these PEFT-based RGB-T tracking methods typ…