7 papers
TTF: Temporal Token Fusion for Efficient Video-Language Model
Simin Huo, Ning LI
Video-language models (VLMs) face rapid inference costs as visual token counts scale with video length. For example, 32 frames at resolution already yield >8,000 v…
Boosting Self-Supervised Tracking with Contextual Prompts and Noise Learning
Yaozong Zheng, Qihua Liang, Bineng Zhong +4
Learning robust contextual knowledge from unlabeled videos is essential for advancing self-supervised tracking. However, conventional self-supervised trackers lack effective contex…
MaMe & MaRe: Matrix-Based Token Merging and Restoration for Efficient Visual Perception and Synthesis
Simin Huo, Ning Li
Token compression is crucial for mitigating the quadratic complexity of self-attention mechanisms in Vision Transformers (ViTs), which often involve numerous input tokens. Existing…
Robust RGB-T Tracking via Learnable Visual Fourier Prompt Fine-tuning and Modality Fusion Prompt Generation
Hongtao Yang, Bineng Zhong, Qihua Liang +3
Recently, visual prompt tuning is introduced to RGB-Thermal (RGB-T) tracking as a parameter-efficient finetuning (PEFT) method. However, these PEFT-based RGB-T tracking methods typ…
Exploring Decoupled Spatio-Temporal Consistency Learning and Self-Prompting Evolution for Self-Supervised Tracking
Yaozong Zheng, Bineng Zhong, Qihua Liang +4
The success of visual tracking has been largely driven by datasets with manual box annotations. However, these box annotations require tremendous human effort, limiting the scale a…
Similarity-Guided Layer-Adaptive Vision Transformer for UAV Tracking
Chaocan Xue, Bineng Zhong, Qihua Liang +4
Vision transformers (ViTs) have emerged as a popular backbone for visual tracking. However, complete ViT architectures are too cumbersome to deploy for unmanned aerial vehicle (UAV…