6 papers
Models as Tools: An Agentic Coordination Framework for Unified Multimodal Visual Tracking
Wenrui Cai, Yuzhe Li, Qingjie Liu +1
Most current visual trackers adopt a matching-based architecture trained exclusively on tracking datasets, whose performance gains depend heavily on the length of the input context…
Uni-MDTrack: Learning Decoupled Memory and Dynamic States for Parameter-Efficient Visual Tracking in All Modality
Wenrui Cai, Zhenyi Lu, Yuzhe Li +4
With the advent of Transformer-based one-stream trackers that possess strong capability in inter-frame relation modeling, recent research has increasingly focused on how to introdu…
SPMTrack: Spatio-Temporal Parameter-Efficient Fine-Tuning with Mixture of Experts for Scalable Visual Tracking
Wenrui Cai, Qingjie Liu, Yunhong Wang
Most state-of-the-art trackers adopt one-stream paradigm, using a single Vision Transformer for joint feature extraction and relation modeling of template and search region images.…
SeeDNorm: Self-Rescaled Dynamic Normalization
Wenrui Cai, Defa Zhu, Qingjie Liu +1
Normalization layer constitutes an essential component in neural networks. In transformers, the predominantly used RMSNorm constrains vectors to a unit hypersphere, followed by dim…
SkeletonX: Data-Efficient Skeleton-based Action Recognition via Cross-sample Feature Aggregation
Zongye Zhang, Wenrui Cai, Qingjie Liu +1
While current skeleton action recognition models demonstrate impressive performance on large-scale datasets, their adaptation to new application scenarios remains challenging. Thes…
Vision-Language Model for Object Detection and Segmentation: A Review and Evaluation
Yongchao Feng, Yajie Liu, Shuai Yang +13
Vision-Language Model (VLM) have gained widespread adoption in Open-Vocabulary (OV) object detection and segmentation tasks. Despite they have shown promise on OV-related tasks, th…