22 papers
Models as Tools: An Agentic Coordination Framework for Unified Multimodal Visual Tracking
Wenrui Cai, Yuzhe Li, Qingjie Liu +1
Most current visual trackers adopt a matching-based architecture trained exclusively on tracking datasets, whose performance gains depend heavily on the length of the input context…
LightOcc: Lightweight Spatial Embedding for Efficient Vision-based 3D Occupancy Prediction
Jinqing Zhang, Yanan Zhang, Wenrui Cai +2
Occupancy prediction has garnered increasing attention in recent years for its comprehensive fine-grained environmental representation and strong generalization to open-set objects…
Semantic-Aware Motion Encoding for Topology-Agnostic Character Animation
Zongye Zhang, Yuzhuo Cui, Qingjie Liu +1
Generalizing motion representation across diverse characters remains challenging due to significant topological variations in skeletal structures across datasets and species, which…
Reasoning-Driven Anomaly Detection and Localization with Image-Level Supervision
Yizhou Jin, Yuezhu Feng, Jinjin Zhang +3
Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning and perceptual abilities for anomaly detection. However, most approaches remain confined to…
Uni-MDTrack: Learning Decoupled Memory and Dynamic States for Parameter-Efficient Visual Tracking in All Modality
Wenrui Cai, Zhenyi Lu, Yuzhe Li +4
With the advent of Transformer-based one-stream trackers that possess strong capability in inter-frame relation modeling, recent research has increasingly focused on how to introdu…
SPMTrack: Spatio-Temporal Parameter-Efficient Fine-Tuning with Mixture of Experts for Scalable Visual Tracking
Wenrui Cai, Qingjie Liu, Yunhong Wang
Most state-of-the-art trackers adopt one-stream paradigm, using a single Vision Transformer for joint feature extraction and relation modeling of template and search region images.…