activity
20242026
collaborators

22 papers

cs.CV2026

Models as Tools: An Agentic Coordination Framework for Unified Multimodal Visual Tracking

Wenrui Cai, Yuzhe Li, Qingjie Liu +1

Most current visual trackers adopt a matching-based architecture trained exclusively on tracking datasets, whose performance gains depend heavily on the length of the input context…

cs.CV2026

LightOcc: Lightweight Spatial Embedding for Efficient Vision-based 3D Occupancy Prediction

Jinqing Zhang, Yanan Zhang, Wenrui Cai +2

Occupancy prediction has garnered increasing attention in recent years for its comprehensive fine-grained environmental representation and strong generalization to open-set objects…

cs.GR2026

Semantic-Aware Motion Encoding for Topology-Agnostic Character Animation

Zongye Zhang, Yuzhuo Cui, Qingjie Liu +1

Generalizing motion representation across diverse characters remains challenging due to significant topological variations in skeletal structures across datasets and species, which…

cs.CV2026

Reasoning-Driven Anomaly Detection and Localization with Image-Level Supervision

Yizhou Jin, Yuezhu Feng, Jinjin Zhang +3

Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning and perceptual abilities for anomaly detection. However, most approaches remain confined to…

cs.CV2026

Uni-MDTrack: Learning Decoupled Memory and Dynamic States for Parameter-Efficient Visual Tracking in All Modality

Wenrui Cai, Zhenyi Lu, Yuzhe Li +4

With the advent of Transformer-based one-stream trackers that possess strong capability in inter-frame relation modeling, recent research has increasingly focused on how to introdu…

cs.CV2026

SPMTrack: Spatio-Temporal Parameter-Efficient Fine-Tuning with Mixture of Experts for Scalable Visual Tracking

Wenrui Cai, Qingjie Liu, Yunhong Wang

Most state-of-the-art trackers adopt one-stream paradigm, using a single Vision Transformer for joint feature extraction and relation modeling of template and search region images.…