3 papers
cs.CV2026
Models as Tools: An Agentic Coordination Framework for Unified Multimodal Visual Tracking
Wenrui Cai, Yuzhe Li, Qingjie Liu +1
Most current visual trackers adopt a matching-based architecture trained exclusively on tracking datasets, whose performance gains depend heavily on the length of the input context…
cs.CV2026
Uni-MDTrack: Learning Decoupled Memory and Dynamic States for Parameter-Efficient Visual Tracking in All Modality
Wenrui Cai, Zhenyi Lu, Yuzhe Li +4
With the advent of Transformer-based one-stream trackers that possess strong capability in inter-frame relation modeling, recent research has increasingly focused on how to introdu…
cs.CV2026
Seeing Straight: Document Orientation Detection for Efficient OCR
Suranjan Goswami, Abhinav Ravi, Raja Kolla +5
Despite significant advances in document understanding, determining the correct orientation of scanned or photographed documents remains a critical pre-processing step in the real…