23 citations · 23 across the 2 of their papers we have counts for
2 papers
cs.CV2023
Divert More Attention to Vision-Language Object Tracking
Mingzhe Guo, Zhipeng Zhang, Liping Jing +2
Multimodal vision-language (VL) learning has noticeably pushed the tendency toward generic intelligence owing to emerging large foundation models. However, tracking, as a fundament…
cs.CV2022★ 23 cited
Divert More Attention to Vision-Language Tracking
Mingzhe Guo, Zhipeng Zhang, Heng Fan +1
Relying on Transformer for complex visual feature learning, object tracking has witnessed the new standard for state-of-the-arts (SOTAs). However, this advancement accompanies by l…