38 citations · 38 across the 2 of their papers we have counts for
2 papers
cs.CV2023
All in One: Exploring Unified Vision-Language Tracking with Multi-Modal Alignment
Chunhui Zhang, Xin Sun, Yiqian Yang +4
Current mainstream vision-language (VL) tracking framework consists of three parts, \ie a visual feature extractor, a language feature extractor, and a fusion model. To pursue bett…
cs.CV2022★ 38 cited
You Need to Read Again: Multi-granularity Perception Network for Moment Retrieval in Videos
Xin Sun, Xuan Wang, Jialin Gao +2
Moment retrieval in videos is a challenging task that aims to retrieve the most relevant video moment in an untrimmed video given a sentence description. Previous methods tend to p…