3 papers
cs.CV2025
Global2Local: A Joint-Hierarchical Attention for Video Captioning
Chengpeng Dai, Fuhai Chen, Xiaoshuai Sun +3
Recently, automatic video captioning has attracted increasing attention, where the core challenge lies in capturing the key semantic items, like objects and actions as well as thei…
cs.CV2024
Rethinking Sampling Strategies for Unsupervised Person Re-identification
Xumeng Han, Xuehui Yu, Guorong Li +5
Unsupervised person re-identification (re-ID) remains a challenging task. While extensive research has focused on the framework design and loss function, this paper shows that samp…
cs.CV2024
ClickTrack: Towards Real-time Interactive Single Object Tracking
Kuiran Wang, Xuehui Yu, Wenwen Yu +5
Single object tracking(SOT) relies on precise object bounding box initialization. In this paper, we reconsidered the deficiencies in the current approaches to initializing single o…