6 papers
An Efficient Token Compression Framework for Visual Object Tracking
Weijing Wu, Qihua Liang, Bineng Zhong +3
Refining visual representations by eliminating their internal feature-level redundancy is crucial for simultaneously optimizing the performance and computational cost of models in…
Learning to Track Instance from Single Nature Language Description
Yaozong Zheng, Bineng Zhong, Qihua Liang +3
How to achieve vision-language (VL) tracking using natural language descriptions from a video sequence \textbf{without relying on any bounding-box ground truth}? In this work, we a…
Let the Model Learn to Feel: Mode-Guided Tonality Injection for Symbolic Music Emotion Recognition
Haiying Xia, Zhongyi Huang, Yumei Tan +1
Music emotion recognition is a key task in symbolic music understanding (SMER). Recent approaches have shown promising results by fine-tuning large-scale pre-trained models (e.g.,…
Explicit Context Reasoning with Supervision for Visual Tracking
Fansheng Zeng, Bineng Zhong, Haiying Xia +4
Contextual reasoning with constraints is crucial for enhancing temporal consistency in cross-frame modeling for visual tracking. However, mainstream tracking algorithms typically a…
Exploring Decoupled Spatio-Temporal Consistency Learning and Self-Prompting Evolution for Self-Supervised Tracking
Yaozong Zheng, Bineng Zhong, Qihua Liang +4
The success of visual tracking has been largely driven by datasets with manual box annotations. However, these box annotations require tremendous human effort, limiting the scale a…
Low-Rank Adaptive Structural Priors for Generalizable Diabetic Retinopathy Grading
Yunxuan Wang, Ray Yin, Yumei Tan +2
Diabetic retinopathy (DR), a serious ocular complication of diabetes, is one of the primary causes of vision loss among retinal vascular diseases. Deep learning methods have been e…