4 papers · 1 filter
VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence
Chenhui Qiang, Zhaoyang Wei, Xumeng Han +5
With the rapid development of MLLMs, evaluating their visual capabilities has become increasingly crucial. Current benchmarks primarily fall into two main types: basic perception b…
Boosting Segment Anything Model Towards Open-Vocabulary Learning
Xumeng Han, Longhui Wei, Xuehui Yu +6
The recent Segment Anything Model (SAM) has emerged as a new paradigmatic vision foundation model, showcasing potent zero-shot generalization and flexible prompting. Despite SAM fi…
Rethinking Sampling Strategies for Unsupervised Person Re-identification
Xumeng Han, Xuehui Yu, Guorong Li +5
Unsupervised person re-identification (re-ID) remains a challenging task. While extensive research has focused on the framework design and loss function, this paper shows that samp…
ClickTrack: Towards Real-time Interactive Single Object Tracking
Kuiran Wang, Xuehui Yu, Wenwen Yu +5
Single object tracking(SOT) relies on precise object bounding box initialization. In this paper, we reconsidered the deficiencies in the current approaches to initializing single o…