5 papers
Towards Interactive Global Geolocation Assistant
Zhiyang Dou, Zipeng Wang, Xumeng Han +3
Global geolocation, which seeks to predict the geographical location of images captured anywhere in the world, is one of the most challenging tasks in the field of computer vision.…
HeroGS: Hierarchical Guidance for Robust 3D Gaussian Splatting under Sparse Views
Jiashu Li, Xumeng Han, Zhaoyang Wei +5
3D Gaussian Splatting (3DGS) has recently emerged as a promising approach in novel view synthesis, combining photorealistic rendering with real-time efficiency. However, its succes…
P2Object: Single Point Supervised Object Detection and Instance Segmentation
Pengfei Chen, Xuehui Yu, Xumeng Han +5
Object recognition using single-point supervision has attracted increasing attention recently. However, the performance gap compared with fully-supervised algorithms remains large.…
Boosting Segment Anything Model Towards Open-Vocabulary Learning
Xumeng Han, Longhui Wei, Xuehui Yu +6
The recent Segment Anything Model (SAM) has emerged as a new paradigmatic vision foundation model, showcasing potent zero-shot generalization and flexible prompting. Despite SAM fi…
ClickTrack: Towards Real-time Interactive Single Object Tracking
Kuiran Wang, Xuehui Yu, Wenwen Yu +5
Single object tracking(SOT) relies on precise object bounding box initialization. In this paper, we reconsidered the deficiencies in the current approaches to initializing single o…