5 papers
Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes
Zhaoyang Wei, Bowen Jiang, Xumeng Han +6
While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deteriorates significantly in complex scenes unde…
Towards Interactive Global Geolocation Assistant
Zhiyang Dou, Zipeng Wang, Xumeng Han +3
Global geolocation, which seeks to predict the geographical location of images captured anywhere in the world, is one of the most challenging tasks in the field of computer vision.…
HeroGS: Hierarchical Guidance for Robust 3D Gaussian Splatting under Sparse Views
Jiashu Li, Xumeng Han, Zhaoyang Wei +5
3D Gaussian Splatting (3DGS) has recently emerged as a promising approach in novel view synthesis, combining photorealistic rendering with real-time efficiency. However, its succes…
SAPNet++: Evolving Point-Prompted Instance Segmentation with Semantic and Spatial Awareness
Zhaoyang Wei, Xumeng Han, Xuehui Yu +4
Single-point annotation is increasingly prominent in visual tasks for labeling cost reduction. However, it challenges tasks requiring high precision, such as the point-prompted ins…
P2Object: Single Point Supervised Object Detection and Instance Segmentation
Pengfei Chen, Xuehui Yu, Xumeng Han +5
Object recognition using single-point supervision has attracted increasing attention recently. However, the performance gap compared with fully-supervised algorithms remains large.…