activity
20222024
most citedReformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

4 citations · 5 across the 6 of their papers we have counts for

collaborators

6 papers

cs.RO2024

Increasing SLAM Pose Accuracy by Ground-to-Satellite Image Registration

Yanhao Zhang, Yujiao Shi, Shan Wang +4

Vision-based localization for autonomous driving has been of great interest among researchers. When a pre-built 3D map is not available, the techniques of visual simultaneous local…

cs.CV2024

Homography Guided Temporal Fusion for Road Line and Marking Segmentation

Shan Wang, Chuong Nguyen, Jiawei Liu +6

Reliable segmentation of road lines and markings is critical to autonomous driving. Our work is motivated by the observations that road lines and markings are (1) frequently occlud…

cs.CV20234 cited

Reformulating Vision-Language Foundation Models and Datasets Towards Universal Multimodal Assistants

Tianyu Yu, Jinyi Hu, Yuan Yao +10

Recent Multimodal Large Language Models (MLLMs) exhibit impressive abilities to perceive images and follow open-ended instructions. The capabilities of MLLMs depend on two crucial…

cs.CV2023

View Consistent Purification for Accurate Cross-View Localization

Shan Wang, Yanhao Zhang, Akhil Perincherry +2

This paper proposes a fine-grained self-localization method for outdoor robotics that utilizes a flexible number of onboard cameras and readily accessible satellite images. The pro…

cs.CV2023

Model Calibration in Dense Classification with Adaptive Label Perturbation

Jiawei Liu, Changkun Ye, Shan Wang +4

For safety-related applications, it is crucial to produce trustworthy deep neural networks whose prediction is associated with confidence that can represent the likelihood of corre…

cs.CV20221 cited

CVLNet: Cross-View Semantic Correspondence Learning for Video-based Camera Localization

Yujiao Shi, Xin Yu, Shan Wang +1

This paper tackles the problem of Cross-view Video-based camera Localization (CVL). The task is to localize a query camera by leveraging information from its past observations, i.e…