4 papers
Geometry-Aware Test-Time Learning for Quantitative Spatial Reasoning
Gege Zhang, Shuaicheng Niu, Gang Dai +2
Quantitative spatial reasoning in visual-language models (VLMs) aims to infer spatial distances and directional relationships among objects in 3D space from a 2D image and a natura…
OrbitNVS: Harnessing Video Diffusion Priors for Novel View Synthesis
Jinglin Liang, Zijian Zhou, Rui Huang +2
Novel View Synthesis (NVS) aims to generate unseen views of a 3D object given a limited number of known views. Existing methods often struggle to synthesize plausible views for uno…
MPDrive: Improving Spatial Understanding with Marker-Based Prompt Learning for Autonomous Driving
Zhiyuan Zhang, Xiaofan Li, Zhihao Xu +4
Autonomous driving visual question answering (AD-VQA) aims to answer questions related to perception, prediction, and planning based on given driving scene images, heavily relying…
SEG-SAM: Semantic-Guided SAM for Unified Medical Image Segmentation
Shuangping Huang, Hao Liang, Qingfeng Wang +3
Recently, developing unified medical image segmentation models gains increasing attention, especially with the advent of the Segment Anything Model (SAM). SAM has shown promising b…