3 papers
cs.CV2026
Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation
Ken Deng, Yifu Qiu, Yoni Kasten +2
We study whether vision-language models (VLMs) can solve relative camera pose estimation (RCPE) from image pairs, a direct test of multi-view spatial reasoning. We cast RCPE as a d…
cs.CV2025
GeoSAM2: Unleashing the Power of SAM2 for 3D Part Segmentation
Ken Deng, Yunhan Yang, Jingxiang Sun +4
We introduce GeoSAM2, a prompt-controllable framework for 3D part segmentation that casts the task as multi-view 2D mask prediction. Given a textureless object, we render normal an…
cs.CV2025
DetailGen3D: Generative 3D Geometry Enhancement via Data-Dependent Flow
Ken Deng, Yuan-Chen Guo, Jingxiang Sun +6
Modern 3D generation methods can rapidly create shapes from sparse or single views, but their outputs often lack geometric detail due to computational constraints. We present Detai…