8 papers
PAGE-4D: Disentangled pose and geometry estimation for vggt-4d perception
Kaichen Zhou, Yuhan Wang, Grace Chen +5
Recent 3D feed-forward models, such as the Visual Geometry Grounded Transformer (VGGT), have shown strong capability in inferring 3D attributes of static scenes. However, since the…
Neural Surface Reconstruction from Sparse Views Using Epipolar Geometry
Xinhai Chang, Kaichen Zhou
Reconstructing accurate surfaces from sparse multi-view images remains challenging due to severe geometric ambiguity and occlusions. Existing generalizable neural surface reconstru…
Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
Yifan Liu, Fangneng Zhan, Kaichen Zhou +3
Vision-language models (VLMs) struggle with 3D-related tasks such as spatial cognition and physical understanding, which are crucial for real-world applications like robotics and e…
WiCompass: Oracle-driven Data Scaling for mmWave Human Pose Estimation
Bo Liang, Chen Gong, Haobo Wang +8
Millimeter-wave Human Pose Estimation (mmWave HPE) promises privacy but suffers from poor generalization under distribution shifts. We demonstrate that brute-force data scaling is…
K-Sort Eval: Efficient Preference Evaluation for Visual Generation via Corrected VLM-as-a-Judge
Zhikai Li, Jiatong Li, Xuewen Liu +7
The rapid development of visual generative models raises the need for more scalable and human-aligned evaluation methods. While the crowdsourced Arena platforms offer human prefere…
Memorization in 3D Shape Generation: An Empirical Study
Shu Pu, Boya Zeng, Kaichen Zhou +2
Generative models are increasingly used in 3D vision to synthesize novel shapes, yet it remains unclear whether their generation relies on memorizing training shapes. Understanding…