2 papers
cs.CV2026
GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
Kaichen Zhou, Yuzhen Chen, Fangneng Zhan +8
Video world models can generate realistic futures from a single instruction, but they often fail to track the same physical points consistently across time. As a result, the genera…
cs.CV2026
PAGE-4D: Disentangled pose and geometry estimation for vggt-4d perception
Kaichen Zhou, Yuhan Wang, Grace Chen +5
Recent 3D feed-forward models, such as the Visual Geometry Grounded Transformer (VGGT), have shown strong capability in inferring 3D attributes of static scenes. However, since the…