2 papers
cs.CV2025
LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging
Zhijian Shu, Cheng Lin, Tao Xie +8
3D vision foundation models like Visual Geometry Grounded Transformer (VGGT) have advanced greatly in geometric perception. However, it is time-consuming and memory-intensive for l…
cs.CV2025
PosePilot: Steering Camera Pose for Generative World Models with Self-supervised Depth
Bu Jin, Weize Li, Baihan Yang +9
Recent advancements in autonomous driving (AD) systems have highlighted the potential of world models in achieving robust and generalizable performance across both ordinary and cha…