1 paper
Haotang Li, Zhenyu Qi, Shaohan Henry Wang +5
Visual Geometry Transformer (VGGT) is a strong feed-forward model for multiple 3D tasks, but its Alternating-Attention (AA) stack scales quadratically in the total token count, mak…