7 papers
Not All Attention Heads Contribute to Critical Visual Token Selection: Head-Aware Pruning Matters More
Chaofang Ma, Lin Jiang, Carol Jingyi Li +4
Vision-Language Models (VLMs) have exhibited impressive performance across diverse visual scenarios. However, this success comes at the cost of explosive growth in visual tokens, w…
Not All Tasks Quantize Equally: Fisher-Guided Quantization for Visual Geometry Transformer
Yipu Zhang, Jintao Cheng, Weilun Feng +5
Feed-forward 3D reconstruction models, represented by Visual Geometry Grounded Transformer (VGGT), jointly predict multiple visual geometry tasks such as depth estimation, camera p…
Training-Free Interaction-Aligned Visual Token Pruning for Efficient Embodied Manipulation
Jintao Cheng, Weibin Li, Haozhe Wang +7
Efficient visual representation is a central image-processing challenge in embodied manipulation, where policies repeatedly process dense visual-token sequences during closed-loop…
Beyond First-Order: Learning Riemannian Geometries for Invariant Visual Place Recognition
Jintao Cheng, Weibin Li, Zhijian He +3
Visual Place Recognition (VPR) demands representations robust to drastic environmental and viewpoint shifts. Existing aggregation paradigms either depend on extensive supervised tr…
VersaQ-3D: Architecture Support for Visual Geometry Grounded Transformers via Versatile Quantization
Yipu Zhang, Jintao Cheng, Xingyu Liu +8
3D reconstruction and view synthesis are fundamental to AR/VR, robotics, and digital twins. The Visual Geometry Grounded Transformer (VGGT) enables strong feed-forward 3D reconstru…
Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time
Jintao Cheng, Weibin Li, Jiehao Luo +5
Visual Place Recognition (VPR) has evolved from handcrafted descriptors to deep learning approaches, yet significant challenges remain. Current approaches, including Vision Foundat…