collaborators

7 papers

cs.CV2026

Not All Attention Heads Contribute to Critical Visual Token Selection: Head-Aware Pruning Matters More

Chaofang Ma, Lin Jiang, Carol Jingyi Li +4

Vision-Language Models (VLMs) have exhibited impressive performance across diverse visual scenarios. However, this success comes at the cost of explosive growth in visual tokens, w…

cs.CV2026

Not All Tasks Quantize Equally: Fisher-Guided Quantization for Visual Geometry Transformer

Yipu Zhang, Jintao Cheng, Weilun Feng +5

Feed-forward 3D reconstruction models, represented by Visual Geometry Grounded Transformer (VGGT), jointly predict multiple visual geometry tasks such as depth estimation, camera p…

cs.CV2026

Training-Free Interaction-Aligned Visual Token Pruning for Efficient Embodied Manipulation

Jintao Cheng, Weibin Li, Haozhe Wang +7

Efficient visual representation is a central image-processing challenge in embodied manipulation, where policies repeatedly process dense visual-token sequences during closed-loop…

cs.CV2026

Beyond First-Order: Learning Riemannian Geometries for Invariant Visual Place Recognition

Jintao Cheng, Weibin Li, Zhijian He +3

Visual Place Recognition (VPR) demands representations robust to drastic environmental and viewpoint shifts. Existing aggregation paradigms either depend on extensive supervised tr…

cs.AR2026

VersaQ-3D: Architecture Support for Visual Geometry Grounded Transformers via Versatile Quantization

Yipu Zhang, Jintao Cheng, Xingyu Liu +8

3D reconstruction and view synthesis are fundamental to AR/VR, robotics, and digital twins. The Visual Geometry Grounded Transformer (VGGT) enables strong feed-forward 3D reconstru…

cs.LG2025

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time

Jintao Cheng, Weibin Li, Jiehao Luo +5

Visual Place Recognition (VPR) has evolved from handcrafted descriptors to deep learning approaches, yet significant challenges remain. Current approaches, including Vision Foundat…