2 papers
cs.CV2026
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference
Xu Li, Yi Zheng, Mengyang Zhao +7
Large Vision-Language Models (LVLMs) typically require processing hundreds to thousands of visual tokens, leading to substantial inference overhead. Existing visual token pruning m…
cs.CV2026
Understanding Geometric Representations in Self-Supervised Vision Transformers via Subspace Intervention
Weichen Zhou, Yawen Zou, Chunzhi Gu +3
We introduce a controlled subspace intervention framework to investigate how self-supervised Vision Transformers (ViTs) encode dense geometric information. While linear probing is…