3 papers
cs.CV2026
OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning
Geng Li, Guohao Chen, Ting Chen +6
Vision-language models (VLMs) rely on long visual token sequences for visual understanding, making the prefill stage expensive in both computation and memory. Most existing pruning…
cs.CV2026
Rethinking VLM Representation for VLA Initialization
Weifeng Lin, Siyuan Huang, Hao Li +5
Vision-Language-Action (VLA) models widely adopt pretrained Vision-Language Models (VLMs) as policy backbones, yet it remains unclear what kind of pretrained VLM representation is…
cs.CV2025
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Weifeng Lin, Xinyu Wei, Ruichuan An +7
We present Perceive Anything Model (PAM), a conceptually straightforward and efficient framework for comprehensive region-level visual understanding in images and videos. Our appro…