Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
When Vision Becomes Text: Visual Token Pruning via Cross-Modal Residual Guidance in VLMs
Congyang Ou, Ruike Song, Yang Zhou +3
Abundant visual information strengthens vision-language model (VLM) perception, yet massive visual tokens raise inference costs. Existing visual token pruning methods rely on simil…
cs.CV2026
When Does Sparse MoE Help in Vision? The Role of Backbone Compute Leverage in Sparse Routing
Libo Sun, Po-wei Harn, Peixiong He +1
Mixture-of-Experts (MoE) networks promise favorable accuracy-compute trade-offs, yet practical vision deployments are hindered by expert collapse and limited end-to-end efficiency…