2 citations · 2 across the 5 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Multi-Image Visual Token Pruning in Large Visual Language Models
Rongyang Zhang, Chengqiang Lu, Cong Li +9
With the growing demand for processing multiple image sequences in real-world applications, various visual token pruning methods have emerged to mitigate the computational and cont…
cs.CV2024★ 2 cited
VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision Computation
Shiwei Wu, Joya Chen, Kevin Qinghong Lin +7
A well-known dilemma in large vision-language models (e.g., GPT-4, LLaVA) is that while increasing the number of vision tokens generally enhances visual understanding, it also sign…