Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Multi-Image Visual Token Pruning in Large Visual Language Models
Rongyang Zhang, Chengqiang Lu, Cong Li +9
With the growing demand for processing multiple image sequences in real-world applications, various visual token pruning methods have emerged to mitigate the computational and cont…
cs.CV2024★ 2 cited
Vript: A Video Is Worth Thousands of Words
Dongjie Yang, Suyuan Huang, Chengqiang Lu +5
Advancements in multimodal learning, particularly in video understanding and generation, require high-quality video-text datasets for improved model performance. Vript addresses th…