Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
MMSpec: Benchmarking Speculative Decoding for Vision-Language Models
Hui Shen, Xin Wang, Ping Zhang +11
Vision-language models (VLMs) achieve strong performance on multimodal tasks but suffer from high inference latency due to large model sizes and long multimodal contexts. Speculati…
cs.CV2024
Famba-V: Fast Vision Mamba with Cross-Layer Token Fusion
Hui Shen, Zhongwei Wan, Xin Wang +1
Mamba and Vision Mamba (Vim) models have shown their potential as an alternative to methods based on Transformer architecture. This work introduces Fast Mamba for Vision (Famba-V),…