21 citations · 23 across the 13 of their papers we have counts for
1 paper · 1 filter
Mingxiao Huo, Jiayi Zhang, Hewei Wang +4
Vision-Language Models (VLMs) enable powerful multimodal reasoning but suffer from slow autoregressive inference, limiting their deployment in real-time applications. We introduce…