5 citations · 13 across the 24 of their papers we have counts for
1 paper · 2 filters
Chengyue Wu, Shiyi Lan, Yonggan Fu +9
Vision-language models (VLMs) predominantly rely on autoregressive decoding, which generates tokens one at a time and fundamentally limits inference throughput. This limitation is…