109 citations · 152 across the 10 of their papers we have counts for
1 paper · 1 filter
Longteng Guo, Jing Liu, Xinxin Zhu +3
Most image captioning models are autoregressive, i.e. they generate each word by conditioning on previously generated words, which leads to heavy latency during inference. Recently…