11 citations · 11 across the 1 of their papers we have counts for
1 paper
Longteng Guo, Jing Liu, Xinxin Zhu +3
Most image captioning models are autoregressive, i.e. they generate each word by conditioning on previously generated words, which leads to heavy latency during inference. Recently…