11 citations · 17 across the 2 of their papers we have counts for
1 paper · 1 filter
Longteng Guo, Jing Liu, Xinxin Zhu +3
Most image captioning models are autoregressive, i.e. they generate each word by conditioning on previously generated words, which leads to heavy latency during inference. Recently…