25 citations · 25 across the 1 of their papers we have counts for
1 paper
Junlong Gao, Xi Meng, Shiqi Wang +4
Existing captioning models often adopt the encoder-decoder architecture, where the decoder uses autoregressive decoding to generate captions, such that each token is generated sequ…