1 paper
Jiajun Du, Yu Qin, Hongtao Lu +1
Most attention-based image captioning models attend to the image once per word. However, attending once per word is rigid and is easy to miss some information. Attending more times…