5 citations · 10 across the 8 of their papers we have counts for
1 paper · 1 filter
Jiajun Du, Yu Qin, Hongtao Lu +1
Most attention-based image captioning models attend to the image once per word. However, attending once per word is rigid and is easy to miss some information. Attending more times…