1 paper
Peter Anderson, Xiaodong He, Chris Buehler +4
Top-down visual attention mechanisms have been used extensively in image captioning and visual question answering (VQA) to enable deeper image understanding through fine-grained an…