4 citations · 5 across the 3 of their papers we have counts for
5 papers
Joint learning of object graph and relation graph for visual question answering
Hao Li, Xu Li, Belhal Karimi +2
Modeling visual question answering(VQA) through scene graphs can significantly improve the reasoning accuracy and interpretability. However, existing models answer poorly for compl…
ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval
Mengjun Cheng, Yipeng Sun, Longchao Wang +8
Visual appearance is considered to be the most important cue to understand images for cross-modal retrieval, while sometimes the scene text appearing in images can provide valuable…
Adaptively Aligned Image Captioning via Adaptive Attention Time
Lun Huang, Wenmin Wang, Yaxian Xia +1
Recent neural models for image captioning usually employ an encoder-decoder framework with an attention mechanism. However, the attention mechanism in such a framework aligns one s…
Attention on Attention for Image Captioning
Lun Huang, Wenmin Wang, Jie Chen +1
Attention mechanisms are widely used in current encoder/decoder frameworks of image captioning, where a weighted average on encoded vectors is generated at each time step to guide…
ParNet: Position-aware Aggregated Relation Network for Image-Text matching
Yaxian Xia, Lun Huang, Wenmin Wang +1
Exploring fine-grained relationship between entities(e.g. objects in image or words in sentence) has great contribution to understand multimedia content precisely. Previous attenti…