most citedViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval

4 citations · 5 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CV20221 cited

Joint learning of object graph and relation graph for visual question answering

Hao Li, Xu Li, Belhal Karimi +2

Modeling visual question answering(VQA) through scene graphs can significantly improve the reasoning accuracy and interpretability. However, existing models answer poorly for compl…

cs.CV20224 cited

ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval

Mengjun Cheng, Yipeng Sun, Longchao Wang +8

Visual appearance is considered to be the most important cue to understand images for cross-modal retrieval, while sometimes the scene text appearing in images can provide valuable…

cs.CV2019

Adaptively Aligned Image Captioning via Adaptive Attention Time

Lun Huang, Wenmin Wang, Yaxian Xia +1

Recent neural models for image captioning usually employ an encoder-decoder framework with an attention mechanism. However, the attention mechanism in such a framework aligns one s…

cs.CV2019

Attention on Attention for Image Captioning

Lun Huang, Wenmin Wang, Jie Chen +1

Attention mechanisms are widely used in current encoder/decoder frameworks of image captioning, where a weighted average on encoded vectors is generated at each time step to guide…

cs.CV2019

ParNet: Position-aware Aggregated Relation Network for Image-Text matching

Yaxian Xia, Lun Huang, Wenmin Wang +1

Exploring fine-grained relationship between entities(e.g. objects in image or words in sentence) has great contribution to understand multimedia content precisely. Previous attenti…