1 paper · 1 filter
Ao Xiang, Zongqing Qi, Han Wang +2
This paper introduces a new multi-modal model based on the Transformer architecture and tensor product fusion strategy, combining BERT's text vectors and ViT's image vectors to cla…