1 paper
Ao Xiang, Zongqing Qi, Han Wang +2
This paper introduces a new multi-modal model based on the Transformer architecture and tensor product fusion strategy, combining BERT's text vectors and ViT's image vectors to cla…