3 citations · 3 across the 1 of their papers we have counts for
1 paper
Lei Shi, Shijie Geng, Kai Shuang +4
Multi-modality fusion technologies have greatly improved the performance of neural network-based Video Description/Caption, Visual Question Answering (VQA) and Audio Visual Scene-a…