12 citations · 12 across the 2 of their papers we have counts for
1 paper · 1 filter
Xichen Pan, Peiyu Chen, Yichen Gong +3
Training Transformer-based models demands a large amount of data, while obtaining aligned and labelled data in multimodality is rather cost-demanding, especially for audio-visual s…