10 citations · 27 across the 10 of their papers we have counts for
1 paper · 1 filter
Minh Tran, Mohammad Soleymani
In this paper, we introduce a pretrained audio-visual Transformer trained on more than 500k utterances from nearly 4000 celebrities from the VoxCeleb2 dataset for human behavior un…