1 citations · 1 across the 1 of their papers we have counts for
1 paper
Minh Tran, Mohammad Soleymani
In this paper, we introduce a pretrained audio-visual Transformer trained on more than 500k utterances from nearly 4000 celebrities from the VoxCeleb2 dataset for human behavior un…