4.8k citations · 4.8k across the 3 of their papers we have counts for
1 paper · 1 filter
Mahmoud Azab, Mingzhe Wang, Max Smith +3
We propose a new model for speaker naming in movies that leverages visual, textual, and acoustic modalities in an unified optimization framework. To evaluate the performance of our…