1 paper
Mahmoud Azab, Mingzhe Wang, Max Smith +3
We propose a new model for speaker naming in movies that leverages visual, textual, and acoustic modalities in an unified optimization framework. To evaluate the performance of our…