1 paper
Xiaoqun Liu, Tanu Mitra, Harshit Rajgarhia +1
Speech-to-speech (S2S) models now run inside dubbing, translation, and voice agents. Unlike text models, they hear the speaker's voice, which carries the speaker's gender. A faithf…