1 paper
Anuj Diwan, Anirudh Srinivasan, David Harwath +1
Existing speech-to-speech translation (S2ST) models fall into two camps: they either leverage text as an intermediate step or require hundreds of hours of parallel speech data. Bot…