3 papers
eess.AS2024
PixIT: Joint Training of Speaker Diarization and Speech Separation from Real-world Multi-speaker Recordings
Joonas Kalda, Clément Pagés, Ricard Marxer +2
A major drawback of supervised speech separation (SSep) systems is their reliance on synthetic data, leading to poor real-world generalization. Mixture invariant training (MixIT) w…
eess.AS2024
TalTech-IRIT-LIS Speaker and Language Diarization Systems for DISPLACE 2024
Joonas Kalda, Tanel Alumäe, Martin Lebourdais +3
This paper describes the submissions of team TalTech-IRIT-LIS to the DISPLACE 2024 challenge. Our team participated in the speaker diarization and language diarization tracks of th…
cs.CL2024
Finetuning End-to-End Models for Estonian Conversational Spoken Language Translation
Tiia Sildam, Andra Velve, Tanel Alumäe
This paper investigates the finetuning of end-to-end models for bidirectional Estonian-English and Estonian-Russian conversational speech-to-text translation. Due to the limited av…