2 papers
eess.AS2025
On the Use of Self-Supervised Representation Learning for Speaker Diarization and Separation
Séverin Baroudi, Hervé Bredin, Joseph Razik +1
Self-supervised speech models such as wav2vec2.0 and WavLM have been shown to significantly improve the performance of many downstream speech tasks, especially in low-resource sett…
eess.AS2024
PixIT: Joint Training of Speaker Diarization and Speech Separation from Real-world Multi-speaker Recordings
Joonas Kalda, Clément Pagés, Ricard Marxer +2
A major drawback of supervised speech separation (SSep) systems is their reliance on synthetic data, leading to poor real-world generalization. Mixture invariant training (MixIT) w…