3 papers
eess.AS2026
Doctor or Patient? Synergizing Diarization and ASR for Code-Switched Hinglish Medical Conditions Extraction
Séverin Baroudi, Yanis Labrak, Shashi Kumar +7
Extracting patient medical conditions from code-switched clinical spoken dialogues is challenging due to rapid turn-taking and highly overlapped speech. We present a robust system…
cs.SD2025
Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models
Sandipana Dowerah, Atharva Kulkarni, Ajinkya Kulkarni +7
Parallel to the development of advanced deepfake audio generation, audio deepfake detection has also seen significant progress. However, a standardized and comprehensive benchmark…
eess.AS2024
PixIT: Joint Training of Speaker Diarization and Speech Separation from Real-world Multi-speaker Recordings
Joonas Kalda, Clément Pagés, Ricard Marxer +2
A major drawback of supervised speech separation (SSep) systems is their reliance on synthetic data, leading to poor real-world generalization. Mixture invariant training (MixIT) w…