6 papers
Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization
Jiangyu Han, Federico Landini, Johan Rohdin +4
Self-supervised learning (SSL) models like WavLM can be effectively utilized when building speaker diarization systems but are often large and slow, limiting their use in resource…
Joint Training of Speaker Embedding Extractor, Speech and Overlap Detection for Diarization
Petr Pálka, Federico Landini, Dominik Klement +4
In spite of the popularity of end-to-end diarization systems nowadays, modular systems comprised of voice activity detection (VAD), speaker embedding extraction plus clustering, an…
Leveraging Self-Supervised Learning for Speaker Diarization
Jiangyu Han, Federico Landini, Johan Rohdin +3
End-to-end neural diarization has evolved considerably over the past few years, but data scarcity is still a major obstacle for further improvements. Self-supervised learning metho…
Do End-to-End Neural Diarization Attractors Need to Encode Speaker Characteristic Information?
Lin Zhang, Themos Stafylakis, Federico Landini +3
In this paper, we apply the variational information bottleneck approach to end-to-end neural diarization with encoder-decoder attractors (EEND-EDA). This allows us to investigate w…
Spoof Diarization: "What Spoofed When" in Partially Spoofed Audio
Lin Zhang, Xin Wang, Erica Cooper +4
This paper defines Spoof Diarization as a novel task in the Partial Spoof (PS) scenario. It aims to determine what spoofed when, which includes not only locating spoof regions but…
DiaPer: End-to-End Neural Diarization with Perceiver-Based Attractors
Federico Landini, Mireia Diez, Themos Stafylakis +1
Until recently, the field of speaker diarization was dominated by cascaded systems. Due to their limitations, mainly regarding overlapped speech and cumbersome pipelines, end-to-en…