2 papers
eess.AS2026
Endpoint Anticipation for Low-Latency Spoken Dialogue
Sathvik Udupa, Shinji Watanabe, Petr Schwarz +1
While low-latency interaction is critical for spoken dialogue, cascaded architectures are often bottlenecked by reactive turn-completion detection. We propose Endpoint Anticipation…
eess.AS2025
Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models
Jiangyu Han, Petr Pálka, Marc Delcroix +4
Self-supervised learning (SSL) models such as WavLM have substantially advanced speaker diarization by providing rich contextual speech representations. However, the high computati…