6 citations · 6 across the 2 of their papers we have counts for
5 papers
Lightweight Target-Speaker-Based Overlap Transcription for Practical Streaming ASR
Aleš Pražák, Marie Kunešová, Josef Psutka
Overlapping speech remains a major challenge for automatic speech recognition (ASR) in real-world applications, particularly in broadcast media with dynamic, multi-speaker interact…
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS
Marie Kunešová, Zdeněk Hanzlíček, Jindřich Matoušek
Zero-shot multi-speaker text-to-speech (TTS) systems rely on speaker embeddings to synthesize speech in the voice of an unseen speaker, using only a short reference utterance. Whil…
Quality Assessment of Noisy and Enhanced Speech with Limited Data: UWB-NTIS System for VoiceMOS 2024
Marie Kunešová, Aleš Pražák, Jan Lehečka
We present a system for non-intrusive prediction of speech quality in noisy and enhanced speech, developed for Track 3 of the VoiceMOS 2024 Challenge. The task required estimating…
Detection of Prosodic Boundaries in Speech Using Wav2Vec 2.0
Marie Kunešová, Markéta Řezáčková
Prosodic boundaries in speech are of great relevance to both speech synthesis and audio annotation. In this paper, we apply the wav2vec 2.0 framework to the task of detecting these…
UWB-NTIS Speaker Diarization System for the DIHARD II 2019 Challenge
Zbyněk Zajíc, Marie Kunešová, Marek Hrúz +1
In this paper, we present our system developed by the team from the New Technologies for the Information Society (NTIS) research center of the University of West Bohemia in Pilsen,…