2 papers
eess.AS2025
Lightweight Target-Speaker-Based Overlap Transcription for Practical Streaming ASR
Aleš Pražák, Marie Kunešová, Josef Psutka
Overlapping speech remains a major challenge for automatic speech recognition (ASR) in real-world applications, particularly in broadcast media with dynamic, multi-speaker interact…
cs.CL2024
A Comparative Analysis of Bilingual and Trilingual Wav2Vec Models for Automatic Speech Recognition in Multilingual Oral History Archives
Jan LeheÄka, Josef V. Psutka, LuboÅ¡ Å mÃdl +2
In this paper, we are comparing monolingual Wav2Vec 2.0 models with various multilingual models to see whether we could improve speech recognition performance on a unique oral hist…