6 papers
Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge
Aivo Olev, Tanel Alumäe
This paper describes TalTech's submissions to the Beyond Transcription Challenge (BeTraC), which requires generating SOAP notes directly from long doctor-patient conversation recor…
Multi-Source Evidence Fusion for Audio Question Answering
Aivo Olev, Tanel Alumäe
Large audio language models (LALMs) can answer questions about speech, music, and environmental sounds, yet their internal reasoning is largely opaque and difficult to validate. We…
Estonian Native Large Language Model Benchmark
Helena Grete Lillepalu, Tanel Alumäe
The availability of LLM benchmarks for the Estonian language is limited, and a comprehensive evaluation comparing the performance of different LLMs on Estonian tasks has yet to be…
Comparison of End-to-end Speech Assessment Models for the NOCASA 2025 Challenge
Aleksei Žavoronkov, Tanel Alumäe
This paper presents an analysis of three end-to-end models developed for the NOCASA 2025 Challenge, aimed at automatic word-level pronunciation assessment for children learning Nor…
TalTech Systems for the Interspeech 2025 ML-SUPERB 2.0 Challenge
Tanel Alumäe, Artem Fedorchenko
This paper describes the language identification and multilingual speech recognition system developed at Tallinn University of Technology for the Interspeech 2025 ML-SUPERB 2.0 Cha…
Optimizing Estonian TV Subtitles with Semi-supervised Learning and LLMs
Artem Fedorchenko, Tanel Alumäe
This paper presents an approach for generating high-quality, same-language subtitles for Estonian TV content. We fine-tune the Whisper model on human-generated Estonian subtitles a…