collaborators

6 papers

cs.CL2026

Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge

Aivo Olev, Tanel Alumäe

This paper describes TalTech's submissions to the Beyond Transcription Challenge (BeTraC), which requires generating SOAP notes directly from long doctor-patient conversation recor…

eess.AS2026

Multi-Source Evidence Fusion for Audio Question Answering

Aivo Olev, Tanel Alumäe

Large audio language models (LALMs) can answer questions about speech, music, and environmental sounds, yet their internal reasoning is largely opaque and difficult to validate. We…

cs.CL2026

Estonian Native Large Language Model Benchmark

Helena Grete Lillepalu, Tanel Alumäe

The availability of LLM benchmarks for the Estonian language is limited, and a comprehensive evaluation comparing the performance of different LLMs on Estonian tasks has yet to be…

cs.CL2025

Comparison of End-to-end Speech Assessment Models for the NOCASA 2025 Challenge

Aleksei Žavoronkov, Tanel Alumäe

This paper presents an analysis of three end-to-end models developed for the NOCASA 2025 Challenge, aimed at automatic word-level pronunciation assessment for children learning Nor…

cs.CL2025

TalTech Systems for the Interspeech 2025 ML-SUPERB 2.0 Challenge

Tanel Alumäe, Artem Fedorchenko

This paper describes the language identification and multilingual speech recognition system developed at Tallinn University of Technology for the Interspeech 2025 ML-SUPERB 2.0 Cha…

cs.CL2025

Optimizing Estonian TV Subtitles with Semi-supervised Learning and LLMs

Artem Fedorchenko, Tanel Alumäe

This paper presents an approach for generating high-quality, same-language subtitles for Estonian TV content. We fine-tune the Whisper model on human-generated Estonian subtitles a…