11 papers
ORCA: Open-ended Response Correctness Assessment for Audio Question Answering
Å imon SedláÄek, Sara Barahona, Bolaji Yusuf +9
Reliable assessment of the abilities of large audio language models (LALMs) is essential to advancing the state of the art. As benchmarks rapidly evolve to incorporate complex reas…
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
Katia Vendrame, Bolaji Yusuf, Santosh Kesiraju +3
End-to-end spoken dialogue state tracking (DST) is made difficult by the tandem of having to handle speech input and data scarcity. Combining speech foundation encoders and large l…
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
Martin Kocour, Martin Karafiat, Alexander Polok +3
We propose a speaker-attributed (SA) Whisper-based model for multi-talker speech recognition that combines target-speaker modeling with serialized output training (SOT). Our approa…
SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper
Alexander Polok, Dominik Klement, Samuele Cornell +4
Speaker-attributed automatic speech recognition (ASR) in multi-speaker environments remains a major challenge. While some approaches achieve strong performance when fine-tuned on s…
BUT Systems for Environmental Sound Deepfake Detection in the ESDD 2026 Challenge
Junyi Peng, Lin Zhang, Jin Li +2
This paper describes the BUT submission to the ESDD 2026 Challenge, specifically focusing on Track 1: Environmental Sound Deepfake Detection with Unseen Generators. To address the…
Unsupervised Speech Enhancement using Data-defined Priors
Dominik Klement, Matthew Maciejewski, Sanjeev Khudanpur +2
The majority of deep learning-based speech enhancement methods require paired clean-noisy speech data. Collecting such data at scale in real-world conditions is infeasible, which h…