activity
20182026
most citedATCO2 corpus: A Large-Scale Dataset for Research on Automatic Speech Recognition and Natural Language Understanding of Air Traffic Control Communications

14 citations · 35 across the 25 of their papers we have counts for

collaborators
Showing 2025Show all

8 papers · 1 filter

eess.AS2025

BUT Systems for Environmental Sound Deepfake Detection in the ESDD 2026 Challenge

Junyi Peng, Lin Zhang, Jin Li +2

This paper describes the BUT submission to the ESDD 2026 Challenge, specifically focusing on Track 1: Environmental Sound Deepfake Detection with Unseen Generators. To address the…

cs.SD2025

ORCA: Open-ended Response Correctness Assessment for Audio Question Answering

Šimon Sedláček, Sara Barahona, Bolaji Yusuf +9

Reliable assessment of the abilities of large audio language models (LALMs) is essential to advancing the state of the art. As benchmarks rapidly evolve to incorporate complex reas…

cs.CL2025

Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking

Katia Vendrame, Bolaji Yusuf, Santosh Kesiraju +3

End-to-end spoken dialogue state tracking (DST) is made difficult by the tandem of having to handle speech input and data scarcity. Combining speech foundation encoders and large l…

eess.AS2025

Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition

Martin Kocour, Martin Karafiat, Alexander Polok +3

We propose a speaker-attributed (SA) Whisper-based model for multi-talker speech recognition that combines target-speaker modeling with serialized output training (SOT). Our approa…

eess.AS2025

Unsupervised Speech Enhancement using Data-defined Priors

Dominik Klement, Matthew Maciejewski, Sanjeev Khudanpur +2

The majority of deep learning-based speech enhancement methods require paired clean-noisy speech data. Collecting such data at scale in real-world conditions is infeasible, which h…

cs.SD2025

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training

Sathvik Udupa, Shinji Watanabe, Petr Schwarz +1

Accurate, low-latency endpointing is crucial for effective spoken dialogue systems. While traditional endpointers often rely on spectrum-based audio features, this work proposes re…