15 papers
ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era
Masao Someki, Alexander Polok, Carlos Carvalho +14
Recent speech research involves increasingly large datasets, complex models, and diverse experimental workflows. However, existing frameworks require substantial engineering effort…
Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning
Alexander Polok, Samuele Cornell, Sathvik Udupa +3
We propose diarization-conditioned spoken language models (SLMs), a strategy for extending SLMs to far-field multi-talker audio. Rather than adapting the decoder via Serialized Out…
CS-YODAS: A Mined Dataset of In-the-Wild Code-Switched Speech
Brian Yan, Qingzheng Wang, Matthew Wiesner +9
We present CS-YODAS, a Creative Commons-licensed dataset of in-the-wild code-switched speech mined from multilingual YouTube data. Code-switching (CS), or the alternation between l…
BUT System Description for CHiME-9 MCoRec Challenge
Dominik Klement, Alexander Polok, Nguyen Hai Phong +2
Multi-talker automatic speech recognition (ASR) in conversational recordings remains an open problem, particularly in scenarios with large portion of overlapping speech where ident…
Who Spoke What When? Evaluating Spoken Language Models for Conversational ASR with Semantic and Overlap-Aware Metrics
Naohiro Tawara, Samuele Cornell, Alexander Polok +3
Conversational automatic speech recognition remains challenging due to overlapping speech, far-field noise, and varying speaker counts. While recent LLM-based systems perform well…
Modeling Overlapped Speech with Shuffles
Matthew Wiesner, Samuele Cornell, Alexander Polok +3
We propose to model parallel streams of data, such as overlapped speech, using shuffles. Specifically, this paper shows how the shuffle product and partial order finite-state autom…