audio reasoning 1audio representation learning 1language sensitivity 1masked token prediction 1multi-domain pretraining 1multilingual 1neural audio codecs 1open-source models 1self-supervised learning 1speech representation 1
From the 2 of 23 linked papers with an AI index.
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Who Spoke What When? Evaluating Spoken Language Models for Conversational ASR with Semantic and Overlap-Aware Metrics
Naohiro Tawara, Samuele Cornell, Alexander Polok +3
Conversational automatic speech recognition remains challenging due to overlapping speech, far-field noise, and varying speaker counts. While recent LLM-based systems perform well…
cs.CL2025
CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
Brian Yan, Injy Hamed, Shuichiro Shimizu +24
We present CS-FLEURS, a new dataset for developing and evaluating code-switched speech recognition and translation systems beyond high-resourced languages. CS-FLEURS consists of 4…
cs.CL2025
ESPnet-SpeechLM: An Open Speech Language Model Toolkit
Jinchuan Tian, Jiatong Shi, William Chen +13
We present ESPnet-SpeechLM, an open toolkit designed to democratize the development of speech language models (SpeechLMs) and voice-driven agentic applications. The toolkit standar…