6 papers
The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge
Shikhar Bharadwaj, Samuele Cornell, Kwanghee Choi +4
This technical report describes our submission to the ICME 2025 audio encoder challenge. Our submitted system is built on BEATs, a masked speech token prediction based audio encode…
OpusLM: A Family of Open Unified Speech Language Models
Jinchuan Tian, William Chen, Yifan Peng +9
This paper presents Open Unified Speech Language Models (OpusLMs), a family of open foundational speech language models (SpeechLMs) up to 7B. Initialized from decoder-only text lan…
EmoNews: A Spoken Dialogue System for Expressive News Conversations
Ryuki Matsuura, Shikhar Bharadwaj, Jiarui Liu +1
We develop a task-oriented spoken dialogue system (SDS) that regulates emotional speech based on contextual cues to enable more empathetic news conversations. Despite advancements…
Context-Driven Dynamic Pruning for Large Speech Foundation Models
Masao Someki, Shikhar Bharadwaj, Atharva Anand Joshi +7
Speech foundation models achieve strong generalization across languages and acoustic conditions, but require significant computational resources for inference. In the context of sp…
Multimodal Modeling For Spoken Language Identification
Shikhar Bharadwaj, Min Ma, Shikhar Vashishth +10
Spoken language identification refers to the task of automatically predicting the spoken language in a given utterance. Conventionally, it is modeled as a speech-based language ide…
Label Aware Speech Representation Learning For Language Identification
Shikhar Vashishth, Shikhar Bharadwaj, Sriram Ganapathy +5
Speech representation learning approaches for non-semantic tasks such as language recognition have either explored supervised embedding extraction methods using a classifier model…