6 papers
Towards Operational Conversational Intelligence: A Speech Intelligence Framework
C. Vishnoi, S. Khurana, A. Timmapur +2
Body-worn camera (BWC) audio presents unique challenges including high ambient noise, variable recording conditions, and multiple overlapping speakers that make automated transcrip…
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis
Wasim Madha, Nityanand Mathur, Hamees Sayed +4
Current text-to-speech systems face a trade-off: autoregres- sive codec language models produce highly intelligible speech but require large-scale models and training data and deco…
How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech
Nityanand Mathur, Hamees Sayed, Wasim Madha +4
Style-captioned text-to-speech systems use natural language to control voice characteristics, but how individual words influence acoustic output remains unclear. Understanding this…
Factorized RVQ-GAN For Disentangled Speech Tokenization
Sameer Khurana, Dominik Klement, Antoine Laurent +13
We propose Hierarchical Audio Codec (HAC), a unified neural speech codec that factorizes its bottleneck into three linguistic levels-acoustic, phonetic, and lexical-within a single…
HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement
Amir Hussein, Sameer Khurana, Gordon Wichern +2
Effective speech representations for spoken language models must balance semantic relevance with acoustic fidelity for high-quality reconstruction. However, existing approaches str…
SMITIN: Self-Monitored Inference-Time INtervention for Generative Music Transformers
Junghyun Koo, Gordon Wichern, Francois G. Germain +2
We introduce Self-Monitored Inference-Time INtervention (SMITIN), an approach for controlling an autoregressive generative music transformer using classifier probes. These simple l…