collaborators

6 papers

eess.AS2026

Towards Operational Conversational Intelligence: A Speech Intelligence Framework

C. Vishnoi, S. Khurana, A. Timmapur +2

Body-worn camera (BWC) audio presents unique challenges including high ambient noise, variable recording conditions, and multiple overlapping speakers that make automated transcrip…

cs.CL2026

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis

Wasim Madha, Nityanand Mathur, Hamees Sayed +4

Current text-to-speech systems face a trade-off: autoregres- sive codec language models produce highly intelligible speech but require large-scale models and training data and deco…

cs.AI2026

How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech

Nityanand Mathur, Hamees Sayed, Wasim Madha +4

Style-captioned text-to-speech systems use natural language to control voice characteristics, but how individual words influence acoustic output remains unclear. Understanding this…

eess.AS2025

Factorized RVQ-GAN For Disentangled Speech Tokenization

Sameer Khurana, Dominik Klement, Antoine Laurent +13

We propose Hierarchical Audio Codec (HAC), a unified neural speech codec that factorizes its bottleneck into three linguistic levels-acoustic, phonetic, and lexical-within a single…

eess.AS2025

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement

Amir Hussein, Sameer Khurana, Gordon Wichern +2

Effective speech representations for spoken language models must balance semantic relevance with acoustic fidelity for high-quality reconstruction. However, existing approaches str…

cs.SD2025

SMITIN: Self-Monitored Inference-Time INtervention for Generative Music Transformers

Junghyun Koo, Gordon Wichern, Francois G. Germain +2

We introduce Self-Monitored Inference-Time INtervention (SMITIN), an approach for controlling an autoregressive generative music transformer using classifier probes. These simple l…