collaborators

5 papers

eess.AS2025

Factorized RVQ-GAN For Disentangled Speech Tokenization

Sameer Khurana, Dominik Klement, Antoine Laurent +13

We propose Hierarchical Audio Codec (HAC), a unified neural speech codec that factorizes its bottleneck into three linguistic levels-acoustic, phonetic, and lexical-within a single…

cs.CL2025

HENT-SRT: Hierarchical Efficient Neural Transducer with Self-Distillation for Joint Speech Recognition and Translation

Amir Hussein, Cihan Xiao, Matthew Wiesner +3

Neural transducers (NT) provide an effective framework for speech streaming, demonstrating strong performance in automatic speech recognition (ASR). However, the application of NT…

eess.AS2025

HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement

Amir Hussein, Sameer Khurana, Gordon Wichern +2

Effective speech representations for spoken language models must balance semantic relevance with acoustic fidelity for high-quality reconstruction. However, existing approaches str…

cs.CL2023

Enhancing End-to-End Conversational Speech Translation Through Target Language Context Utilization

Amir Hussein, Brian Yan, Antonios Anastasopoulos +2

Incorporating longer context has been shown to benefit machine translation, but the inclusion of context in end-to-end speech translation (E2E-ST) remains under-studied. To bridge…

cs.SD2023

Speech collage: code-switched audio generation by collaging monolingual corpora

Amir Hussein, Dorsa Zeinali, Ondřej Klejch +6

Designing effective automatic speech recognition (ASR) systems for Code-Switching (CS) often depends on the availability of the transcribed CS resources. To address data scarcity,…