collaborators

17 papers

eess.AS2026

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis

Zhisheng Zheng, Xiaohang Sun, Zhu Liu +5

Recent Text-To-Speech (TTS) systems have achieved strong naturalness and zero-shot voice cloning performance, but fine-grained control of expressive speech at the word or phoneme l…

eess.AS2026

Phone Segmentation and Recognition through Phonological Activation Mapping

Shikhar Bharadwaj, Kwanghee Choi, Stephen McIntosh +8

Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the re…

eess.AS2026

RosettaSpeech: Zero-Shot Speech-to-Speech Translation without Parallel Speech

Zhisheng Zheng, Xiaohang Sun, Tuan Dinh +8

End-to-end speech-to-speech translation (S2ST) systems typically struggle with a critical data bottleneck: the scarcity of parallel speech-to-speech corpora. To overcome this, we i…

cs.CL2026

Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration

Ryan Soh-Eun Shim, Kwanghee Choi, Kalvin Chang +8

Multilingual speech foundation models such as Whisper are trained on web-scale data, where data for each language consists of a myriad of regional varieties. However, different reg…

cs.CV2025

FacEDiT: Unified Talking Face Editing and Generation via Facial Motion Infilling

Kim Sung-Bin, Joohyun Chang, David Harwath +1

Talking face editing and face generation have often been studied as distinct problems. In this work, we propose viewing both not as separate tasks but as subtasks of a unifying for…

eess.AS2025

Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs

Wei-Cheng Tseng, David Harwath

Recent advancements in neural audio codecs have not only enabled superior audio compression but also enhanced speech synthesis techniques. Researchers are now exploring their poten…