collaborators

6 papers

cs.CL2026

HK-LegiCoST: Leveraging Non-Verbatim Transcripts for Speech Translation

Cihan Xiao, Henry Li Xinyuan, Jinyi Yang +4

We introduce HK-LegiCoST, a new three-way parallel corpus of Cantonese-English translations, containing 600+ hours of Cantonese audio, its standard traditional Chinese transcript,…

eess.AS2026

Universal Speech Content Factorization

Henry Li Xinyuan, Zexin Cai, Lin Zhang +5

We propose Universal Speech Content Factorization (USCF), a simple and invertible linear method for extracting a low-rank speech representation in which speaker timbre is suppresse…

eess.AS2025

GenVC: Self-Supervised Zero-Shot Voice Conversion

Zexin Cai, Henry Li Xinyuan, Ashi Garg +5

Most current zero-shot voice conversion methods rely on externally supervised components, particularly speaker encoders, for training. To explore alternatives that eliminate this d…

eess.AS2025

Rapidly Adapting to New Voice Spoofing: Few-Shot Detection of Synthesized Speech Under Distribution Shifts

Ashi Garg, Zexin Cai, Henry Li Xinyuan +5

We address the challenge of detecting synthesized speech under distribution shifts -- arising from unseen synthesis methods, speakers, languages, or audio conditions -- relative to…

eess.AS2025

Scalable Controllable Accented TTS

Henry Li Xinyuan, Zexin Cai, Ashi Garg +5

We tackle the challenge of scaling accented TTS systems, expanding their capabilities to include much larger amounts of training data and a wider variety of accent labels, even for…

eess.AS2025

ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts

Ashi Garg, Zexin Cai, Lin Zhang +6

The problem of synthetic speech detection has enjoyed considerable attention, with recent methods achieving low error rates across several established benchmarks. However, to what…