5 citations · 6 across the 12 of their papers we have counts for
8 papers · 1 filter
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation
Haowei Lou, Hye-Young Paik, Dai Jia +2
Human voice generation has made rapid progress in speech generation, singing voice generation, voice cloning, and voice editing. However, most existing systems are designed for spe…
AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling
Haowei Lou, Junda Wu, Chengkai Huang +4
State-of-the-art text-to-speech (TTS) models achieve impressive naturalness and expressiveness, yet fine-grained, disentangled control over speaking styles remains challenging. In…
ParaMETA: Towards Learning Disentangled Paralinguistic Speaking Styles Representations from Speech
Haowei Lou, Hye-young Paik, Wen Hu +1
Learning representative embeddings for different types of speaking styles, such as emotion, age, and gender, is critical for both recognition tasks (e.g., cognitive computing and h…
ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation
Haowei Lou, Hye-Young Paik, Wen Hu +1
Controlling speaking style in text-to-speech (TTS) systems has become a growing focus in both academia and industry. While many existing approaches rely on reference audio to guide…
Generalized Multilingual Text-to-Speech Generation with Language-Aware Style Adaptation
Haowei Lou, Hye-young Paik, Sheng Li +2
Text-to-Speech (TTS) models can generate natural, human-like speech across multiple languages by transforming phonemes into waveforms. However, multilingual TTS remains challenging…
LatentSpeech: Latent Diffusion for Text-To-Speech Generation
Haowei Lou, Helen Paik, Pari Delir Haghighi +2
Diffusion-based Generative AI gains significant attention for its superior performance over other generative techniques like Generative Adversarial Networks and Variational Autoenc…