works on

From the 1 of 12 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.SDShow all

7 papers · 1 filter

cs.SD2026

AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling

Haowei Lou, Junda Wu, Chengkai Huang +4

AutoSIFT is a text-to-speech framework that separates speaking style into explicit categories (e.g., emotion, age) and residual prosodic details, allowing users to edit specific st…

cs.SD2026

ParaMETA: Towards Learning Disentangled Paralinguistic Speaking Styles Representations from Speech

Haowei Lou, Hye-young Paik, Wen Hu +1

Learning representative embeddings for different types of speaking styles, such as emotion, age, and gender, is critical for both recognition tasks (e.g., cognitive computing and h…

cs.SD2025

ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation

Haowei Lou, Hye-Young Paik, Wen Hu +1

Controlling speaking style in text-to-speech (TTS) systems has become a growing focus in both academia and industry. While many existing approaches rely on reference audio to guide…

cs.SD2025

Generalized Multilingual Text-to-Speech Generation with Language-Aware Style Adaptation

Haowei Lou, Hye-young Paik, Sheng Li +2

Text-to-Speech (TTS) models can generate natural, human-like speech across multiple languages by transforming phonemes into waveforms. However, multilingual TTS remains challenging…

cs.SD2024

LatentSpeech: Latent Diffusion for Text-To-Speech Generation

Haowei Lou, Helen Paik, Pari Delir Haghighi +2

Diffusion-based Generative AI gains significant attention for its superior performance over other generative techniques like Generative Adversarial Networks and Variational Autoenc…

cs.SD2024

Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration

Haowei Lou, Helen Paik, Wen Hu +1

Recent advancements in text-to-speech (TTS) systems, such as FastSpeech and StyleSpeech, have significantly improved speech generation quality. However, these models often rely on…