1 paper
Haitao Li, Chunxiang Jin, Chenglin Li +3
Zero-shot text-to-speech models can clone a speaker's timbre from a short reference audio, but they also strongly inherit the speaking style present in the reference. As a result,…