1 paper
Chihiro Arata, Kiyoshi Kurihara
Autoregressive neural codec language models have shown strong zero-shot voice cloning ability, but decoder-only architectures treat input text as a prefix that competes with the gr…