From the 1 of 8 linked papers with an AI index.
9 papers
Qwen-Audio-3.0-Gen-Preview Technical Report
Junyu Dai, Xiaoyue Duan, Xinyue Fan +14
The paper introduces Qwen-Audio-3.0-Gen-Preview, a unified non‑autoregressive model that uses a diffusion transformer and a shared VAE to generate complete mixed‑waveform audio fro…
Harness TTS: Towards Context-Aware Expressive Speech Synthesis with Harness Layer
Shengfan Shen, Di Wu, Xingchen Song +5
Expressive speech synthesis for voice assistants requires flexible style control that adapts to explicit requests and broader interaction context. We propose Harness TTS, a lightwe…
F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation
Dinghao Zhou, Xingchen Song, Di Wu +3
Continuous audio autoencoders reconstruct waveforms well but often produce latents with weak structure for understanding, while self-supervised audio encoders capture semantics but…
TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis
Xi Wang, Jie Wang, Xingchen Song +8
While generative text-to-speech (TTS) models approach human-level quality, monolithic metrics fail to diagnose fine-grained acoustic artifacts or explain perceptual collapse. To ad…
Borderless Long Speech Synthesis
Xingchen Song, Di Wu, Dinghao Zhou +12
Most existing text-to-speech (TTS) systems either synthesize speech sentence by sentence and stitch the results together, or drive synthesis from plain-text dialogues alone. Both a…
Iterate to Differentiate: Enhancing Discriminability and Reliability in Zero-Shot TTS Evaluation
Shengfan Shen, Di Wu, Xingchen Song +5
Reliable evaluation of modern zero-shot text-to-speech (TTS) models remains challenging. Subjective tests are costly and hard to reproduce, while objective metrics often saturate,…