works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.SD2026

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization

Peijie Chen, Zhuanling Zha, Zhipeng Nie +8

In current zero-shot text-to-speech systems, conventional semantic tokenizers are typically optimized using supervised automatic speech recognition or self-supervised learning obje…

cs.SD2026

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

Weijie Wu, Junbo Li, Lin Li +2

The paper introduces MMAC, a large benchmark of 5,638 audio clips designed to evaluate audio captioning models across multiple capability categories and evaluation dimensions, focu…

cs.SD2026

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations

Peijie Chen, Wenhao Guan, Weijie Wu +7

Zero-shot text-to-speech (TTS) relies on robust speech representations. However, current speech tokenizers face a fundamental trade-off: acoustic codecs preserve high-fidelity audi…

eess.AS2025

SyncVoice: Towards Video Dubbing with Vision-Augmented Pretrained TTS Model

Kaidi Wang, Yi He, Wenhao Guan +8

Video dubbing aims to generate high-fidelity speech that is precisely temporally aligned with the visual content. Existing methods still suffer from limitations in speech naturalne…

eess.AS2025

Phoenix-VAD: Streaming Semantic Endpoint Detection for Full-Duplex Speech Interaction

Weijie Wu, Wenhao Guan, Kaidi Wang +6

Spoken dialogue models have significantly advanced intelligent human-computer interaction, yet they lack a plug-and-play full-duplex prediction module for semantic endpoint detecti…

cs.SD2025

DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec

Peijie Chen, Wenhao Guan, Kaidi Wang +4

Neural speech codecs are essential for advancing text-to-speech (TTS) systems. With the recent success of large language models in text generation, developing high-quality speech t…