4 papers
SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity
Hanke Xie, Haopeng Lin, Wenxiao Cao +17
Recent advances in text-to-speech (TTS) synthesis have significantly improved speech expressiveness and naturalness. However, most existing systems are tailored for single-speaker…
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
He Wang, Linhan Ma, Dake Guo +4
Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluati…
StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding
Dake Guo, Jixun Yao, Linhan Ma +2
Recent advancements in discrete token-based speech generation have highlighted the importance of token-to-waveform generation for audio quality, particularly in real-time interacti…
FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech
Linhan Ma, Dake Guo, He Wang +2
Current speech generation research can be categorized into two primary classes: non-autoregressive and autoregressive. The fundamental distinction between these approaches lies in…