4 papers
CoVoMix2: Advancing Zero-Shot Dialogue Generation with Fully Non-Autoregressive Flow Matching
Leying Zhang, Yao Qian, Xiaofei Wang +8
Generating natural-sounding, multi-speaker dialogue is crucial for applications such as podcast creation, virtual agents, and multimedia content generation. However, existing syste…
Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation
Peidong Wang, Naoyuki Kanda, Jian Xue +7
Streaming multi-talker speech translation is a task that involves not only generating accurate and fluent translations with low latency but also recognizing when a speaker change o…
Isochrony-Controlled Speech-to-Text Translation: A study on translating from Sino-Tibetan to Indo-European Languages
Midia Yousefi, Yao Qian, Junkun Chen +5
End-to-end speech translation (ST), which translates source language speech directly into target language text, has garnered significant attention in recent years. Many ST applicat…
Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
Jiaqi Li, Dongmei Wang, Xiaofei Wang +13
Neural audio codec tokens serve as the fundamental building blocks for speech language model (SLM)-based speech generation. However, there is no systematic understanding on how the…