26 citations · 32 across the 6 of their papers we have counts for
4 papers · 1 filter
X2-NativeCursor: Native-Token Text Progress Tracking for Incremental-Text Streaming Codec TTS
Zehan Liu, Carl Chen, Rime Wen +7
Incremental-text streaming text-to-speech (TTS) needs online text progress tracking for synchronized highlighting, interruption handling, and dialogue-history updates. Input text a…
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction
Kaiqi Fu, Rime Wen, Altman Lin +4
Accurate and responsive turn-taking is essential for spoken dialogue systems, which must distinguish in real time between user interruptions, backchannels that should be ignored, a…
Pronunciation Assessment with Multi-modal Large Language Models
Kaiqi Fu, Linkai Peng, Nan Yang +1
Large language models (LLMs), renowned for their powerful conversational abilities, are widely recognized as exceptional tools in the field of education, particularly in the contex…
A Full Text-Dependent End to End Mispronunciation Detection and Diagnosis with Easy Data Augmentation Techniques
Kaiqi Fu, Jones Lin, Dengfeng Ke +3
Recently, end-to-end mispronunciation detection and diagnosis (MD&D) systems has become a popular alternative to greatly simplify the model-building process of conventional hybrid…