4 papers
From Oracle to Noisy Context: Mitigating Contextual Exposure Bias in Speech-LLMs
Xiaoyong Guo, Nanjie Li, Zijie Zeng +4
Contextual automatic speech recognition (ASR) with Speech-LLMs is typically trained with oracle conversation history, but relies on error-prone history at inference, causing a trai…
Phoneme-based speech recognition driven by large language models and sampling marginalization
Te Ma, Nanjie Li, Hao Huang +1
Recently, the Large Language Model-based Phoneme-to-Grapheme (LLM-P2G) method has shown excellent performance in speech recognition tasks and has become a feasible direction to rep…
Kling-Foley: Multimodal Diffusion Transformer for High-Quality Video-to-Audio Generation
Jun Wang, Xijuan Zeng, Chunyu Qiang +19
We propose Kling-Foley, a large-scale multimodal Video-to-Audio generation model that synthesizes high-quality audio synchronized with video content. In Kling-Foley, we introduce m…
Towards Flow-Matching-based TTS without Classifier-Free Guidance
Yuzhe Liang, Wenzhe Liu, Chunyu Qiang +7
Flow matching has demonstrated strong generative capabilities and has become a core component in modern Text-to-Speech (TTS) systems. To ensure high-quality speech synthesis, Class…